System

A system for tracking individuals in security camera footage by facial and clothing recognition, behavioral analysis, and police database verification addresses the challenge of efficient video data analysis, enabling rapid and accurate investigations.

JP2026022353APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123870
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Efficiently tracking specific individuals from vast amounts of video data and analyzing their behavioral patterns in security camera footage is technically challenging, requiring significant time and effort, which hinders rapid and accurate investigations.

Method used

A system that recognizes faces and extracts clothing features from photographic data, tracks individuals using a security camera video database, records timestamps and camera locations, analyzes behavioral patterns, and verifies identity by comparing with a police database, presenting comprehensive analysis results.

Benefits of technology

Enables accurate and rapid tracking and analysis of individuals by extracting facial and clothing features, analyzing behavioral patterns, and confirming identity, facilitating timely and efficient crime prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022353000001_ABST
    Figure 2026022353000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for facial recognition; means for clothes feature extraction; means for accessing a security camera video database; means for retrieving video based on identified feature data; means for recording timestamps and camera locations; means for analyzing behavioral patterns; means for verifying identity against a police database; and means for generating and presenting analysis results as a report.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, security camera footage has played an important role in crime prevention and investigation. However, efficiently tracking specific individuals from vast amounts of video data and analyzing their behavioral patterns is technically challenging and requires a huge amount of time and effort. Under these circumstances, rapid and accurate investigations are difficult to carry out, which can hinder efforts to uncover and prevent crimes. Therefore, there is a need for technology that can efficiently track specific individuals via security cameras, analyze their behavior, and enable accurate identification. [Means for solving the problem]

[0005] The present invention provides a system that recognizes faces and extracts clothing features from photographic data and tracks target individuals in conjunction with a security camera video database. Specifically, feature data is generated using a face recognition unit and a clothing feature extraction unit, and the security camera video database is accessed based on this data to search for footage. Based on the identified feature data, timestamps and camera locations are recorded and behavioral patterns are analyzed. Furthermore, the analysis results are compared with a police database to confirm the individual's identity. Finally, the analysis results are presented to the user as a comprehensive report, enabling accurate and rapid tracking and analysis.

[0006] "Facial recognition means" refers to a system or algorithm that extracts specific facial feature points from photographic data and matches them with people in a video database.

[0007] The "clothing feature extraction means" is a system or algorithm that analyzes clothing features such as color, pattern, and shape from photographic data and matches them with images in the database.

[0008] A "security camera video database" is a server or system for storing and managing video data obtained from multiple security cameras.

[0009] "Means for searching for video based on feature data" refers to a system or algorithm for analyzing and searching for video in a security camera video database based on extracted facial feature data and clothing feature data.

[0010] A "timestamp" is data that indicates the date and time when a particular video frame was recorded.

[0011] "Camera location" is data indicating the installation location or location information of the security camera that captured the video.

[0012] A "means for analyzing behavioral patterns" is a system or algorithm that uses multiple timestamps and camera location information to analyze the behavioral characteristics of a specific person, such as their route of travel and length of stay.

[0013] A "police database" is a system for storing and managing personal identification information managed by law enforcement agencies.

[0014] "Means for generating and presenting analysis results as a report" means a system or algorithm for organizing the results of the behavioral analysis and the results of the identity verification and presenting them to the user in a visual or digital format. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The present invention is a system that extracts facial recognition and clothing characteristics from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and further verifies their identity in cooperation with a police database. Specific embodiments of the present invention are described in detail below.

[0037] System Overview

[0038] The system consists of the following main components:

[0039] 1. User terminal: The interface where the user inputs photo data and receives the analysis results.

[0040] 2. Server: Responsible for central processing such as data processing, facial recognition, behavioral analysis, and database matching.

[0041] 3. Security camera video database: A database that stores video data collected from multiple security cameras.

[0042] 4. Police databases: Databases containing personally identifiable information.

[0043] System operation and program processing

[0044] 1. Data Entry

[0045] The user uploads a photo of the individual to be tracked to the device. This photo data includes both the face and clothing. The device then sends the photo data to the server.

[0046] 2. Pretreatment

[0047] The server receives this photo data and first extracts the facial region using a facial recognition algorithm, from which it identifies the following feature points:

[0048] Eye, nose, and mouth position

[0049] Facial contours

[0050] Other features (mole, scar, etc.)

[0051] At the same time, the server extracts the clothing's characteristics, specifically analyzing its color, pattern, shape, etc., and generates feature data.

[0052] 3. Video search and feature data matching

[0053] The server accesses a security camera video database and searches for video data based on the identified feature data. From the security camera video, it performs facial recognition and clothing feature matching to identify matching individuals.

[0054] 4. Recording timestamps and camera positions

[0055] For each frame in which a matching person is identified, the server records the timestamp and camera position, allowing the path of the person to be traced.

[0056] 5. Behavioral Pattern Analysis

[0057] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, the following information is generated:

[0058] Travel route

[0059] Stay time

[0060] Frequently visited places

[0061] This information is organized as a result of behavioral pattern analysis.

[0062] 6. Identity Verification

[0063] After the behavioral pattern analysis is completed, the server will compare the target person's facial feature data with the police database, and based on the matching results, the target person's identity will be confirmed.

[0064] 7. Report generation and presentation

[0065] Finally, the server generates a detailed report containing the results of the analysis and identification, including the following information:

[0066] Travel route

[0067] Behavioral patterns

[0068] Timestamp and camera position

[0069] Identity verification results

[0070] The terminal presents this report to the user and provides the necessary information.

[0071] Specific examples

[0072] For example, to track a suspicious person in a public facility, a user uploads a photo of the suspicious person to the system. The server extracts facial and clothing features from the photo and accesses a security camera footage database to detect the person. The server then analyzes the person's behavioral patterns to identify their route of travel and where they are staying. Finally, the server checks the data against a police database to confirm their identity, and a detailed report is provided to the user. This allows for accurate and rapid tracking of people.

[0073] The processing flow will be explained below.

[0074] Step 1:

[0075] The user uploads a photo of the individual to be tracked to the device, which receives and stores the photo data.

[0076] Step 2:

[0077] The device sends the photo data to the server, which receives it and applies a facial recognition algorithm.

[0078] Step 3:

[0079] The server uses a facial recognition algorithm to extract facial regions from the photograph, identify specific feature points (such as the eyes, nose, and mouth) from the facial regions, and generate facial feature data.

[0080] Step 4:

[0081] At the same time, the server analyzes the clothing characteristics (color, pattern, shape, etc.) from the photograph data and generates clothing characteristic data.

[0082] Step 5:

[0083] The server accesses the security camera video database and obtains the latest video data.

[0084] Step 6:

[0085] The server performs facial recognition and clothing feature matching on the acquired video data to search for people who match specific feature data.

[0086] Step 7:

[0087] The server records the timestamp and camera position for each frame in which a matching person is identified.

[0088] Step 8:

[0089] The server uses multiple timestamps and camera location information to analyze the target person's behavioral characteristics, such as their route of travel and length of stay.

[0090] Step 9:

[0091] The server organizes the analysis results and extracts behavioral patterns (frequently visited places, places where people spend long periods of time, etc.).

[0092] Step 10:

[0093] The server uses the analyzed facial feature data to access police databases and verify the identity of the person in question.

[0094] Step 11:

[0095] The server adds the matching results to the report and combines all the analysis results to generate the final report.

[0096] Step 12:

[0097] The terminal presents the final report to the user, who then checks the report and obtains the necessary information.

[0098] Example 1

[0099] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0100] In modern society, numerous security cameras and surveillance devices are installed, helping to prevent and investigate crime. However, quickly and accurately identifying and tracking specific individuals from large amounts of video data is difficult, time-consuming, and laborious. To solve this problem, there is a need for a system that can extract facial recognition and clothing characteristics from photographic data, track and analyze the behavior of target individuals using a security camera video database, and further verify their identity in conjunction with an identification database.

[0101] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0102] In this invention, the server includes a face recognition means, a clothing feature extraction means, a means for accessing a video database, a means for searching video data based on the identified feature data, a means for recording timestamps and location information, a means for analyzing behavioral patterns, a means for verifying identity by comparing with an identification database, a means for generating and presenting analysis results as a report, a user interface including a means for uploading photo data, and a means for using a generative AI model as part of the analysis process. This makes it possible to quickly and accurately identify a target person from photo data, analyze the person's behavioral patterns, and verify the person's identity in cooperation with the identification database.

[0103] "Facial recognition means" refers to a device or program that uses technology or algorithms to detect facial feature points and identify individuals.

[0104] "Clothing feature extraction means" refers to a device or program that uses technology or algorithms to analyze and extract clothing features such as color, pattern, and shape from image data.

[0105] "Means for accessing a video database" refers to a device or program that connects to a database in which video data collected from multiple cameras is stored and acquires the necessary data.

[0106] "Means for searching video data based on identified feature data" refers to devices or programs that use technology or algorithms to search for matching people from videos in a video database using extracted facial and clothing feature data.

[0107] The "means for recording timestamp and location information" refers to a device or program that acquires and records the shooting date and time of the searched video frame and the installation location of the camera.

[0108] "Means for analyzing behavioral patterns" refers to devices or programs that use technologies or algorithms to analyze a target person's behavioral patterns, such as their travel routes, length of stay, and frequently visited places, based on recorded timestamps and location information.

[0109] "Means for verifying identity by matching against an identification database" refers to devices or programs that use technology or algorithms to compare analyzed facial feature information with an identification database and verify the identity of the person based on matching data.

[0110] The "means for generating and presenting the analysis results as a report" refers to a device or program that organizes the results of behavioral pattern analysis and identity verification, and generates a report that can be visually presented to the user.

[0111] The "user interface including a means for uploading photographic data" is an interface that allows a user to input photographic data into the system via a terminal.

[0112] "Means of using a generative AI model as part of the analytical process" refers to devices or programs that utilize generative AI models as part of the analysis, and use technologies or algorithms to achieve more advanced data processing and pattern recognition.

[0113] The present invention is a system that extracts facial recognition and clothing characteristics from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and verifies their identity in cooperation with an identification database. To implement this system, the following hardware and software are used.

[0114] System Overview

[0115] The system consists of the following main components:

[0116] 1. User terminal: An interface where the user inputs photo data and receives the analysis results.

[0117] 2. Server: Responsible for central processing such as data processing, facial recognition, behavior analysis, and database matching.

[0118] 3. Security camera video database: A database that stores video data collected from multiple security cameras.

[0119] 4. Identification Database: A database containing personally identifiable information.

[0120] Hardware and software used

[0121] Face recognition algorithm: Uses OpenCV and dlib libraries.

[0122] Clothing feature extraction algorithm: Uses color analysis using RGB values, texture analysis, and contour detection algorithms.

[0123] Video search algorithm: Facenet and DeepFace are used for face recognition, and Elasticsearch's image search plugin is used for clothing matching.

[0124] Behavioral analysis algorithms: Use custom algorithms to plot movement paths and analyze dwell times.

[0125] Identity database matching algorithm: Uses face identification model to match with identity database.

[0126] Detailed procedure

[0127] The user uploads photo data of the object to be tracked to the device, which then sends the uploaded photo data to the server. The server then uses a facial recognition algorithm to detect the facial area and extract feature points such as the eyes, nose, mouth, facial contours, moles, and scars.

[0128] At the same time, the server analyzes the color, pattern, and shape of the clothing to generate feature data. Based on the analyzed feature data, it accesses a security camera video database to search for matching video data. The server then obtains and records the timestamp and camera position of the video frame of the matching person.

[0129] The system then analyzes the person's behavioral patterns based on timestamps and camera location data to identify their route and location. Finally, the system compares the facial feature data with an identification database to confirm their identity. The server then processes the analysis results and generates a detailed report that is presented to the user's device.

[0130] Specific examples

[0131] For example, when tracking a suspicious person in a public facility, the user uploads photo data to the system. The server extracts facial and clothing features, searches a security camera footage database, and analyzes the suspicious person's behavioral patterns. The server then verifies the person's identity against an identification database and provides a detailed report to the user.

[0132] Prompt Sentence Examples

[0133] "Using this photo, please briefly explain the entire process of identifying the person in question by extracting facial and clothing characteristics, searching a security camera footage database, and analyzing their behavioral patterns to confirm their identity."

[0134] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0135] Step 1: Data entry

[0136] The user uploads photographic data of the individual to be tracked to the device, which includes information on their face and clothing.

[0137] Input: Photo data (face and clothing information)

[0138] Output: Photo data sent to the server

[0139] Specifically, the user selects photo data using the file upload function of the device and clicks the upload button. The device then sends the selected photo data to the server as an HTTP request.

[0140] Step 2: Face recognition and clothing feature extraction (preprocessing)

[0141] The server receives this photo data and extracts the face region using a facial recognition algorithm, for example, using the face detection function of OpenCV or dlib.

[0142] Input: Submitted photo data

[0143] Output: Facial feature points (eyes, nose, mouth position, facial contours, moles, scars, etc.)

[0144] Specifically, the server reads the photo data using an image processing library, detects the facial area, and extracts feature points.

[0145] At the same time, the server extracts the clothing's features by analyzing its color (RGB value), pattern (texture analysis), and shape (contour detection) to generate feature data.

[0146] Input: Submitted photo data

[0147] Output: Clothing feature data (color, pattern, shape)

[0148] Specifically, the server extracts clothing feature data using a color analysis algorithm, a texture analysis algorithm, and a shape analysis algorithm.

[0149] Step 3: Feature data matching and video search

[0150] The server then accesses a security camera video database based on the identified facial and clothing feature data and searches for the relevant video data, using Facenet and DeepFace for facial recognition and Elasticsearch's image search plugin for clothing matching.

[0151] Input: facial feature points, clothing feature data

[0152] Output: Video data of the matching person

[0153] Specifically, the server queries a security camera video database to search for video frames that match the facial features and clothing data.

[0154] Step 4: Record timestamps and camera positions

[0155] For each frame in which a matching person is identified, the server captures and records the timestamp and camera position from the video data.

[0156] Input: Video data of the person to match

[0157] Output: timestamp and camera position information

[0158] Specifically, the server extracts and records the timestamp and camera position from the metadata of the matched video frame.

[0159] Step 5: Behavioral pattern analysis

[0160] The server analyzes the target person's behavioral patterns based on the recorded timestamps and camera location information, plotting their movement paths, analyzing their stay times, and identifying frequently visited locations.

[0161] Input: timestamp, camera location information

[0162] Output: Behavioral pattern analysis results (route, stay time, frequently visited places)

[0163] Specifically, the server analyzes the timestamp and camera position data, plots the target person's movement path, and calculates the length of time they stayed there.

[0164] Step 6: Identity Verification

[0165] After completing the behavioral pattern analysis, the server compares the target person's facial feature data with the identification database using a facial identification model (such as Facenet or DeepFace).

[0166] Input: Facial feature data

[0167] Output: Identity verification result

[0168] Specifically, the server accesses an identification database, matches the facial feature data, and confirms the identity of the matching individual.

[0169] Step 7: Report generation and presentation

[0170] Finally, the server generates a detailed report containing the analysis results and the identification results, including the route of travel, behavioral patterns, timestamps, camera locations, and the identification results. The device then presents this report to the user.

[0171] Input: behavioral pattern analysis results, identity verification results

[0172] Output: Detailed report

[0173] Specifically, the server compiles the movement route, behavioral patterns, timestamps, camera locations, and identity verification results to generate a report, which is then displayed on the device for the user to view.

[0174] Prompt Sentence Examples

[0175] "Using this photo, please briefly explain the entire process of identifying the person in question by extracting facial and clothing characteristics, searching a security camera footage database, and analyzing their behavioral patterns to confirm their identity."

[0176] (Application example 1)

[0177] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0178] Existing systems that use security camera footage and police databases to track and identify suspicious individuals have difficulty informing administrators of suspicious individuals' movements in real time. Furthermore, delays in reporting the results of the analysis mean that appropriate countermeasures cannot be implemented promptly.

[0179] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0180] In this invention, the server includes a facial recognition unit, a clothing feature extraction unit, a unit for accessing a security camera video database, a unit for searching for video based on the identified feature data, a unit for recording timestamps and camera locations, a unit for analyzing behavioral patterns, a unit for verifying identity by checking against a police database, a unit for generating and presenting a report of the analysis results, and a unit for sending data to a smartphone application to notify the administrator. This makes it possible to notify the location and behavioral patterns of suspicious individuals in real time.

[0181] "Facial recognition means" is a technology that extracts facial feature points from an input facial image and identifies an individual using a specific algorithm.

[0182] "Clothing feature extraction means" is a technology that analyzes features such as clothing color, shape, and pattern from an input image and extracts them as data.

[0183] "Means for accessing a security camera video database" refers to technology that accesses a database that stores video data collected from multiple security cameras, and performs searches and retrieval.

[0184] "Means for searching for footage based on identified feature data" refers to a technology that searches security camera footage in a database based on extracted facial and clothing feature data, and identifies the person in question.

[0185] The "means for recording timestamps and camera positions" is a technology for recording the time information of a specified video frame and the location where the camera is installed.

[0186] "Means for analyzing behavioral patterns" refers to technology that analyzes behavioral patterns such as a person's movement route and length of stay based on timestamps and camera position information.

[0187] "Means of confirming identity by comparing with police database" refers to a technology that compares extracted facial feature data with a database held by the police to confirm the identity of the target person.

[0188] "Means for generating and presenting analysis results as a report" refers to a technology that creates a detailed report based on the results of behavioral pattern analysis and identity verification and provides it to the user.

[0189] "Means for sending data to a smartphone application and notifying the user" refers to a technology that sends generated reports and analysis results to a smartphone application in real time and notifies the user.

[0190] The present invention is a system including a facial recognition unit, a clothing feature extraction unit, a security camera video database access unit, a video search unit based on the identified feature data, a timestamp and camera location recording unit, a behavioral pattern analysis unit, a police database comparison unit to confirm identity, a report generating and presenting the analysis results, and a smartphone application that notifies the user of the analysis results. This system is suitable for quickly and accurately tracking and identifying suspicious individuals using security camera video and the police database.

[0191] Face recognition and clothing feature extraction

[0192] The server receives the photo data uploaded by the user to the device and extracts the face area using a face recognition algorithm (e.g., OpenCV or Google Vision API), identifies facial feature points (positions of eyes, nose, mouth, facial contours, and other features), and generates data on clothing features (color, pattern, shape, etc.).

[0193] Video database search and matching

[0194] The server accesses a security camera video database and searches for video data based on the extracted feature data. It then performs facial recognition and clothing feature matching on the security camera footage to identify matching individuals. For each identified frame, the server records the timestamp and camera location, and tracks the target individual's movement path.

[0195] Behavioral pattern analysis and identity verification

[0196] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, it generates information such as the person's route of travel, length of stay, and frequently visited places. The server then organizes the analysis results. The server then compares the target person's facial feature data with the police database to confirm their identity.

[0197] Report generation and notification

[0198] Finally, the server generates a detailed report containing the analysis and identity verification results and sends a notification to the smartphone application, where the user can view the report and take any necessary measures.

[0199] For example, if a public facility manager spots someone behaving suspiciously in a parking lot and uploads a photo of them to a smartphone application, the system will review security camera footage and analyze the suspicious person's movement path. It will then compare the person's identity with a police database and confirm their identity. The manager will receive a real-time notification, enabling them to respond quickly.

[0200] Prompt Sentence Examples

[0201] "Upload a photo of someone behaving suspiciously in your parking lot. We will then extract the suspicious person's facial and clothing characteristics from the photo and search our security camera footage database for matching footage. We will then confirm the person's behavioral patterns and identity, and generate a detailed report to notify you."

[0202] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0203] Step 1:

[0204] A user uploads a photo of a suspicious person using a smartphone application. The input is the photo data of the suspicious person, and the output is the photo data sent to the server.

[0205] Step 2:

[0206] The server analyzes the received photo data and extracts the face region using a face recognition algorithm (e.g., OpenCV or Google Vision API). The input is the photo data, and the output is facial feature points (positions of eyes, nose, mouth, facial contours, and other features).

[0207] Step 3:

[0208] At the same time, the server extracts clothing features from the photo data. Specifically, it analyzes color, pattern, and shape and generates feature data. The input is the photo data, and the output is clothing feature data.

[0209] Step 4:

[0210] The server accesses the security camera video database and searches the video data based on the extracted facial and clothing feature data. The input is the facial and clothing feature data, and the output is the video frame of the matching person.

[0211] Step 5:

[0212] The server records the timestamp and camera position for each frame in which a matching person is identified. The input is the matching video frame, and the output is a record of the time and camera position information.

[0213] Step 6:

[0214] The server analyzes the target person's behavioral patterns based on timestamps and camera location information. The input is time information and camera location information, and the output is behavioral pattern information such as travel route, length of stay, and frequently visited places.

[0215] Step 7:

[0216] The server compares the facial feature data with the police database to confirm the identity of the person. The input is the facial feature data, and the output is the identity confirmation result.

[0217] Step 8:

[0218] The server generates a detailed report including the analysis results and identity verification results, and notifies the smartphone application of the report. The input is behavioral pattern information and identity verification results, and the output is a report sent to the user's device.

[0219] Step 9:

[0220] The user checks the received report through a smartphone application and takes necessary measures. The input is the report from the server, and the output is the countermeasure action taken by the user.

[0221] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0222] The present invention combines a system that extracts facial recognition and clothing features from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and verifies their identity in cooperation with a police database with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention will be described in detail below.

[0223] System Overview

[0224] The system consists of the following main components:

[0225] 1. User Device:

[0226] An interface where users can input photo data and receive analysis results.

[0227] An emotion engine that recognizes user emotions and collects emotion data.

[0228] 2. Server:

[0229] It handles core processing such as data processing, facial recognition, behavioral analysis, and database matching.

[0230] Emotional data from the emotion engine is used for analysis to optimize the way information is presented to users.

[0231] 3. Security camera footage database:

[0232] A database that stores video data collected from multiple security cameras.

[0233] 4. Police Database:

[0234] Databases containing personally identifiable information.

[0235] System operation and program processing

[0236] 1. Data Entry

[0237] The user uploads a photo of the individual to be tracked to the device. This photo data includes both the face and clothing. The device sends the photo data to the server and recognizes the user's emotions to generate emotion data. This emotion data is also sent to the server.

[0238] 2. Pretreatment

[0239] The server receives the photo data and first extracts the facial region using a facial recognition algorithm. From this facial region, it identifies specific feature points such as the eyes, nose, and mouth, and generates facial feature data. At the same time, the server analyzes the clothing features (color, pattern, shape, etc.) from the photo data and generates clothing feature data.

[0240] 3. Video search and feature data matching

[0241] The server accesses a security camera video database and searches for video data based on the identified feature data. From the security camera video, it performs facial recognition and clothing feature matching to identify matching individuals.

[0242] 4. Recording timestamps and camera positions

[0243] For each frame in which a matching person is identified, the server records the timestamp and camera position, allowing the path of the person to be traced.

[0244] 5. Behavioral Pattern Analysis

[0245] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, it generates information such as the person's route, length of stay, and frequently visited places. This analysis also takes into account the user's emotional data.

[0246] 6. Identity Verification

[0247] After the behavioral pattern analysis is completed, the server will compare the target person's facial feature data with the police database, and based on the matching results, the target person's identity will be confirmed.

[0248] 7. Report generation and presentation

[0249] Finally, the server generates a detailed report containing the analysis results, emotion data, and identity verification results. This report includes behavioral patterns, timestamps, camera locations, identity verification results, and information presentation methods adjusted according to the emotion data. The device then presents this report to the user and provides the necessary information.

[0250] Specific examples

[0251] For example, to track a suspicious person in a public facility, a user uploads a photo of the suspicious person to the system. The device receives the photo data and sends it to the server, recognizing the user's emotions to generate emotion data. The server extracts facial and clothing features from the photo and accesses a security camera footage database to detect the person. The server then analyzes the person's behavioral patterns and identifies their route and where they stayed. Finally, it compares the person with a police database to confirm their identity, and a detailed report is provided to the user. This report is based on the user's emotion data and presented in an easy-to-understand format. This enables accurate and fast person tracking.

[0252] In this way, the present invention improves the efficiency and accuracy of crime prevention activities and enables the provision of information that takes into consideration the user's emotions.

[0253] The processing flow will be explained below.

[0254] Step 1:

[0255] The user uploads a photo of the individual to be tracked to the device, which receives the photo data and sends it to the server.

[0256] Step 2:

[0257] The device activates an emotion engine to recognize the user's emotions, collecting emotion data from the user's voice, facial expressions, input data, etc.

[0258] Step 3:

[0259] The device sends the collected emotion data to the server, which receives the photo data and emotion data and begins processing.

[0260] Step 4:

[0261] The server uses a facial recognition algorithm to extract facial regions from the photograph, and then identifies feature points such as the eyes, nose, and mouth from the facial regions to generate facial feature data.

[0262] Step 5:

[0263] The server analyzes the clothing characteristics (color, pattern, shape, etc.) from the photograph data and generates clothing characteristic data.

[0264] Step 6:

[0265] The server accesses the security camera video database and searches the video data based on the acquired feature data. It then performs facial recognition and clothing feature matching on the security camera video to identify matching individuals.

[0266] Step 7:

[0267] The server records the timestamp and camera position for each frame in which a matching person is identified.

[0268] Step 8:

[0269] The server uses the recorded timestamps and camera positions to analyze the target person's movement route, duration of stay, frequently visited places, etc.

[0270] Step 9:

[0271] The server organizes the analysis results and extracts behavioral patterns, summarizing the information in a format that is easy for users to understand.

[0272] Step 10:

[0273] The server uses the facial feature data of the target person to compare it with the police database to confirm their identity, and obtains the matching results.

[0274] Step 11:

[0275] The server integrates the results of behavioral analysis and identity verification into a report, adjusts the report content based on the user's emotional data, and optimizes the information presentation method as needed.

[0276] Step 12:

[0277] The terminal presents the final report to the user, who then checks the report and obtains the necessary information.

[0278] Example 2

[0279] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0280] Conventional security systems are capable of tracking individuals and analyzing their behavior, but they do not provide information that takes into account the user's emotional state, and they are unable to present information in an efficient and easy-to-understand manner. Furthermore, they do not adequately utilize user emotional data to improve the accuracy of behavioral pattern analysis. As a result, tracking results and analytical information are sometimes not practical for users, and there is a need to improve the efficiency and accuracy of security activities.

[0281] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0282] In this invention, the server includes a facial recognition means, a clothing feature extraction means, a means for accessing a security camera video database, a means for searching for video based on the identified feature data, a means for recording a timestamp and camera location, a means for analyzing behavioral patterns, a means for verifying identity by checking against a police database, a means for generating and presenting a report of the analysis results, a means for recognizing a user's emotions using an emotion recognition engine and generating emotion data, and a means for optimizing information presentation by utilizing the generated emotion data in analysis, thereby enabling improved efficiency and accuracy in crime prevention activities.

[0283] The "face recognition means" is a means for detecting an individual's face area from image data, identifying feature points (eyes, nose, mouth, etc.), and extracting facial feature data.

[0284] The "clothing feature extraction means" is a means for analyzing clothing features (color, pattern, shape, etc.) from image data and generating clothing feature data.

[0285] "Means for accessing a security camera video database" refers to a means for connecting to a database that stores video data collected from a large number of security cameras and obtaining the necessary video data from there.

[0286] The "means for searching for video based on identified feature data" refers to a means for searching for video in a security camera video database based on the extracted facial feature data and clothing feature data, and identifying matching individuals.

[0287] The "means for recording timestamps and camera locations" refers to a means for recording the time of capture and the camera location information for each video frame in which a matching person is identified.

[0288] "Means for analyzing behavioral patterns" refers to a method for analyzing a target person's movement route, length of stay, frequently visited places, etc. from recorded timestamps and camera location information to analyze behavioral patterns.

[0289] "Means for confirming identity by checking against a police database" refers to a means for comparing facial feature data after behavioral pattern analysis with a police database and confirming the identity of the target person based on the matching results.

[0290] The "means for generating and presenting analysis results as a report" refers to a means for compiling behavioral analysis results, analysis results based on emotional data, and comparison results with police databases to create a detailed report and provide it to the user.

[0291] The "means for recognizing a user's emotions using an emotion recognition engine and generating emotion data" refers to a means for taking a picture of the user's face, analyzing their emotional state through an emotion recognition algorithm, and generating the results as data.

[0292] "Means for optimizing information presentation by utilizing generated emotional data in analysis" refers to means for adjusting the analysis results and information presentation method based on the acquired emotional data, and providing information in a form that is most easily understandable for the user.

[0293] System Overview

[0294] This system extracts facial recognition and clothing characteristics from personal photo data, and tracks and analyzes the behavior of target individuals using a security camera video database. Furthermore, by linking with a police database to verify identity and combining it with an emotion engine that recognizes user emotions, it achieves efficient and highly accurate crime prevention activities.

[0295] Key Components

[0296] User device:

[0297] It provides an interface where users can input photo data and receive analysis results.

[0298] An emotion recognition engine is used to collect user emotion data and send it to a server.

[0299] server:

[0300] It handles core processing such as data processing, facial recognition, behavioral analysis, and database matching.

[0301] Emotional data from the emotion engine is used for analysis to optimize the way information is presented to users.

[0302] Security camera footage database:

[0303] It is a database that stores video data collected from multiple security cameras.

[0304] Police Database:

[0305] It is a database containing personally identifiable information.

[0306] Specific actions

[0307] 1. Data Entry and Emotion Recognition

[0308] The user uploads a photo of the individual to be tracked to the device. The device receives the photo data and temporarily stores it in local storage. The device's built-in emotion engine then captures the user's face with a camera and generates emotion data in real time using an emotion recognition algorithm, such as TensorFlow. This emotion data is also sent to the server.

[0309] 2. Preprocessing of photo data

[0310] The server preprocesses the received photo data using the OpenCV library. Specifically, it converts the image to grayscale and performs face detection. It then uses the Haar Cascade and Dlib libraries to detect the face area and identify feature points such as the eyes, nose, and mouth to generate facial feature data. At the same time, it analyzes clothing features (color, pattern, shape, etc.) from the photo data and generates clothing feature data. Clothing analysis libraries such as DeepFashion are used for this.

[0311] 3. Video Search and Database Matching

[0312] The server accesses the security camera video database and utilizes SQL queries and NoSQL database technologies (e.g., MongoDB). It searches the video data based on facial and clothing feature data. It uses FaceNet or DeepFace models for facial recognition and similarity search algorithms (e.g., Cosine Similarity) for clothing matching. The server identifies matching individuals from the search results and extracts their video frames.

[0313] 4. Recording timestamps and camera positions

[0314] The server records the timestamp and camera location information for each video frame of the identified person, and stores the person's movement path in a detailed database.The server efficiently indexes the timestamp and location information using Elasticsearch.

[0315] 5. Behavioral Pattern Analysis

[0316] The server uses the collected timestamps and location information to analyze the target person's behavioral patterns using machine learning models (e.g., time series analysis methods and LSTM). It generates detailed information such as the person's route, length of stay, and frequently visited places. The server also incorporates the user's emotional data into the analysis and adjusts the analysis results.

[0317] 6. Identity Verification

[0318] After completing the behavioral pattern analysis, the server compares the facial feature data of the target person with the police database, using REST API or SOAP to access the police database and confirm the target person's identity based on the comparison results.

[0319] 7. Report generation and presentation

[0320] The server compiles the analysis results, emotion data, and identity verification results to generate a detailed report, possibly using Jupyter Notebook or ReportLab. The device receives the generated report and presents it to the user in the form of a dashboard or PDF report.

[0321] Specific examples

[0322] For example, to track a suspicious individual in a public facility, a user uploads a photo of the suspicious individual to the system. The device receives the photo data and sends it to the server, generating the user's emotional data using an emotion engine. The server extracts facial and clothing features from the uploaded photo and searches a security camera footage database to detect the individual. The server then analyzes the identified individual's behavioral patterns, determines their route of travel and where they stayed, and finally verifies their identity by comparing it with a police database. A detailed report is generated and provided to the user via the device. This report is presented in an easy-to-read format, taking into account the user's emotional data.

[0323] Example prompts for generative AI models

[0324] "We are developing a system to track suspicious individuals in public facilities. Please tell us how to build a program that inputs photographic data and searches a security camera footage database using the facial and clothing data obtained. Also, please explain in detail the procedure for analyzing the individual's behavioral patterns obtained in this way and comparing them with the police database to confirm their identity."

[0325] In this way, the system improves the efficiency and accuracy of crime prevention activities while providing information that takes into consideration the user's emotions.

[0326] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0327] Step 1:

[0328] The user uploads a photo of the individual to be tracked to the device. The device receives this photo data and temporarily stores it in local storage. Next, the device's built-in emotion recognition engine captures the user's face with a camera and generates emotion data using an emotion recognition algorithm. TensorFlow is used for this. This emotion data is sent to the server along with the photo data. The input is the individual's photo and the user's facial image, and the output is the photo data and emotion data.

[0329] Step 2:

[0330] The server preprocesses the received photo data using the OpenCV library. Specifically, it converts the image to grayscale and performs face detection. Here, the Haar Cascade and Dlib libraries are used. The face area is detected and feature points such as the eyes, nose, and mouth are identified to generate facial feature data. At the same time, the server uses a clothing analysis library such as DeepFashion to analyze clothing features (color, pattern, shape, etc.) from the photo data and generate clothing feature data. The input is photo data, and the output is facial feature data and clothing feature data.

[0331] Step 3:

[0332] The server accesses the security camera video database and utilizes SQL queries and NoSQL database technology (e.g., MongoDB). It searches security camera footage based on facial feature data and clothing feature data. It uses FaceNet or DeepFace models for facial recognition and a similarity search algorithm (e.g., Cosine Similarity) for clothing matching. It identifies matching people from the search results and extracts their video frames. The input is facial feature data and clothing feature data, and the output is video frames of matching people.

[0333] Step 4:

[0334] The server records the timestamp and camera location information for each video frame of the identified person and stores the person's movement path in a database. Here, we use Elasticsearch to index the timestamp and location information. The input is the video frame of the matching person, and the output is a record of the timestamp and camera location information.

[0335] Step 5:

[0336] The server analyzes the target person's behavioral patterns based on the collected timestamps and location information. This uses machine learning models (e.g., time series analysis methods and LSTM). The behavioral pattern analysis includes information such as travel routes, length of stay, and frequently visited places. In addition, the analysis results are adjusted taking into account the user's emotional data. The inputs are timestamps, camera location information, and emotional data, and the output is the analyzed behavioral patterns.

[0337] Step 6:

[0338] After completing the behavioral pattern analysis, the server compares the facial feature data of the target person with the police database. It accesses the police database using REST API or SOAP. Based on the comparison results, the target person's identity is confirmed. The input is facial feature data, and the output is the identity confirmation result.

[0339] Step 7:

[0340] The server compiles the analysis results, emotion data, and identity verification results to generate a detailed report. Jupyter Notebook or ReportLab can be used for this. The generated report is then presented to the user via their terminal. The information is presented in the form of a dashboard or PDF report. The input is the analysis results, emotion data, and identity verification results, and the output is a detailed report.

[0341] (Application example 2)

[0342] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0343] Conventional security systems were able to track individuals by facial recognition and extracting clothing characteristics, but they were unable to take the user's emotional state into account, making it difficult for them to provide optimal information to users in stressful situations. There was also a need for a system that could track individuals in real time and respond appropriately based on the user's emotions.

[0344] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a clothing feature extraction means, a means for accessing a security camera video database, a means for recording timestamps and camera locations, a means for analyzing behavioral patterns, a means for verifying identity by checking against a police database, a means for incorporating an emotion engine into the user terminal, and a means for utilizing the user's emotion data for analysis and presenting optimized information. This enables accurate tracking and analysis of individuals as well as the provision of information that takes the user's emotions into consideration.

[0345] "Facial recognition means" refers to devices and technology that detect an individual's face from input photographic data, analyze its features, and digitize them.

[0346] "Clothing feature extraction means" refers to a device and technology for analyzing the features of an individual's clothing, such as color, pattern, and shape, from input photographic data and converting them into data.

[0347] "Means for accessing a security camera video database" refers to devices and technologies for accessing video data collected from security cameras and obtaining the necessary information.

[0348] "Means for searching for video based on identified feature data" refers to devices and technologies for searching for matching video within a security camera video database based on feature data obtained through facial recognition or clothing feature extraction.

[0349] "Means for recording timestamps and camera locations" refers to devices and techniques for recording the date and time (timestamp) when the image was recorded and the location of the camera for each frame in which an identified person appears.

[0350] "Means for analyzing behavioral patterns" refers to devices and technology for analyzing behavioral patterns, such as the route a target person took and how long they stayed in a particular location, based on timestamps and camera position data.

[0351] "Means for verifying identity by comparing with police databases" refers to devices and technologies for verifying the identity of a person by comparing the acquired facial feature data with personal identification information held by the police.

[0352] The "means for generating and presenting the analysis results as a report" refers to a device and technology for integrating the above analysis results and generating and presenting them as a report in a format that is easy for the user to understand.

[0353] "Means for incorporating an emotion engine into a user terminal" refers to devices and techniques for incorporating emotion recognition software and hardware into a user terminal.

[0354] "Means for utilizing user emotional data for analysis and presenting optimized information" refers to devices and technologies that analyze the user's emotional state, adjust the way information is presented based on the results, and provide data in a form that is more suitable for the user.

[0355] The present invention is a system that includes a facial recognition unit, a clothing feature extraction unit, a security camera video database access unit, a video search unit based on identified feature data, a timestamp and camera location recorder, a behavioral pattern analysis unit, a police database for identifying the user, a report generating and presenting the analysis results, an emotion engine embedded in the user's device, and a unit for analyzing and presenting optimized information based on the user's emotion data. This system enables accurate tracking of individuals and the provision of information based on the user's emotion.

[0356] Hardware and software used

[0357] Hardware: Smartphones, smart glasses

[0358] Software: OpenCV (image processing library), EmotionRecognition (emotion analysis engine), police database API client

[0359] Data processing and calculation flow

[0360] A user uploads a photo of a specific individual to the device and simultaneously captures the user's emotional state. When the device transmits the photo data to the server, it uses an emotion engine to generate the user's emotional data, which is also transmitted to the server.

[0361] The server uses OpenCV to perform facial recognition based on the received photo data. It extracts the facial area from the photo and digitizes its feature points (the positions of the eyes, nose, mouth, etc.). At the same time, it extracts clothing features from the photo data and generates clothing feature data. This data is used to search the security camera video database and identify footage of matching people.

[0362] The system records the timestamp and camera position for each frame in which the identified person appears, thereby clarifying the person's movement path. The server then performs behavioral analysis to analyze the person's movement patterns and the length of time they stayed in the area. Finally, the results of this analysis are compared with a police database to confirm the person's identity.

[0363] The analysis results are integrated with the user's emotional data to generate a final report, which is presented in an easy-to-understand format for the user, enabling appropriate responses that take into account the user's emotional state.

[0364] Specific example explanation

[0365] For example, if a suspicious person appears at a concert venue, a security officer can upload a photo of the person using their smartphone to the system. The system then sends the photo to a server, which performs facial recognition and clothing feature extraction, then searches the security camera footage database to analyze the person's behavioral patterns.

[0366] If a security officer experiences high stress levels, the system can provide additional support in real time to allow for a rapid response. This process can be facilitated by prompts such as:

[0367] Prompt Sentence Examples

[0368] "Track the person in this photo, verify their identity against police databases, and suggest appropriate actions based on the user's current emotions."

[0369] In this way, the present invention improves the efficiency and accuracy of crime prevention activities and enables the provision of information that takes into consideration the user's emotions.

[0370] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0371] Step 1:

[0372] A user uploads a photo of a specific individual to a smartphone. The input is photo data that clearly shows the face. The output is the photo data that is saved on the device.

[0373] Step 2:

[0374] The device recognizes the user's emotional state from video or image data and generates emotional data using an emotion engine. The input is the user's video or image data, which is analyzed and emotional data (e.g., stress level, tension, etc.) is output.

[0375] Step 3:

[0376] The terminal sends photo data and emotion data to the server. The input at this time is the photo data and emotion data. The output is the data received by the server.

[0377] Step 4:

[0378] The server performs face recognition using the received photo data. The software used is OpenCV. The input is the photo data, the face is detected, and data on specific feature points (eyes, nose, mouth) is output.

[0379] Step 5:

[0380] The server extracts clothing features from the photo data, including information on color, pattern, shape, etc. The input is the photo data, and the output is clothing feature data.

[0381] Step 6:

[0382] The server searches a security camera video database for matching video based on the facial feature data and clothing feature data. The input is the facial feature data and clothing feature data, and the output is the matching video data.

[0383] Step 7:

[0384] The server records the timestamp and camera location of the video in which the matching person appears. The input is the matching video data, and the output is the timestamp and camera location information.

[0385] Step 8:

[0386] The server analyzes the target person's behavioral patterns based on the timestamp and camera position information. For example, the server outputs the person's route and duration of stay. The input is the timestamp and camera position information.

[0387] Step 9:

[0388] The server compares the results of behavioral pattern analysis with the police database to confirm the identity of the person. The input is facial feature data, and the matching results from the police database are output.

[0389] Step 10:

[0390] The server integrates the analysis results with the user's emotional data and generates a final report. The inputs are the behavioral pattern analysis results and the emotional data, and the output is a report.

[0391] Step 11:

[0392] The terminal presents this generated report to the user. The input is the report sent from the server, and the output is to display it in a user-friendly format.

[0393] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0394] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0395] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0396] [Second embodiment]

[0397] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0398] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0399] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0400] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0401] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0402] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0403] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0404] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0405] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0406] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0407] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0408] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0409] The present invention is a system that extracts facial recognition and clothing characteristics from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and further verifies their identity in cooperation with a police database. Specific embodiments of the present invention are described in detail below.

[0410] System Overview

[0411] The system consists of the following main components:

[0412] 1. User terminal: The interface where the user inputs photo data and receives the analysis results.

[0413] 2. Server: Responsible for central processing such as data processing, facial recognition, behavioral analysis, and database matching.

[0414] 3. Security camera video database: A database that stores video data collected from multiple security cameras.

[0415] 4. Police databases: Databases containing personally identifiable information.

[0416] System operation and program processing

[0417] 1. Data Entry

[0418] The user uploads a photo of the individual to be tracked to the device. This photo data includes both the face and clothing. The device then sends the photo data to the server.

[0419] 2. Pretreatment

[0420] The server receives this photo data and first extracts the facial region using a facial recognition algorithm, from which it identifies the following feature points:

[0421] Eye, nose, and mouth position

[0422] Facial contours

[0423] Other features (mole, scar, etc.)

[0424] At the same time, the server extracts the clothing's characteristics, specifically analyzing its color, pattern, shape, etc., and generates feature data.

[0425] 3. Video search and feature data matching

[0426] The server accesses a security camera video database and searches for video data based on the identified feature data. From the security camera video, it performs facial recognition and clothing feature matching to identify matching individuals.

[0427] 4. Recording timestamps and camera positions

[0428] For each frame in which a matching person is identified, the server records the timestamp and camera position, allowing the path of the person to be traced.

[0429] 5. Behavioral Pattern Analysis

[0430] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, the following information is generated:

[0431] Travel route

[0432] Stay time

[0433] Frequently visited places

[0434] This information is organized as a result of behavioral pattern analysis.

[0435] 6. Identity Verification

[0436] After the behavioral pattern analysis is completed, the server will compare the target person's facial feature data with the police database, and based on the matching results, the target person's identity will be confirmed.

[0437] 7. Report generation and presentation

[0438] Finally, the server generates a detailed report containing the results of the analysis and identification, including the following information:

[0439] Travel route

[0440] Behavioral patterns

[0441] Timestamp and camera position

[0442] Identity verification results

[0443] The terminal presents this report to the user and provides the necessary information.

[0444] Specific examples

[0445] For example, to track a suspicious person in a public facility, a user uploads a photo of the suspicious person to the system. The server extracts facial and clothing features from the photo and accesses a security camera footage database to detect the person. The server then analyzes the person's behavioral patterns to identify their route of travel and where they are staying. Finally, the server checks the data against a police database to confirm their identity, and a detailed report is provided to the user. This allows for accurate and rapid tracking of people.

[0446] The processing flow will be explained below.

[0447] Step 1:

[0448] The user uploads a photo of the individual to be tracked to the device, which receives and stores the photo data.

[0449] Step 2:

[0450] The device sends the photo data to the server, which receives it and applies a facial recognition algorithm.

[0451] Step 3:

[0452] The server uses a facial recognition algorithm to extract facial regions from the photograph, identify specific feature points (such as the eyes, nose, and mouth) from the facial regions, and generate facial feature data.

[0453] Step 4:

[0454] At the same time, the server analyzes the clothing characteristics (color, pattern, shape, etc.) from the photograph data and generates clothing characteristic data.

[0455] Step 5:

[0456] The server accesses the security camera video database and obtains the latest video data.

[0457] Step 6:

[0458] The server performs facial recognition and clothing feature matching on the acquired video data to search for people who match specific feature data.

[0459] Step 7:

[0460] The server records the timestamp and camera position for each frame in which a matching person is identified.

[0461] Step 8:

[0462] The server uses multiple timestamps and camera location information to analyze the target person's behavioral characteristics, such as their route of travel and length of stay.

[0463] Step 9:

[0464] The server organizes the analysis results and extracts behavioral patterns (frequently visited places, places where people spend long periods of time, etc.).

[0465] Step 10:

[0466] The server uses the analyzed facial feature data to access police databases and verify the identity of the person in question.

[0467] Step 11:

[0468] The server adds the matching results to the report and combines all the analysis results to generate the final report.

[0469] Step 12:

[0470] The terminal presents the final report to the user, who then checks the report and obtains the necessary information.

[0471] Example 1

[0472] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0473] In modern society, numerous security cameras and surveillance devices are installed, helping to prevent and investigate crime. However, quickly and accurately identifying and tracking specific individuals from large amounts of video data is difficult, time-consuming, and laborious. To solve this problem, there is a need for a system that can extract facial recognition and clothing characteristics from photographic data, track and analyze the behavior of target individuals using a security camera video database, and further verify their identity in conjunction with an identification database.

[0474] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0475] In this invention, the server includes a face recognition means, a clothing feature extraction means, a means for accessing a video database, a means for searching video data based on the identified feature data, a means for recording timestamps and location information, a means for analyzing behavioral patterns, a means for verifying identity by comparing with an identification database, a means for generating and presenting analysis results as a report, a user interface including a means for uploading photo data, and a means for using a generative AI model as part of the analysis process. This makes it possible to quickly and accurately identify a target person from photo data, analyze the person's behavioral patterns, and verify the person's identity in cooperation with the identification database.

[0476] "Facial recognition means" refers to a device or program that uses technology or algorithms to detect facial feature points and identify individuals.

[0477] "Clothing feature extraction means" refers to a device or program that uses technology or algorithms to analyze and extract clothing features such as color, pattern, and shape from image data.

[0478] "Means for accessing a video database" refers to a device or program that connects to a database in which video data collected from multiple cameras is stored and acquires the necessary data.

[0479] "Means for searching video data based on identified feature data" refers to devices or programs that use technology or algorithms to search for matching people from videos in a video database using extracted facial and clothing feature data.

[0480] The "means for recording timestamp and location information" refers to a device or program that acquires and records the shooting date and time of the searched video frame and the installation location of the camera.

[0481] "Means for analyzing behavioral patterns" refers to devices or programs that use technologies or algorithms to analyze a target person's behavioral patterns, such as their travel routes, length of stay, and frequently visited places, based on recorded timestamps and location information.

[0482] "Means for verifying identity by matching against an identification database" refers to devices or programs that use technology or algorithms to compare analyzed facial feature information with an identification database and verify the identity of the person based on matching data.

[0483] The "means for generating and presenting the analysis results as a report" refers to a device or program that organizes the results of behavioral pattern analysis and identity verification, and generates a report that can be visually presented to the user.

[0484] The "user interface including a means for uploading photographic data" is an interface that allows a user to input photographic data into the system via a terminal.

[0485] "Means of using a generative AI model as part of the analytical process" refers to devices or programs that utilize generative AI models as part of the analysis, and use technologies or algorithms to achieve more advanced data processing and pattern recognition.

[0486] The present invention is a system that extracts facial recognition and clothing characteristics from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and verifies their identity in cooperation with an identification database. To implement this system, the following hardware and software are used.

[0487] System Overview

[0488] The system consists of the following main components:

[0489] 1. User terminal: An interface where the user inputs photo data and receives the analysis results.

[0490] 2. Server: Responsible for central processing such as data processing, facial recognition, behavior analysis, and database matching.

[0491] 3. Security camera video database: A database that stores video data collected from multiple security cameras.

[0492] 4. Identification Database: A database containing personally identifiable information.

[0493] Hardware and software used

[0494] Face recognition algorithm: Uses OpenCV and dlib libraries.

[0495] Clothing feature extraction algorithm: Uses color analysis using RGB values, texture analysis, and contour detection algorithms.

[0496] Video search algorithm: Facenet and DeepFace are used for face recognition, and Elasticsearch's image search plugin is used for clothing matching.

[0497] Behavioral analysis algorithms: Use custom algorithms to plot movement paths and analyze dwell times.

[0498] Identity database matching algorithm: Uses face identification model to match with identity database.

[0499] Detailed procedure

[0500] The user uploads photo data of the object to be tracked to the device, which then sends the uploaded photo data to the server. The server then uses a facial recognition algorithm to detect the facial area and extract feature points such as the eyes, nose, mouth, facial contours, moles, and scars.

[0501] At the same time, the server analyzes the color, pattern, and shape of the clothing to generate feature data. Based on the analyzed feature data, it accesses a security camera video database to search for matching video data. The server then obtains and records the timestamp and camera position of the video frame of the matching person.

[0502] The system then analyzes the person's behavioral patterns based on timestamps and camera location data to identify their route and location. Finally, the system compares the facial feature data with an identification database to confirm their identity. The server then processes the analysis results and generates a detailed report that is presented to the user's device.

[0503] Specific examples

[0504] For example, when tracking a suspicious person in a public facility, the user uploads photo data to the system. The server extracts facial and clothing features, searches a security camera footage database, and analyzes the suspicious person's behavioral patterns. The server then verifies the person's identity against an identification database and provides a detailed report to the user.

[0505] Prompt Sentence Examples

[0506] "Using this photo, please briefly explain the entire process of identifying the person in question by extracting facial and clothing characteristics, searching a security camera footage database, and analyzing their behavioral patterns to confirm their identity."

[0507] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0508] Step 1: Data entry

[0509] The user uploads photographic data of the individual to be tracked to the device, which includes information on their face and clothing.

[0510] Input: Photo data (face and clothing information)

[0511] Output: Photo data sent to the server

[0512] Specifically, the user selects photo data using the file upload function of the device and clicks the upload button. The device then sends the selected photo data to the server as an HTTP request.

[0513] Step 2: Face recognition and clothing feature extraction (preprocessing)

[0514] The server receives this photo data and extracts the face region using a facial recognition algorithm, for example, using the face detection function of OpenCV or dlib.

[0515] Input: Submitted photo data

[0516] Output: Facial feature points (eyes, nose, mouth position, facial contours, moles, scars, etc.)

[0517] Specifically, the server reads the photo data using an image processing library, detects the facial area, and extracts feature points.

[0518] At the same time, the server extracts the clothing's features by analyzing its color (RGB value), pattern (texture analysis), and shape (contour detection) to generate feature data.

[0519] Input: Submitted photo data

[0520] Output: Clothing feature data (color, pattern, shape)

[0521] Specifically, the server extracts clothing feature data using a color analysis algorithm, a texture analysis algorithm, and a shape analysis algorithm.

[0522] Step 3: Feature data matching and video search

[0523] The server then accesses a security camera video database based on the identified facial and clothing feature data and searches for the relevant video data, using Facenet and DeepFace for facial recognition and Elasticsearch's image search plugin for clothing matching.

[0524] Input: facial feature points, clothing feature data

[0525] Output: Video data of the matching person

[0526] Specifically, the server queries a security camera video database to search for video frames that match the facial features and clothing data.

[0527] Step 4: Record timestamps and camera positions

[0528] For each frame in which a matching person is identified, the server captures and records the timestamp and camera position from the video data.

[0529] Input: Video data of the person to match

[0530] Output: timestamp and camera position information

[0531] Specifically, the server extracts and records the timestamp and camera position from the metadata of the matched video frame.

[0532] Step 5: Behavioral pattern analysis

[0533] The server analyzes the target person's behavioral patterns based on the recorded timestamps and camera location information, plotting their movement paths, analyzing their stay times, and identifying frequently visited locations.

[0534] Input: timestamp, camera location information

[0535] Output: Behavioral pattern analysis results (route, stay time, frequently visited places)

[0536] Specifically, the server analyzes the timestamp and camera position data, plots the target person's movement path, and calculates the length of time they stayed there.

[0537] Step 6: Identity Verification

[0538] After completing the behavioral pattern analysis, the server compares the target person's facial feature data with the identification database using a facial identification model (such as Facenet or DeepFace).

[0539] Input: Facial feature data

[0540] Output: Identity verification result

[0541] Specifically, the server accesses an identification database, matches the facial feature data, and confirms the identity of the matching individual.

[0542] Step 7: Report generation and presentation

[0543] Finally, the server generates a detailed report containing the analysis results and the identification results, including the route of travel, behavioral patterns, timestamps, camera locations, and the identification results. The device then presents this report to the user.

[0544] Input: behavioral pattern analysis results, identity verification results

[0545] Output: Detailed report

[0546] Specifically, the server compiles the movement route, behavioral patterns, timestamps, camera locations, and identity verification results to generate a report, which is then displayed on the device for the user to view.

[0547] Prompt Sentence Examples

[0548] "Using this photo, please briefly explain the entire process of identifying the person in question by extracting facial and clothing characteristics, searching a security camera footage database, and analyzing their behavioral patterns to confirm their identity."

[0549] (Application example 1)

[0550] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0551] Existing systems that use security camera footage and police databases to track and identify suspicious individuals have difficulty informing administrators of suspicious individuals' movements in real time. Furthermore, delays in reporting the results of the analysis mean that appropriate countermeasures cannot be implemented promptly.

[0552] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0553] In this invention, the server includes a facial recognition unit, a clothing feature extraction unit, a unit for accessing a security camera video database, a unit for searching for video based on the identified feature data, a unit for recording timestamps and camera locations, a unit for analyzing behavioral patterns, a unit for verifying identity by checking against a police database, a unit for generating and presenting a report of the analysis results, and a unit for sending data to a smartphone application to notify the administrator. This makes it possible to notify the location and behavioral patterns of suspicious individuals in real time.

[0554] "Facial recognition means" is a technology that extracts facial feature points from an input facial image and identifies an individual using a specific algorithm.

[0555] "Clothing feature extraction means" is a technology that analyzes features such as clothing color, shape, and pattern from an input image and extracts them as data.

[0556] "Means for accessing a security camera video database" refers to technology that accesses a database that stores video data collected from multiple security cameras, and performs searches and retrieval.

[0557] "Means for searching for footage based on identified feature data" refers to a technology that searches security camera footage in a database based on extracted facial and clothing feature data, and identifies the person in question.

[0558] The "means for recording timestamps and camera positions" is a technology for recording the time information of a specified video frame and the location where the camera is installed.

[0559] "Means for analyzing behavioral patterns" refers to technology that analyzes behavioral patterns such as a person's movement route and length of stay based on timestamps and camera position information.

[0560] "Means of confirming identity by comparing with police database" refers to a technology that compares extracted facial feature data with a database held by the police to confirm the identity of the target person.

[0561] "Means for generating and presenting analysis results as a report" refers to a technology that creates a detailed report based on the results of behavioral pattern analysis and identity verification and provides it to the user.

[0562] "Means for sending data to a smartphone application and notifying the user" refers to a technology that sends generated reports and analysis results to a smartphone application in real time and notifies the user.

[0563] The present invention is a system including a facial recognition unit, a clothing feature extraction unit, a security camera video database access unit, a video search unit based on the identified feature data, a timestamp and camera location recording unit, a behavioral pattern analysis unit, a police database comparison unit to confirm identity, a report generating and presenting the analysis results, and a smartphone application that notifies the user of the analysis results. This system is suitable for quickly and accurately tracking and identifying suspicious individuals using security camera video and the police database.

[0564] Face recognition and clothing feature extraction

[0565] The server receives the photo data uploaded by the user to the device and extracts the face area using a face recognition algorithm (e.g., OpenCV or Google Vision API), identifies facial feature points (positions of eyes, nose, mouth, facial contours, and other features), and generates data on clothing features (color, pattern, shape, etc.).

[0566] Video database search and matching

[0567] The server accesses a security camera video database and searches for video data based on the extracted feature data. It then performs facial recognition and clothing feature matching on the security camera footage to identify matching individuals. For each identified frame, the server records the timestamp and camera location, and tracks the target individual's movement path.

[0568] Behavioral pattern analysis and identity verification

[0569] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, it generates information such as the person's route of travel, length of stay, and frequently visited places. The server then organizes the analysis results. The server then compares the target person's facial feature data with the police database to confirm their identity.

[0570] Report generation and notification

[0571] Finally, the server generates a detailed report containing the analysis and identity verification results and sends a notification to the smartphone application, where the user can view the report and take any necessary measures.

[0572] For example, if a public facility manager spots someone behaving suspiciously in a parking lot and uploads a photo of them to a smartphone application, the system will review security camera footage and analyze the suspicious person's movement path. It will then compare the person's identity with a police database and confirm their identity. The manager will receive a real-time notification, enabling them to respond quickly.

[0573] Prompt Sentence Examples

[0574] "Upload a photo of someone behaving suspiciously in your parking lot. We will then extract the suspicious person's facial and clothing characteristics from the photo and search our security camera footage database for matching footage. We will then confirm the person's behavioral patterns and identity, and generate a detailed report to notify you."

[0575] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0576] Step 1:

[0577] A user uploads a photo of a suspicious person using a smartphone application. The input is the photo data of the suspicious person, and the output is the photo data sent to the server.

[0578] Step 2:

[0579] The server analyzes the received photo data and extracts the face region using a face recognition algorithm (e.g., OpenCV or Google Vision API). The input is the photo data, and the output is facial feature points (positions of eyes, nose, mouth, facial contours, and other features).

[0580] Step 3:

[0581] At the same time, the server extracts clothing features from the photo data. Specifically, it analyzes color, pattern, and shape and generates feature data. The input is the photo data, and the output is clothing feature data.

[0582] Step 4:

[0583] The server accesses the security camera video database and searches the video data based on the extracted facial and clothing feature data. The input is the facial and clothing feature data, and the output is the video frame of the matching person.

[0584] Step 5:

[0585] The server records the timestamp and camera position for each frame in which a matching person is identified. The input is the matching video frame, and the output is a record of the time and camera position information.

[0586] Step 6:

[0587] The server analyzes the target person's behavioral patterns based on timestamps and camera location information. The input is time information and camera location information, and the output is behavioral pattern information such as travel route, length of stay, and frequently visited places.

[0588] Step 7:

[0589] The server compares the facial feature data with the police database to confirm the identity of the person. The input is the facial feature data, and the output is the identity confirmation result.

[0590] Step 8:

[0591] The server generates a detailed report including the analysis results and identity verification results, and notifies the smartphone application of the report. The input is behavioral pattern information and identity verification results, and the output is a report sent to the user's device.

[0592] Step 9:

[0593] The user checks the received report through a smartphone application and takes necessary measures. The input is the report from the server, and the output is the countermeasure action taken by the user.

[0594] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0595] The present invention combines a system that extracts facial recognition and clothing features from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and verifies their identity in cooperation with a police database with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention will be described in detail below.

[0596] System Overview

[0597] The system consists of the following main components:

[0598] 1. User Device:

[0599] An interface where users can input photo data and receive analysis results.

[0600] An emotion engine that recognizes user emotions and collects emotion data.

[0601] 2. Server:

[0602] It handles core processing such as data processing, facial recognition, behavioral analysis, and database matching.

[0603] Emotional data from the emotion engine is used for analysis to optimize the way information is presented to users.

[0604] 3. Security camera footage database:

[0605] A database that stores video data collected from multiple security cameras.

[0606] 4. Police Database:

[0607] Databases containing personally identifiable information.

[0608] System operation and program processing

[0609] 1. Data Entry

[0610] The user uploads a photo of the individual to be tracked to the device. This photo data includes both the face and clothing. The device sends the photo data to the server and recognizes the user's emotions to generate emotion data. This emotion data is also sent to the server.

[0611] 2. Pretreatment

[0612] The server receives the photo data and first extracts the facial region using a facial recognition algorithm. From this facial region, it identifies specific feature points such as the eyes, nose, and mouth, and generates facial feature data. At the same time, the server analyzes the clothing features (color, pattern, shape, etc.) from the photo data and generates clothing feature data.

[0613] 3. Video search and feature data matching

[0614] The server accesses a security camera video database and searches for video data based on the identified feature data. From the security camera video, it performs facial recognition and clothing feature matching to identify matching individuals.

[0615] 4. Recording timestamps and camera positions

[0616] For each frame in which a matching person is identified, the server records the timestamp and camera position, allowing the path of the person to be traced.

[0617] 5. Behavioral Pattern Analysis

[0618] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, it generates information such as the person's route, length of stay, and frequently visited places. This analysis also takes into account the user's emotional data.

[0619] 6. Identity Verification

[0620] After the behavioral pattern analysis is completed, the server will compare the target person's facial feature data with the police database, and based on the matching results, the target person's identity will be confirmed.

[0621] 7. Report generation and presentation

[0622] Finally, the server generates a detailed report containing the analysis results, emotion data, and identity verification results. This report includes behavioral patterns, timestamps, camera locations, identity verification results, and information presentation methods adjusted according to the emotion data. The device then presents this report to the user and provides the necessary information.

[0623] Specific examples

[0624] For example, to track a suspicious person in a public facility, a user uploads a photo of the suspicious person to the system. The device receives the photo data and sends it to the server, recognizing the user's emotions to generate emotion data. The server extracts facial and clothing features from the photo and accesses a security camera footage database to detect the person. The server then analyzes the person's behavioral patterns and identifies their route and where they stayed. Finally, it compares the person with a police database to confirm their identity, and a detailed report is provided to the user. This report is based on the user's emotion data and presented in an easy-to-understand format. This enables accurate and fast person tracking.

[0625] In this way, the present invention improves the efficiency and accuracy of crime prevention activities and enables the provision of information that takes into consideration the user's emotions.

[0626] The processing flow will be explained below.

[0627] Step 1:

[0628] The user uploads a photo of the individual to be tracked to the device, which receives the photo data and sends it to the server.

[0629] Step 2:

[0630] The device activates an emotion engine to recognize the user's emotions, collecting emotion data from the user's voice, facial expressions, input data, etc.

[0631] Step 3:

[0632] The device sends the collected emotion data to the server, which receives the photo data and emotion data and begins processing.

[0633] Step 4:

[0634] The server uses a facial recognition algorithm to extract facial regions from the photograph, and then identifies feature points such as the eyes, nose, and mouth from the facial regions to generate facial feature data.

[0635] Step 5:

[0636] The server analyzes the clothing characteristics (color, pattern, shape, etc.) from the photograph data and generates clothing characteristic data.

[0637] Step 6:

[0638] The server accesses the security camera video database and searches the video data based on the acquired feature data. It then performs facial recognition and clothing feature matching on the security camera video to identify matching individuals.

[0639] Step 7:

[0640] The server records the timestamp and camera position for each frame in which a matching person is identified.

[0641] Step 8:

[0642] The server uses the recorded timestamps and camera positions to analyze the target person's movement route, duration of stay, frequently visited places, etc.

[0643] Step 9:

[0644] The server organizes the analysis results and extracts behavioral patterns, summarizing the information in a format that is easy for users to understand.

[0645] Step 10:

[0646] The server uses the facial feature data of the target person to compare it with the police database to confirm their identity, and obtains the matching results.

[0647] Step 11:

[0648] The server integrates the results of behavioral analysis and identity verification into a report, adjusts the report content based on the user's emotional data, and optimizes the information presentation method as needed.

[0649] Step 12:

[0650] The terminal presents the final report to the user, who then checks the report and obtains the necessary information.

[0651] Example 2

[0652] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0653] Conventional security systems are capable of tracking individuals and analyzing their behavior, but they do not provide information that takes into account the user's emotional state, and they are unable to present information in an efficient and easy-to-understand manner. Furthermore, they do not adequately utilize user emotional data to improve the accuracy of behavioral pattern analysis. As a result, tracking results and analytical information are sometimes not practical for users, and there is a need to improve the efficiency and accuracy of security activities.

[0654] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0655] In this invention, the server includes a facial recognition means, a clothing feature extraction means, a means for accessing a security camera video database, a means for searching for video based on the identified feature data, a means for recording a timestamp and camera location, a means for analyzing behavioral patterns, a means for verifying identity by checking against a police database, a means for generating and presenting a report of the analysis results, a means for recognizing a user's emotions using an emotion recognition engine and generating emotion data, and a means for optimizing information presentation by utilizing the generated emotion data in analysis, thereby enabling improved efficiency and accuracy in crime prevention activities.

[0656] The "face recognition means" is a means for detecting an individual's face area from image data, identifying feature points (eyes, nose, mouth, etc.), and extracting facial feature data.

[0657] The "clothing feature extraction means" is a means for analyzing clothing features (color, pattern, shape, etc.) from image data and generating clothing feature data.

[0658] "Means for accessing a security camera video database" refers to a means for connecting to a database that stores video data collected from a large number of security cameras and obtaining the necessary video data from there.

[0659] The "means for searching for video based on identified feature data" refers to a means for searching for video in a security camera video database based on the extracted facial feature data and clothing feature data, and identifying matching individuals.

[0660] The "means for recording timestamps and camera locations" refers to a means for recording the time of capture and the camera location information for each video frame in which a matching person is identified.

[0661] "Means for analyzing behavioral patterns" refers to a method for analyzing a target person's movement route, length of stay, frequently visited places, etc. from recorded timestamps and camera location information to analyze behavioral patterns.

[0662] "Means for confirming identity by checking against a police database" refers to a means for comparing facial feature data after behavioral pattern analysis with a police database and confirming the identity of the target person based on the matching results.

[0663] The "means for generating and presenting analysis results as a report" refers to a means for compiling behavioral analysis results, analysis results based on emotional data, and comparison results with police databases to create a detailed report and provide it to the user.

[0664] The "means for recognizing a user's emotions using an emotion recognition engine and generating emotion data" refers to a means for taking a picture of the user's face, analyzing their emotional state through an emotion recognition algorithm, and generating the results as data.

[0665] "Means for optimizing information presentation by utilizing generated emotional data in analysis" refers to means for adjusting the analysis results and information presentation method based on the acquired emotional data, and providing information in a form that is most easily understandable for the user.

[0666] System Overview

[0667] This system extracts facial recognition and clothing characteristics from personal photo data, and tracks and analyzes the behavior of target individuals using a security camera video database. Furthermore, by linking with a police database to verify identity and combining it with an emotion engine that recognizes user emotions, it achieves efficient and highly accurate crime prevention activities.

[0668] Key Components

[0669] User device:

[0670] It provides an interface where users can input photo data and receive analysis results.

[0671] An emotion recognition engine is used to collect user emotion data and send it to a server.

[0672] server:

[0673] It handles core processing such as data processing, facial recognition, behavioral analysis, and database matching.

[0674] Emotional data from the emotion engine is used for analysis to optimize the way information is presented to users.

[0675] Security camera footage database:

[0676] It is a database that stores video data collected from multiple security cameras.

[0677] Police Database:

[0678] It is a database containing personally identifiable information.

[0679] Specific actions

[0680] 1. Data Entry and Emotion Recognition

[0681] The user uploads a photo of the individual to be tracked to the device. The device receives the photo data and temporarily stores it in local storage. The device's built-in emotion engine then captures the user's face with a camera and generates emotion data in real time using an emotion recognition algorithm, such as TensorFlow. This emotion data is also sent to the server.

[0682] 2. Preprocessing of photo data

[0683] The server preprocesses the received photo data using the OpenCV library. Specifically, it converts the image to grayscale and performs face detection. It then uses the Haar Cascade and Dlib libraries to detect the face area and identify feature points such as the eyes, nose, and mouth to generate facial feature data. At the same time, it analyzes clothing features (color, pattern, shape, etc.) from the photo data and generates clothing feature data. Clothing analysis libraries such as DeepFashion are used for this.

[0684] 3. Video Search and Database Matching

[0685] The server accesses the security camera video database and utilizes SQL queries and NoSQL database technologies (e.g., MongoDB). It searches the video data based on facial and clothing feature data. It uses FaceNet or DeepFace models for facial recognition and similarity search algorithms (e.g., Cosine Similarity) for clothing matching. The server identifies matching individuals from the search results and extracts their video frames.

[0686] 4. Recording timestamps and camera positions

[0687] The server records the timestamp and camera location information for each video frame of the identified person, and stores the person's movement path in a detailed database.The server efficiently indexes the timestamp and location information using Elasticsearch.

[0688] 5. Behavioral Pattern Analysis

[0689] The server uses the collected timestamps and location information to analyze the target person's behavioral patterns using machine learning models (e.g., time series analysis methods and LSTM). It generates detailed information such as the person's route, length of stay, and frequently visited places. The server also incorporates the user's emotional data into the analysis and adjusts the analysis results.

[0690] 6. Identity Verification

[0691] After completing the behavioral pattern analysis, the server compares the facial feature data of the target person with the police database, using REST API or SOAP to access the police database and confirm the target person's identity based on the comparison results.

[0692] 7. Report generation and presentation

[0693] The server compiles the analysis results, emotion data, and identity verification results to generate a detailed report, possibly using Jupyter Notebook or ReportLab. The device receives the generated report and presents it to the user in the form of a dashboard or PDF report.

[0694] Specific examples

[0695] For example, to track a suspicious individual in a public facility, a user uploads a photo of the suspicious individual to the system. The device receives the photo data and sends it to the server, generating the user's emotional data using an emotion engine. The server extracts facial and clothing features from the uploaded photo and searches a security camera footage database to detect the individual. The server then analyzes the identified individual's behavioral patterns, determines their route of travel and where they stayed, and finally verifies their identity by comparing it with a police database. A detailed report is generated and provided to the user via the device. This report is presented in an easy-to-read format, taking into account the user's emotional data.

[0696] Example prompts for generative AI models

[0697] "We are developing a system to track suspicious individuals in public facilities. Please tell us how to build a program that inputs photographic data and searches a security camera footage database using the facial and clothing data obtained. Also, please explain in detail the procedure for analyzing the individual's behavioral patterns obtained in this way and comparing them with the police database to confirm their identity."

[0698] In this way, the system improves the efficiency and accuracy of crime prevention activities while providing information that takes into consideration the user's emotions.

[0699] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0700] Step 1:

[0701] The user uploads a photo of the individual to be tracked to the device. The device receives this photo data and temporarily stores it in local storage. Next, the device's built-in emotion recognition engine captures the user's face with a camera and generates emotion data using an emotion recognition algorithm. TensorFlow is used for this. This emotion data is sent to the server along with the photo data. The input is the individual's photo and the user's facial image, and the output is the photo data and emotion data.

[0702] Step 2:

[0703] The server preprocesses the received photo data using the OpenCV library. Specifically, it converts the image to grayscale and performs face detection. Here, the Haar Cascade and Dlib libraries are used. The face area is detected and feature points such as the eyes, nose, and mouth are identified to generate facial feature data. At the same time, the server uses a clothing analysis library such as DeepFashion to analyze clothing features (color, pattern, shape, etc.) from the photo data and generate clothing feature data. The input is photo data, and the output is facial feature data and clothing feature data.

[0704] Step 3:

[0705] The server accesses the security camera video database and utilizes SQL queries and NoSQL database technology (e.g., MongoDB). It searches security camera footage based on facial feature data and clothing feature data. It uses FaceNet or DeepFace models for facial recognition and a similarity search algorithm (e.g., Cosine Similarity) for clothing matching. It identifies matching people from the search results and extracts their video frames. The input is facial feature data and clothing feature data, and the output is video frames of matching people.

[0706] Step 4:

[0707] The server records the timestamp and camera location information for each video frame of the identified person and stores the person's movement path in a database. Here, we use Elasticsearch to index the timestamp and location information. The input is the video frame of the matching person, and the output is a record of the timestamp and camera location information.

[0708] Step 5:

[0709] The server analyzes the target person's behavioral patterns based on the collected timestamps and location information. This uses machine learning models (e.g., time series analysis methods and LSTM). The behavioral pattern analysis includes information such as travel routes, length of stay, and frequently visited places. In addition, the analysis results are adjusted taking into account the user's emotional data. The inputs are timestamps, camera location information, and emotional data, and the output is the analyzed behavioral patterns.

[0710] Step 6:

[0711] After completing the behavioral pattern analysis, the server compares the facial feature data of the target person with the police database. It accesses the police database using REST API or SOAP. Based on the comparison results, the target person's identity is confirmed. The input is facial feature data, and the output is the identity confirmation result.

[0712] Step 7:

[0713] The server compiles the analysis results, emotion data, and identity verification results to generate a detailed report. Jupyter Notebook or ReportLab can be used for this. The generated report is then presented to the user via their terminal. The information is presented in the form of a dashboard or PDF report. The input is the analysis results, emotion data, and identity verification results, and the output is a detailed report.

[0714] (Application example 2)

[0715] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0716] Conventional security systems were able to track individuals by facial recognition and extracting clothing characteristics, but they were unable to take the user's emotional state into account, making it difficult for them to provide optimal information to users in stressful situations. There was also a need for a system that could track individuals in real time and respond appropriately based on the user's emotions.

[0717] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a clothing feature extraction means, a means for accessing a security camera video database, a means for recording timestamps and camera locations, a means for analyzing behavioral patterns, a means for verifying identity by checking against a police database, a means for incorporating an emotion engine into the user terminal, and a means for utilizing the user's emotion data for analysis and presenting optimized information. This enables accurate tracking and analysis of individuals as well as the provision of information that takes the user's emotions into consideration.

[0718] "Facial recognition means" refers to devices and technology that detect an individual's face from input photographic data, analyze its features, and digitize them.

[0719] "Clothing feature extraction means" refers to a device and technology for analyzing the features of an individual's clothing, such as color, pattern, and shape, from input photographic data and converting them into data.

[0720] "Means for accessing a security camera video database" refers to devices and technologies for accessing video data collected from security cameras and obtaining the necessary information.

[0721] "Means for searching for video based on identified feature data" refers to devices and technologies for searching for matching video within a security camera video database based on feature data obtained through facial recognition or clothing feature extraction.

[0722] "Means for recording timestamps and camera locations" refers to devices and techniques for recording the date and time (timestamp) when the image was recorded and the location of the camera for each frame in which an identified person appears.

[0723] "Means for analyzing behavioral patterns" refers to devices and technology for analyzing behavioral patterns, such as the route a target person took and how long they stayed in a particular location, based on timestamps and camera position data.

[0724] "Means for verifying identity by comparing with police databases" refers to devices and technologies for verifying the identity of a person by comparing the acquired facial feature data with personal identification information held by the police.

[0725] The "means for generating and presenting the analysis results as a report" refers to a device and technology for integrating the above analysis results and generating and presenting them as a report in a format that is easy for the user to understand.

[0726] "Means for incorporating an emotion engine into a user terminal" refers to devices and techniques for incorporating emotion recognition software and hardware into a user terminal.

[0727] "Means for utilizing user emotional data for analysis and presenting optimized information" refers to devices and technologies that analyze the user's emotional state, adjust the way information is presented based on the results, and provide data in a form that is more suitable for the user.

[0728] The present invention is a system that includes a facial recognition unit, a clothing feature extraction unit, a security camera video database access unit, a video search unit based on identified feature data, a timestamp and camera location recorder, a behavioral pattern analysis unit, a police database for identifying the user, a report generating and presenting the analysis results, an emotion engine embedded in the user's device, and a unit for analyzing and presenting optimized information based on the user's emotion data. This system enables accurate tracking of individuals and the provision of information based on the user's emotion.

[0729] Hardware and software used

[0730] Hardware: Smartphones, smart glasses

[0731] Software: OpenCV (image processing library), EmotionRecognition (emotion analysis engine), police database API client

[0732] Data processing and calculation flow

[0733] A user uploads a photo of a specific individual to the device and simultaneously captures the user's emotional state. When the device transmits the photo data to the server, it uses an emotion engine to generate the user's emotional data, which is also transmitted to the server.

[0734] The server uses OpenCV to perform facial recognition based on the received photo data. It extracts the facial area from the photo and digitizes its feature points (the positions of the eyes, nose, mouth, etc.). At the same time, it extracts clothing features from the photo data and generates clothing feature data. This data is used to search the security camera video database and identify footage of matching people.

[0735] The system records the timestamp and camera position for each frame in which the identified person appears, thereby clarifying the person's movement path. The server then performs behavioral analysis to analyze the person's movement patterns and the length of time they stayed in the area. Finally, the results of this analysis are compared with a police database to confirm the person's identity.

[0736] The analysis results are integrated with the user's emotional data to generate a final report, which is presented in an easy-to-understand format for the user, enabling appropriate responses that take into account the user's emotional state.

[0737] Specific example explanation

[0738] For example, if a suspicious person appears at a concert venue, a security officer can upload a photo of the person using their smartphone to the system. The system then sends the photo to a server, which performs facial recognition and clothing feature extraction, then searches the security camera footage database to analyze the person's behavioral patterns.

[0739] If a security officer experiences high stress levels, the system can provide additional support in real time to allow for a rapid response. This process can be facilitated by prompts such as:

[0740] Prompt Sentence Examples

[0741] "Track the person in this photo, verify their identity against police databases, and suggest appropriate actions based on the user's current emotions."

[0742] In this way, the present invention improves the efficiency and accuracy of crime prevention activities and enables the provision of information that takes into consideration the user's emotions.

[0743] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0744] Step 1:

[0745] A user uploads a photo of a specific individual to a smartphone. The input is photo data that clearly shows the face. The output is the photo data that is saved on the device.

[0746] Step 2:

[0747] The device recognizes the user's emotional state from video or image data and generates emotional data using an emotion engine. The input is the user's video or image data, which is analyzed and emotional data (e.g., stress level, tension, etc.) is output.

[0748] Step 3:

[0749] The terminal sends photo data and emotion data to the server. The input at this time is the photo data and emotion data. The output is the data received by the server.

[0750] Step 4:

[0751] The server performs face recognition using the received photo data. The software used is OpenCV. The input is the photo data, the face is detected, and data on specific feature points (eyes, nose, mouth) is output.

[0752] Step 5:

[0753] The server extracts clothing features from the photo data, including information on color, pattern, shape, etc. The input is the photo data, and the output is clothing feature data.

[0754] Step 6:

[0755] The server searches a security camera video database for matching video based on the facial feature data and clothing feature data. The input is the facial feature data and clothing feature data, and the output is the matching video data.

[0756] Step 7:

[0757] The server records the timestamp and camera location of the video in which the matching person appears. The input is the matching video data, and the output is the timestamp and camera location information.

[0758] Step 8:

[0759] The server analyzes the target person's behavioral patterns based on the timestamp and camera position information. For example, the server outputs the person's route and duration of stay. The input is the timestamp and camera position information.

[0760] Step 9:

[0761] The server compares the results of behavioral pattern analysis with the police database to confirm the identity of the person. The input is facial feature data, and the matching results from the police database are output.

[0762] Step 10:

[0763] The server integrates the analysis results with the user's emotional data and generates a final report. The inputs are the behavioral pattern analysis results and the emotional data, and the output is a report.

[0764] Step 11:

[0765] The terminal presents this generated report to the user. The input is the report sent from the server, and the output is to display it in a user-friendly format.

[0766] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0767] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0768] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0769] [Third embodiment]

[0770] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0771] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0772] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0773] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0774] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0775] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0776] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0777] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0778] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0779] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0780] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0781] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0782] The present invention is a system that extracts facial recognition and clothing characteristics from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and further verifies their identity in cooperation with a police database. Specific embodiments of the present invention are described in detail below.

[0783] System Overview

[0784] The system consists of the following main components:

[0785] 1. User terminal: The interface where the user inputs photo data and receives the analysis results.

[0786] 2. Server: Responsible for central processing such as data processing, facial recognition, behavioral analysis, and database matching.

[0787] 3. Security camera video database: A database that stores video data collected from multiple security cameras.

[0788] 4. Police databases: Databases containing personally identifiable information.

[0789] System operation and program processing

[0790] 1. Data Entry

[0791] The user uploads a photo of the individual to be tracked to the device. This photo data includes both the face and clothing. The device then sends the photo data to the server.

[0792] 2. Pretreatment

[0793] The server receives this photo data and first extracts the facial region using a facial recognition algorithm, from which it identifies the following feature points:

[0794] Eye, nose, and mouth position

[0795] Facial contours

[0796] Other features (mole, scar, etc.)

[0797] At the same time, the server extracts the clothing's characteristics, specifically analyzing its color, pattern, shape, etc., and generates feature data.

[0798] 3. Video search and feature data matching

[0799] The server accesses a security camera video database and searches for video data based on the identified feature data. From the security camera video, it performs facial recognition and clothing feature matching to identify matching individuals.

[0800] 4. Recording timestamps and camera positions

[0801] For each frame in which a matching person is identified, the server records the timestamp and camera position, allowing the path of the person to be traced.

[0802] 5. Behavioral Pattern Analysis

[0803] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, the following information is generated:

[0804] Travel route

[0805] Stay time

[0806] Frequently visited places

[0807] This information is organized as a result of behavioral pattern analysis.

[0808] 6. Identity Verification

[0809] After the behavioral pattern analysis is completed, the server will compare the target person's facial feature data with the police database, and based on the matching results, the target person's identity will be confirmed.

[0810] 7. Report generation and presentation

[0811] Finally, the server generates a detailed report containing the results of the analysis and identification, including the following information:

[0812] Travel route

[0813] Behavioral patterns

[0814] Timestamp and camera position

[0815] Identity verification results

[0816] The terminal presents this report to the user and provides the necessary information.

[0817] Specific examples

[0818] For example, to track a suspicious person in a public facility, a user uploads a photo of the suspicious person to the system. The server extracts facial and clothing features from the photo and accesses a security camera footage database to detect the person. The server then analyzes the person's behavioral patterns to identify their route of travel and where they are staying. Finally, the server checks the data against a police database to confirm their identity, and a detailed report is provided to the user. This allows for accurate and rapid tracking of people.

[0819] The processing flow will be explained below.

[0820] Step 1:

[0821] The user uploads a photo of the individual to be tracked to the device, which receives and stores the photo data.

[0822] Step 2:

[0823] The device sends the photo data to the server, which receives it and applies a facial recognition algorithm.

[0824] Step 3:

[0825] The server uses a facial recognition algorithm to extract facial regions from the photograph, identify specific feature points (such as the eyes, nose, and mouth) from the facial regions, and generate facial feature data.

[0826] Step 4:

[0827] At the same time, the server analyzes the clothing characteristics (color, pattern, shape, etc.) from the photograph data and generates clothing characteristic data.

[0828] Step 5:

[0829] The server accesses the security camera video database and obtains the latest video data.

[0830] Step 6:

[0831] The server performs facial recognition and clothing feature matching on the acquired video data to search for people who match specific feature data.

[0832] Step 7:

[0833] The server records the timestamp and camera position for each frame in which a matching person is identified.

[0834] Step 8:

[0835] The server uses multiple timestamps and camera location information to analyze the target person's behavioral characteristics, such as their route of travel and length of stay.

[0836] Step 9:

[0837] The server organizes the analysis results and extracts behavioral patterns (frequently visited places, places where people spend long periods of time, etc.).

[0838] Step 10:

[0839] The server uses the analyzed facial feature data to access police databases and verify the identity of the person in question.

[0840] Step 11:

[0841] The server adds the matching results to the report and combines all the analysis results to generate the final report.

[0842] Step 12:

[0843] The terminal presents the final report to the user, who then checks the report and obtains the necessary information.

[0844] Example 1

[0845] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0846] In modern society, numerous security cameras and surveillance devices are installed, helping to prevent and investigate crime. However, quickly and accurately identifying and tracking specific individuals from large amounts of video data is difficult, time-consuming, and laborious. To solve this problem, there is a need for a system that can extract facial recognition and clothing characteristics from photographic data, track and analyze the behavior of target individuals using a security camera video database, and further verify their identity in conjunction with an identification database.

[0847] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0848] In this invention, the server includes a face recognition means, a clothing feature extraction means, a means for accessing a video database, a means for searching video data based on the identified feature data, a means for recording timestamps and location information, a means for analyzing behavioral patterns, a means for verifying identity by comparing with an identification database, a means for generating and presenting analysis results as a report, a user interface including a means for uploading photo data, and a means for using a generative AI model as part of the analysis process. This makes it possible to quickly and accurately identify a target person from photo data, analyze the person's behavioral patterns, and verify the person's identity in cooperation with the identification database.

[0849] "Facial recognition means" refers to a device or program that uses technology or algorithms to detect facial feature points and identify individuals.

[0850] "Clothing feature extraction means" refers to a device or program that uses technology or algorithms to analyze and extract clothing features such as color, pattern, and shape from image data.

[0851] "Means for accessing a video database" refers to a device or program that connects to a database in which video data collected from multiple cameras is stored and acquires the necessary data.

[0852] "Means for searching video data based on identified feature data" refers to devices or programs that use technology or algorithms to search for matching people from videos in a video database using extracted facial and clothing feature data.

[0853] The "means for recording timestamp and location information" refers to a device or program that acquires and records the shooting date and time of the searched video frame and the installation location of the camera.

[0854] "Means for analyzing behavioral patterns" refers to devices or programs that use technologies or algorithms to analyze a target person's behavioral patterns, such as their travel routes, length of stay, and frequently visited places, based on recorded timestamps and location information.

[0855] "Means for verifying identity by matching against an identification database" refers to devices or programs that use technology or algorithms to compare analyzed facial feature information with an identification database and verify the identity of the person based on matching data.

[0856] The "means for generating and presenting the analysis results as a report" refers to a device or program that organizes the results of behavioral pattern analysis and identity verification, and generates a report that can be visually presented to the user.

[0857] The "user interface including a means for uploading photographic data" is an interface that allows a user to input photographic data into the system via a terminal.

[0858] "Means of using a generative AI model as part of the analytical process" refers to devices or programs that utilize generative AI models as part of the analysis, and use technologies or algorithms to achieve more advanced data processing and pattern recognition.

[0859] The present invention is a system that extracts facial recognition and clothing characteristics from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and verifies their identity in cooperation with an identification database. To implement this system, the following hardware and software are used.

[0860] System Overview

[0861] The system consists of the following main components:

[0862] 1. User terminal: An interface where the user inputs photo data and receives the analysis results.

[0863] 2. Server: Responsible for central processing such as data processing, facial recognition, behavior analysis, and database matching.

[0864] 3. Security camera video database: A database that stores video data collected from multiple security cameras.

[0865] 4. Identification Database: A database containing personally identifiable information.

[0866] Hardware and software used

[0867] Face recognition algorithm: Uses OpenCV and dlib libraries.

[0868] Clothing feature extraction algorithm: Uses color analysis using RGB values, texture analysis, and contour detection algorithms.

[0869] Video search algorithm: Facenet and DeepFace are used for face recognition, and Elasticsearch's image search plugin is used for clothing matching.

[0870] Behavioral analysis algorithms: Use custom algorithms to plot movement paths and analyze dwell times.

[0871] Identity database matching algorithm: Uses face identification model to match with identity database.

[0872] Detailed procedure

[0873] The user uploads photo data of the object to be tracked to the device, which then sends the uploaded photo data to the server. The server then uses a facial recognition algorithm to detect the facial area and extract feature points such as the eyes, nose, mouth, facial contours, moles, and scars.

[0874] At the same time, the server analyzes the color, pattern, and shape of the clothing to generate feature data. Based on the analyzed feature data, it accesses a security camera video database to search for matching video data. The server then obtains and records the timestamp and camera position of the video frame of the matching person.

[0875] The system then analyzes the person's behavioral patterns based on timestamps and camera location data to identify their route and location. Finally, the system compares the facial feature data with an identification database to confirm their identity. The server then processes the analysis results and generates a detailed report that is presented to the user's device.

[0876] Specific examples

[0877] For example, when tracking a suspicious person in a public facility, the user uploads photo data to the system. The server extracts facial and clothing features, searches a security camera footage database, and analyzes the suspicious person's behavioral patterns. The server then verifies the person's identity against an identification database and provides a detailed report to the user.

[0878] Prompt Sentence Examples

[0879] "Using this photo, please briefly explain the entire process of identifying the person in question by extracting facial and clothing characteristics, searching a security camera footage database, and analyzing their behavioral patterns to confirm their identity."

[0880] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0881] Step 1: Data entry

[0882] The user uploads photographic data of the individual to be tracked to the device, which includes information on their face and clothing.

[0883] Input: Photo data (face and clothing information)

[0884] Output: Photo data sent to the server

[0885] Specifically, the user selects photo data using the file upload function of the device and clicks the upload button. The device then sends the selected photo data to the server as an HTTP request.

[0886] Step 2: Face recognition and clothing feature extraction (preprocessing)

[0887] The server receives this photo data and extracts the face region using a facial recognition algorithm, for example, using the face detection function of OpenCV or dlib.

[0888] Input: Submitted photo data

[0889] Output: Facial feature points (eyes, nose, mouth position, facial contours, moles, scars, etc.)

[0890] Specifically, the server reads the photo data using an image processing library, detects the facial area, and extracts feature points.

[0891] At the same time, the server extracts the clothing's features by analyzing its color (RGB value), pattern (texture analysis), and shape (contour detection) to generate feature data.

[0892] Input: Submitted photo data

[0893] Output: Clothing feature data (color, pattern, shape)

[0894] Specifically, the server extracts clothing feature data using a color analysis algorithm, a texture analysis algorithm, and a shape analysis algorithm.

[0895] Step 3: Feature data matching and video search

[0896] The server then accesses a security camera video database based on the identified facial and clothing feature data and searches for the relevant video data, using Facenet and DeepFace for facial recognition and Elasticsearch's image search plugin for clothing matching.

[0897] Input: facial feature points, clothing feature data

[0898] Output: Video data of the matching person

[0899] Specifically, the server queries a security camera video database to search for video frames that match the facial features and clothing data.

[0900] Step 4: Record timestamps and camera positions

[0901] For each frame in which a matching person is identified, the server captures and records the timestamp and camera position from the video data.

[0902] Input: Video data of the person to match

[0903] Output: timestamp and camera position information

[0904] Specifically, the server extracts and records the timestamp and camera position from the metadata of the matched video frame.

[0905] Step 5: Behavioral pattern analysis

[0906] The server analyzes the target person's behavioral patterns based on the recorded timestamps and camera location information, plotting their movement paths, analyzing their stay times, and identifying frequently visited locations.

[0907] Input: timestamp, camera location information

[0908] Output: Behavioral pattern analysis results (route, stay time, frequently visited places)

[0909] Specifically, the server analyzes the timestamp and camera position data, plots the target person's movement path, and calculates the length of time they stayed there.

[0910] Step 6: Identity Verification

[0911] After completing the behavioral pattern analysis, the server compares the target person's facial feature data with the identification database using a facial identification model (such as Facenet or DeepFace).

[0912] Input: Facial feature data

[0913] Output: Identity verification result

[0914] Specifically, the server accesses an identification database, matches the facial feature data, and confirms the identity of the matching individual.

[0915] Step 7: Report generation and presentation

[0916] Finally, the server generates a detailed report containing the analysis results and the identification results, including the route of travel, behavioral patterns, timestamps, camera locations, and the identification results. The device then presents this report to the user.

[0917] Input: behavioral pattern analysis results, identity verification results

[0918] Output: Detailed report

[0919] Specifically, the server compiles the movement route, behavioral patterns, timestamps, camera locations, and identity verification results to generate a report, which is then displayed on the device for the user to view.

[0920] Prompt Sentence Examples

[0921] "Using this photo, please briefly explain the entire process of identifying the person in question by extracting facial and clothing characteristics, searching a security camera footage database, and analyzing their behavioral patterns to confirm their identity."

[0922] (Application example 1)

[0923] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0924] Existing systems that use security camera footage and police databases to track and identify suspicious individuals have difficulty informing administrators of suspicious individuals' movements in real time. Furthermore, delays in reporting the results of the analysis mean that appropriate countermeasures cannot be implemented promptly.

[0925] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0926] In this invention, the server includes a facial recognition unit, a clothing feature extraction unit, a unit for accessing a security camera video database, a unit for searching for video based on the identified feature data, a unit for recording timestamps and camera locations, a unit for analyzing behavioral patterns, a unit for verifying identity by checking against a police database, a unit for generating and presenting a report of the analysis results, and a unit for sending data to a smartphone application to notify the administrator. This makes it possible to notify the location and behavioral patterns of suspicious individuals in real time.

[0927] "Facial recognition means" is a technology that extracts facial feature points from an input facial image and identifies an individual using a specific algorithm.

[0928] "Clothing feature extraction means" is a technology that analyzes features such as clothing color, shape, and pattern from an input image and extracts them as data.

[0929] "Means for accessing a security camera video database" refers to technology that accesses a database that stores video data collected from multiple security cameras, and performs searches and retrieval.

[0930] "Means for searching for footage based on identified feature data" refers to a technology that searches security camera footage in a database based on extracted facial and clothing feature data, and identifies the person in question.

[0931] The "means for recording timestamps and camera positions" is a technology for recording the time information of a specified video frame and the location where the camera is installed.

[0932] "Means for analyzing behavioral patterns" refers to technology that analyzes behavioral patterns such as a person's movement route and length of stay based on timestamps and camera position information.

[0933] "Means of confirming identity by comparing with police database" refers to a technology that compares extracted facial feature data with a database held by the police to confirm the identity of the target person.

[0934] "Means for generating and presenting analysis results as a report" refers to a technology that creates a detailed report based on the results of behavioral pattern analysis and identity verification and provides it to the user.

[0935] "Means for sending data to a smartphone application and notifying the user" refers to a technology that sends generated reports and analysis results to a smartphone application in real time and notifies the user.

[0936] The present invention is a system including a facial recognition unit, a clothing feature extraction unit, a security camera video database access unit, a video search unit based on the identified feature data, a timestamp and camera location recording unit, a behavioral pattern analysis unit, a police database comparison unit to confirm identity, a report generating and presenting the analysis results, and a smartphone application that notifies the user of the analysis results. This system is suitable for quickly and accurately tracking and identifying suspicious individuals using security camera video and the police database.

[0937] Face recognition and clothing feature extraction

[0938] The server receives the photo data uploaded by the user to the device and extracts the face area using a face recognition algorithm (e.g., OpenCV or Google Vision API), identifies facial feature points (positions of eyes, nose, mouth, facial contours, and other features), and generates data on clothing features (color, pattern, shape, etc.).

[0939] Video database search and matching

[0940] The server accesses a security camera video database and searches for video data based on the extracted feature data. It then performs facial recognition and clothing feature matching on the security camera footage to identify matching individuals. For each identified frame, the server records the timestamp and camera location, and tracks the target individual's movement path.

[0941] Behavioral pattern analysis and identity verification

[0942] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, it generates information such as the person's route of travel, length of stay, and frequently visited places. The server then organizes the analysis results. The server then compares the target person's facial feature data with the police database to confirm their identity.

[0943] Report generation and notification

[0944] Finally, the server generates a detailed report containing the analysis and identity verification results and sends a notification to the smartphone application, where the user can view the report and take any necessary measures.

[0945] For example, if a public facility manager spots someone behaving suspiciously in a parking lot and uploads a photo of them to a smartphone application, the system will review security camera footage and analyze the suspicious person's movement path. It will then compare the person's identity with a police database and confirm their identity. The manager will receive a real-time notification, enabling them to respond quickly.

[0946] Prompt Sentence Examples

[0947] "Upload a photo of someone behaving suspiciously in your parking lot. We will then extract the suspicious person's facial and clothing characteristics from the photo and search our security camera footage database for matching footage. We will then confirm the person's behavioral patterns and identity, and generate a detailed report to notify you."

[0948] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0949] Step 1:

[0950] A user uploads a photo of a suspicious person using a smartphone application. The input is the photo data of the suspicious person, and the output is the photo data sent to the server.

[0951] Step 2:

[0952] The server analyzes the received photo data and extracts the face region using a face recognition algorithm (e.g., OpenCV or Google Vision API). The input is the photo data, and the output is facial feature points (positions of eyes, nose, mouth, facial contours, and other features).

[0953] Step 3:

[0954] At the same time, the server extracts clothing features from the photo data. Specifically, it analyzes color, pattern, and shape and generates feature data. The input is the photo data, and the output is clothing feature data.

[0955] Step 4:

[0956] The server accesses the security camera video database and searches the video data based on the extracted facial and clothing feature data. The input is the facial and clothing feature data, and the output is the video frame of the matching person.

[0957] Step 5:

[0958] The server records the timestamp and camera position for each frame in which a matching person is identified. The input is the matching video frame, and the output is a record of the time and camera position information.

[0959] Step 6:

[0960] The server analyzes the target person's behavioral patterns based on timestamps and camera location information. The input is time information and camera location information, and the output is behavioral pattern information such as travel route, length of stay, and frequently visited places.

[0961] Step 7:

[0962] The server compares the facial feature data with the police database to confirm the identity of the person. The input is the facial feature data, and the output is the identity confirmation result.

[0963] Step 8:

[0964] The server generates a detailed report including the analysis results and identity verification results, and notifies the smartphone application of the report. The input is behavioral pattern information and identity verification results, and the output is a report sent to the user's device.

[0965] Step 9:

[0966] The user checks the received report through a smartphone application and takes necessary measures. The input is the report from the server, and the output is the countermeasure action taken by the user.

[0967] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0968] The present invention combines a system that extracts facial recognition and clothing features from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and verifies their identity in cooperation with a police database with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention will be described in detail below.

[0969] System Overview

[0970] The system consists of the following main components:

[0971] 1. User Device:

[0972] An interface where users can input photo data and receive analysis results.

[0973] An emotion engine that recognizes user emotions and collects emotion data.

[0974] 2. Server:

[0975] It handles core processing such as data processing, facial recognition, behavioral analysis, and database matching.

[0976] Emotional data from the emotion engine is used for analysis to optimize the way information is presented to users.

[0977] 3. Security camera footage database:

[0978] A database that stores video data collected from multiple security cameras.

[0979] 4. Police Database:

[0980] Databases containing personally identifiable information.

[0981] System operation and program processing

[0982] 1. Data Entry

[0983] The user uploads a photo of the individual to be tracked to the device. This photo data includes both the face and clothing. The device sends the photo data to the server and recognizes the user's emotions to generate emotion data. This emotion data is also sent to the server.

[0984] 2. Pretreatment

[0985] The server receives the photo data and first extracts the facial region using a facial recognition algorithm. From this facial region, it identifies specific feature points such as the eyes, nose, and mouth, and generates facial feature data. At the same time, the server analyzes the clothing features (color, pattern, shape, etc.) from the photo data and generates clothing feature data.

[0986] 3. Video search and feature data matching

[0987] The server accesses a security camera video database and searches for video data based on the identified feature data. From the security camera video, it performs facial recognition and clothing feature matching to identify matching individuals.

[0988] 4. Recording timestamps and camera positions

[0989] For each frame in which a matching person is identified, the server records the timestamp and camera position, allowing the path of the person to be traced.

[0990] 5. Behavioral Pattern Analysis

[0991] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, it generates information such as the person's route, length of stay, and frequently visited places. This analysis also takes into account the user's emotional data.

[0992] 6. Identity Verification

[0993] After the behavioral pattern analysis is completed, the server will compare the target person's facial feature data with the police database, and based on the matching results, the target person's identity will be confirmed.

[0994] 7. Report generation and presentation

[0995] Finally, the server generates a detailed report containing the analysis results, emotion data, and identity verification results. This report includes behavioral patterns, timestamps, camera locations, identity verification results, and information presentation methods adjusted according to the emotion data. The device then presents this report to the user and provides the necessary information.

[0996] Specific examples

[0997] For example, to track a suspicious person in a public facility, a user uploads a photo of the suspicious person to the system. The device receives the photo data and sends it to the server, recognizing the user's emotions to generate emotion data. The server extracts facial and clothing features from the photo and accesses a security camera footage database to detect the person. The server then analyzes the person's behavioral patterns and identifies their route and where they stayed. Finally, it compares the person with a police database to confirm their identity, and a detailed report is provided to the user. This report is based on the user's emotion data and presented in an easy-to-understand format. This enables accurate and fast person tracking.

[0998] In this way, the present invention improves the efficiency and accuracy of crime prevention activities and enables the provision of information that takes into consideration the user's emotions.

[0999] The processing flow will be explained below.

[1000] Step 1:

[1001] The user uploads a photo of the individual to be tracked to the device, which receives the photo data and sends it to the server.

[1002] Step 2:

[1003] The device activates an emotion engine to recognize the user's emotions, collecting emotion data from the user's voice, facial expressions, input data, etc.

[1004] Step 3:

[1005] The device sends the collected emotion data to the server, which receives the photo data and emotion data and begins processing.

[1006] Step 4:

[1007] The server uses a facial recognition algorithm to extract facial regions from the photograph, and then identifies feature points such as the eyes, nose, and mouth from the facial regions to generate facial feature data.

[1008] Step 5:

[1009] The server analyzes the clothing characteristics (color, pattern, shape, etc.) from the photograph data and generates clothing characteristic data.

[1010] Step 6:

[1011] The server accesses the security camera video database and searches the video data based on the acquired feature data. It then performs facial recognition and clothing feature matching on the security camera video to identify matching individuals.

[1012] Step 7:

[1013] The server records the timestamp and camera position for each frame in which a matching person is identified.

[1014] Step 8:

[1015] The server uses the recorded timestamps and camera positions to analyze the target person's movement route, duration of stay, frequently visited places, etc.

[1016] Step 9:

[1017] The server organizes the analysis results and extracts behavioral patterns, summarizing the information in a format that is easy for users to understand.

[1018] Step 10:

[1019] The server uses the facial feature data of the target person to compare it with the police database to confirm their identity, and obtains the matching results.

[1020] Step 11:

[1021] The server integrates the results of behavioral analysis and identity verification into a report, adjusts the report content based on the user's emotional data, and optimizes the information presentation method as needed.

[1022] Step 12:

[1023] The terminal presents the final report to the user, who then checks the report and obtains the necessary information.

[1024] Example 2

[1025] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1026] Conventional security systems are capable of tracking individuals and analyzing their behavior, but they do not provide information that takes into account the user's emotional state, and they are unable to present information in an efficient and easy-to-understand manner. Furthermore, they do not adequately utilize user emotional data to improve the accuracy of behavioral pattern analysis. As a result, tracking results and analytical information are sometimes not practical for users, and there is a need to improve the efficiency and accuracy of security activities.

[1027] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1028] In this invention, the server includes a facial recognition means, a clothing feature extraction means, a means for accessing a security camera video database, a means for searching for video based on the identified feature data, a means for recording a timestamp and camera location, a means for analyzing behavioral patterns, a means for verifying identity by checking against a police database, a means for generating and presenting a report of the analysis results, a means for recognizing a user's emotions using an emotion recognition engine and generating emotion data, and a means for optimizing information presentation by utilizing the generated emotion data in analysis, thereby enabling improved efficiency and accuracy in crime prevention activities.

[1029] The "face recognition means" is a means for detecting an individual's face area from image data, identifying feature points (eyes, nose, mouth, etc.), and extracting facial feature data.

[1030] The "clothing feature extraction means" is a means for analyzing clothing features (color, pattern, shape, etc.) from image data and generating clothing feature data.

[1031] "Means for accessing a security camera video database" refers to a means for connecting to a database that stores video data collected from a large number of security cameras and obtaining the necessary video data from there.

[1032] The "means for searching for video based on identified feature data" refers to a means for searching for video in a security camera video database based on the extracted facial feature data and clothing feature data, and identifying matching individuals.

[1033] The "means for recording timestamps and camera locations" refers to a means for recording the time of capture and the camera location information for each video frame in which a matching person is identified.

[1034] "Means for analyzing behavioral patterns" refers to a method for analyzing a target person's movement route, length of stay, frequently visited places, etc. from recorded timestamps and camera location information to analyze behavioral patterns.

[1035] "Means for confirming identity by checking against a police database" refers to a means for comparing facial feature data after behavioral pattern analysis with a police database and confirming the identity of the target person based on the matching results.

[1036] The "means for generating and presenting analysis results as a report" refers to a means for compiling behavioral analysis results, analysis results based on emotional data, and comparison results with police databases to create a detailed report and provide it to the user.

[1037] The "means for recognizing a user's emotions using an emotion recognition engine and generating emotion data" refers to a means for taking a picture of the user's face, analyzing their emotional state through an emotion recognition algorithm, and generating the results as data.

[1038] "Means for optimizing information presentation by utilizing generated emotional data in analysis" refers to means for adjusting the analysis results and information presentation method based on the acquired emotional data, and providing information in a form that is most easily understandable for the user.

[1039] System Overview

[1040] This system extracts facial recognition and clothing characteristics from personal photo data, and tracks and analyzes the behavior of target individuals using a security camera video database. Furthermore, by linking with a police database to verify identity and combining it with an emotion engine that recognizes user emotions, it achieves efficient and highly accurate crime prevention activities.

[1041] Key Components

[1042] User device:

[1043] It provides an interface where users can input photo data and receive analysis results.

[1044] An emotion recognition engine is used to collect user emotion data and send it to a server.

[1045] server:

[1046] It handles core processing such as data processing, facial recognition, behavioral analysis, and database matching.

[1047] Emotional data from the emotion engine is used for analysis to optimize the way information is presented to users.

[1048] Security camera footage database:

[1049] It is a database that stores video data collected from multiple security cameras.

[1050] Police Database:

[1051] It is a database containing personally identifiable information.

[1052] Specific actions

[1053] 1. Data Entry and Emotion Recognition

[1054] The user uploads a photo of the individual to be tracked to the device. The device receives the photo data and temporarily stores it in local storage. The device's built-in emotion engine then captures the user's face with a camera and generates emotion data in real time using an emotion recognition algorithm, such as TensorFlow. This emotion data is also sent to the server.

[1055] 2. Preprocessing of photo data

[1056] The server preprocesses the received photo data using the OpenCV library. Specifically, it converts the image to grayscale and performs face detection. It then uses the Haar Cascade and Dlib libraries to detect the face area and identify feature points such as the eyes, nose, and mouth to generate facial feature data. At the same time, it analyzes clothing features (color, pattern, shape, etc.) from the photo data and generates clothing feature data. Clothing analysis libraries such as DeepFashion are used for this.

[1057] 3. Video Search and Database Matching

[1058] The server accesses the security camera video database and utilizes SQL queries and NoSQL database technologies (e.g., MongoDB). It searches the video data based on facial and clothing feature data. It uses FaceNet or DeepFace models for facial recognition and similarity search algorithms (e.g., Cosine Similarity) for clothing matching. The server identifies matching individuals from the search results and extracts their video frames.

[1059] 4. Recording timestamps and camera positions

[1060] The server records the timestamp and camera location information for each video frame of the identified person, and stores the person's movement path in a detailed database.The server efficiently indexes the timestamp and location information using Elasticsearch.

[1061] 5. Behavioral Pattern Analysis

[1062] The server uses the collected timestamps and location information to analyze the target person's behavioral patterns using machine learning models (e.g., time series analysis methods and LSTM). It generates detailed information such as the person's route, length of stay, and frequently visited places. The server also incorporates the user's emotional data into the analysis and adjusts the analysis results.

[1063] 6. Identity Verification

[1064] After completing the behavioral pattern analysis, the server compares the facial feature data of the target person with the police database, using REST API or SOAP to access the police database and confirm the target person's identity based on the comparison results.

[1065] 7. Report generation and presentation

[1066] The server compiles the analysis results, emotion data, and identity verification results to generate a detailed report, possibly using Jupyter Notebook or ReportLab. The device receives the generated report and presents it to the user in the form of a dashboard or PDF report.

[1067] Specific examples

[1068] For example, to track a suspicious individual in a public facility, a user uploads a photo of the suspicious individual to the system. The device receives the photo data and sends it to the server, generating the user's emotional data using an emotion engine. The server extracts facial and clothing features from the uploaded photo and searches a security camera footage database to detect the individual. The server then analyzes the identified individual's behavioral patterns, determines their route of travel and where they stayed, and finally verifies their identity by comparing it with a police database. A detailed report is generated and provided to the user via the device. This report is presented in an easy-to-read format, taking into account the user's emotional data.

[1069] Example prompts for generative AI models

[1070] "We are developing a system to track suspicious individuals in public facilities. Please tell us how to build a program that inputs photographic data and searches a security camera footage database using the facial and clothing data obtained. Also, please explain in detail the procedure for analyzing the individual's behavioral patterns obtained in this way and comparing them with the police database to confirm their identity."

[1071] In this way, the system improves the efficiency and accuracy of crime prevention activities while providing information that takes into consideration the user's emotions.

[1072] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1073] Step 1:

[1074] The user uploads a photo of the individual to be tracked to the device. The device receives this photo data and temporarily stores it in local storage. Next, the device's built-in emotion recognition engine captures the user's face with a camera and generates emotion data using an emotion recognition algorithm. TensorFlow is used for this. This emotion data is sent to the server along with the photo data. The input is the individual's photo and the user's facial image, and the output is the photo data and emotion data.

[1075] Step 2:

[1076] The server preprocesses the received photo data using the OpenCV library. Specifically, it converts the image to grayscale and performs face detection. Here, the Haar Cascade and Dlib libraries are used. The face area is detected and feature points such as the eyes, nose, and mouth are identified to generate facial feature data. At the same time, the server uses a clothing analysis library such as DeepFashion to analyze clothing features (color, pattern, shape, etc.) from the photo data and generate clothing feature data. The input is photo data, and the output is facial feature data and clothing feature data.

[1077] Step 3:

[1078] The server accesses the security camera video database and utilizes SQL queries and NoSQL database technology (e.g., MongoDB). It searches security camera footage based on facial feature data and clothing feature data. It uses FaceNet or DeepFace models for facial recognition and a similarity search algorithm (e.g., Cosine Similarity) for clothing matching. It identifies matching people from the search results and extracts their video frames. The input is facial feature data and clothing feature data, and the output is video frames of matching people.

[1079] Step 4:

[1080] The server records the timestamp and camera location information for each video frame of the identified person and stores the person's movement path in a database. Here, we use Elasticsearch to index the timestamp and location information. The input is the video frame of the matching person, and the output is a record of the timestamp and camera location information.

[1081] Step 5:

[1082] The server analyzes the target person's behavioral patterns based on the collected timestamps and location information. This uses machine learning models (e.g., time series analysis methods and LSTM). The behavioral pattern analysis includes information such as travel routes, length of stay, and frequently visited places. In addition, the analysis results are adjusted taking into account the user's emotional data. The inputs are timestamps, camera location information, and emotional data, and the output is the analyzed behavioral patterns.

[1083] Step 6:

[1084] After completing the behavioral pattern analysis, the server compares the facial feature data of the target person with the police database. It accesses the police database using REST API or SOAP. Based on the comparison results, the target person's identity is confirmed. The input is facial feature data, and the output is the identity confirmation result.

[1085] Step 7:

[1086] The server compiles the analysis results, emotion data, and identity verification results to generate a detailed report. Jupyter Notebook or ReportLab can be used for this. The generated report is then presented to the user via their terminal. The information is presented in the form of a dashboard or PDF report. The input is the analysis results, emotion data, and identity verification results, and the output is a detailed report.

[1087] (Application example 2)

[1088] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1089] Conventional security systems were able to track individuals by facial recognition and extracting clothing characteristics, but they were unable to take the user's emotional state into account, making it difficult for them to provide optimal information to users in stressful situations. There was also a need for a system that could track individuals in real time and respond appropriately based on the user's emotions.

[1090] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a clothing feature extraction means, a means for accessing a security camera video database, a means for recording timestamps and camera locations, a means for analyzing behavioral patterns, a means for verifying identity by checking against a police database, a means for incorporating an emotion engine into the user terminal, and a means for utilizing the user's emotion data for analysis and presenting optimized information. This enables accurate tracking and analysis of individuals as well as the provision of information that takes the user's emotions into consideration.

[1091] "Facial recognition means" refers to devices and technology that detect an individual's face from input photographic data, analyze its features, and digitize them.

[1092] "Clothing feature extraction means" refers to a device and technology for analyzing the features of an individual's clothing, such as color, pattern, and shape, from input photographic data and converting them into data.

[1093] "Means for accessing a security camera video database" refers to devices and technologies for accessing video data collected from security cameras and obtaining the necessary information.

[1094] "Means for searching for video based on identified feature data" refers to devices and technologies for searching for matching video within a security camera video database based on feature data obtained through facial recognition or clothing feature extraction.

[1095] "Means for recording timestamps and camera locations" refers to devices and techniques for recording the date and time (timestamp) when the image was recorded and the location of the camera for each frame in which an identified person appears.

[1096] "Means for analyzing behavioral patterns" refers to devices and technology for analyzing behavioral patterns, such as the route a target person took and how long they stayed in a particular location, based on timestamps and camera position data.

[1097] "Means for verifying identity by comparing with police databases" refers to devices and technologies for verifying the identity of a person by comparing the acquired facial feature data with personal identification information held by the police.

[1098] The "means for generating and presenting the analysis results as a report" refers to a device and technology for integrating the above analysis results and generating and presenting them as a report in a format that is easy for the user to understand.

[1099] "Means for incorporating an emotion engine into a user terminal" refers to devices and techniques for incorporating emotion recognition software and hardware into a user terminal.

[1100] "Means for utilizing user emotional data for analysis and presenting optimized information" refers to devices and technologies that analyze the user's emotional state, adjust the way information is presented based on the results, and provide data in a form that is more suitable for the user.

[1101] The present invention is a system that includes a facial recognition unit, a clothing feature extraction unit, a security camera video database access unit, a video search unit based on identified feature data, a timestamp and camera location recorder, a behavioral pattern analysis unit, a police database for identifying the user, a report generating and presenting the analysis results, an emotion engine embedded in the user's device, and a unit for analyzing and presenting optimized information based on the user's emotion data. This system enables accurate tracking of individuals and the provision of information based on the user's emotion.

[1102] Hardware and software used

[1103] Hardware: Smartphones, smart glasses

[1104] Software: OpenCV (image processing library), EmotionRecognition (emotion analysis engine), police database API client

[1105] Data processing and calculation flow

[1106] A user uploads a photo of a specific individual to the device and simultaneously captures the user's emotional state. When the device transmits the photo data to the server, it uses an emotion engine to generate the user's emotional data, which is also transmitted to the server.

[1107] The server uses OpenCV to perform facial recognition based on the received photo data. It extracts the facial area from the photo and digitizes its feature points (the positions of the eyes, nose, mouth, etc.). At the same time, it extracts clothing features from the photo data and generates clothing feature data. This data is used to search the security camera video database and identify footage of matching people.

[1108] The system records the timestamp and camera position for each frame in which the identified person appears, thereby clarifying the person's movement path. The server then performs behavioral analysis to analyze the person's movement patterns and the length of time they stayed in the area. Finally, the results of this analysis are compared with a police database to confirm the person's identity.

[1109] The analysis results are integrated with the user's emotional data to generate a final report, which is presented in an easy-to-understand format for the user, enabling appropriate responses that take into account the user's emotional state.

[1110] Specific example explanation

[1111] For example, if a suspicious person appears at a concert venue, a security officer can upload a photo of the person using their smartphone to the system. The system then sends the photo to a server, which performs facial recognition and clothing feature extraction, then searches the security camera footage database to analyze the person's behavioral patterns.

[1112] If a security officer experiences high stress levels, the system can provide additional support in real time to allow for a rapid response. This process can be facilitated by prompts such as:

[1113] Prompt Sentence Examples

[1114] "Track the person in this photo, verify their identity against police databases, and suggest appropriate actions based on the user's current emotions."

[1115] In this way, the present invention improves the efficiency and accuracy of crime prevention activities and enables the provision of information that takes into consideration the user's emotions.

[1116] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1117] Step 1:

[1118] A user uploads a photo of a specific individual to a smartphone. The input is photo data that clearly shows the face. The output is the photo data that is saved on the device.

[1119] Step 2:

[1120] The device recognizes the user's emotional state from video or image data and generates emotional data using an emotion engine. The input is the user's video or image data, which is analyzed and emotional data (e.g., stress level, tension, etc.) is output.

[1121] Step 3:

[1122] The terminal sends photo data and emotion data to the server. The input at this time is the photo data and emotion data. The output is the data received by the server.

[1123] Step 4:

[1124] The server performs face recognition using the received photo data. The software used is OpenCV. The input is the photo data, the face is detected, and data on specific feature points (eyes, nose, mouth) is output.

[1125] Step 5:

[1126] The server extracts clothing features from the photo data, including information on color, pattern, shape, etc. The input is the photo data, and the output is clothing feature data.

[1127] Step 6:

[1128] The server searches a security camera video database for matching video based on the facial feature data and clothing feature data. The input is the facial feature data and clothing feature data, and the output is the matching video data.

[1129] Step 7:

[1130] The server records the timestamp and camera location of the video in which the matching person appears. The input is the matching video data, and the output is the timestamp and camera location information.

[1131] Step 8:

[1132] The server analyzes the target person's behavioral patterns based on the timestamp and camera position information. For example, the server outputs the person's route and duration of stay. The input is the timestamp and camera position information.

[1133] Step 9:

[1134] The server compares the results of behavioral pattern analysis with the police database to confirm the identity of the person. The input is facial feature data, and the matching results from the police database are output.

[1135] Step 10:

[1136] The server integrates the analysis results with the user's emotional data and generates a final report. The inputs are the behavioral pattern analysis results and the emotional data, and the output is a report.

[1137] Step 11:

[1138] The terminal presents this generated report to the user. The input is the report sent from the server, and the output is to display it in a user-friendly format.

[1139] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1140] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1141] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1142] [Fourth embodiment]

[1143] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1144] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1145] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1146] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1147] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1148] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1149] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1150] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1151] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1152] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1153] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1154] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1155] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1156] The present invention is a system that extracts facial recognition and clothing characteristics from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and further verifies their identity in cooperation with a police database. Specific embodiments of the present invention are described in detail below.

[1157] System Overview

[1158] The system consists of the following main components:

[1159] 1. User terminal: The interface where the user inputs photo data and receives the analysis results.

[1160] 2. Server: Responsible for central processing such as data processing, facial recognition, behavioral analysis, and database matching.

[1161] 3. Security camera video database: A database that stores video data collected from multiple security cameras.

[1162] 4. Police databases: Databases containing personally identifiable information.

[1163] System operation and program processing

[1164] 1. Data Entry

[1165] The user uploads a photo of the individual to be tracked to the device. This photo data includes both the face and clothing. The device then sends the photo data to the server.

[1166] 2. Pretreatment

[1167] The server receives this photo data and first extracts the facial region using a facial recognition algorithm, from which it identifies the following feature points:

[1168] Eye, nose, and mouth position

[1169] Facial contours

[1170] Other features (mole, scar, etc.)

[1171] At the same time, the server extracts the clothing's characteristics, specifically analyzing its color, pattern, shape, etc., and generates feature data.

[1172] 3. Video search and feature data matching

[1173] The server accesses a security camera video database and searches for video data based on the identified feature data. From the security camera video, it performs facial recognition and clothing feature matching to identify matching individuals.

[1174] 4. Recording timestamps and camera positions

[1175] For each frame in which a matching person is identified, the server records the timestamp and camera position, allowing the path of the person to be traced.

[1176] 5. Behavioral Pattern Analysis

[1177] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, the following information is generated:

[1178] Travel route

[1179] Stay time

[1180] Frequently visited places

[1181] This information is organized as a result of behavioral pattern analysis.

[1182] 6. Identity Verification

[1183] After the behavioral pattern analysis is completed, the server will compare the target person's facial feature data with the police database, and based on the matching results, the target person's identity will be confirmed.

[1184] 7. Report generation and presentation

[1185] Finally, the server generates a detailed report containing the results of the analysis and identification, including the following information:

[1186] Travel route

[1187] Behavioral patterns

[1188] Timestamp and camera position

[1189] Identity verification results

[1190] The terminal presents this report to the user and provides the necessary information.

[1191] Specific examples

[1192] For example, to track a suspicious person in a public facility, a user uploads a photo of the suspicious person to the system. The server extracts facial and clothing features from the photo and accesses a security camera footage database to detect the person. The server then analyzes the person's behavioral patterns to identify their route of travel and where they are staying. Finally, the server checks the data against a police database to confirm their identity, and a detailed report is provided to the user. This allows for accurate and rapid tracking of people.

[1193] The processing flow will be explained below.

[1194] Step 1:

[1195] The user uploads a photo of the individual to be tracked to the device, which receives and stores the photo data.

[1196] Step 2:

[1197] The device sends the photo data to the server, which receives it and applies a facial recognition algorithm.

[1198] Step 3:

[1199] The server uses a facial recognition algorithm to extract facial regions from the photograph, identify specific feature points (such as the eyes, nose, and mouth) from the facial regions, and generate facial feature data.

[1200] Step 4:

[1201] At the same time, the server analyzes the clothing characteristics (color, pattern, shape, etc.) from the photograph data and generates clothing characteristic data.

[1202] Step 5:

[1203] The server accesses the security camera video database and obtains the latest video data.

[1204] Step 6:

[1205] The server performs facial recognition and clothing feature matching on the acquired video data to search for people who match specific feature data.

[1206] Step 7:

[1207] The server records the timestamp and camera position for each frame in which a matching person is identified.

[1208] Step 8:

[1209] The server uses multiple timestamps and camera location information to analyze the target person's behavioral characteristics, such as their route of travel and length of stay.

[1210] Step 9:

[1211] The server organizes the analysis results and extracts behavioral patterns (frequently visited places, places where people spend long periods of time, etc.).

[1212] Step 10:

[1213] The server uses the analyzed facial feature data to access police databases and verify the identity of the person in question.

[1214] Step 11:

[1215] The server adds the matching results to the report and combines all the analysis results to generate the final report.

[1216] Step 12:

[1217] The terminal presents the final report to the user, who then checks the report and obtains the necessary information.

[1218] Example 1

[1219] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1220] In modern society, numerous security cameras and surveillance devices are installed, helping to prevent and investigate crime. However, quickly and accurately identifying and tracking specific individuals from large amounts of video data is difficult, time-consuming, and laborious. To solve this problem, there is a need for a system that can extract facial recognition and clothing characteristics from photographic data, track and analyze the behavior of target individuals using a security camera video database, and further verify their identity in conjunction with an identification database.

[1221] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1222] In this invention, the server includes a face recognition means, a clothing feature extraction means, a means for accessing a video database, a means for searching video data based on the identified feature data, a means for recording timestamps and location information, a means for analyzing behavioral patterns, a means for verifying identity by comparing with an identification database, a means for generating and presenting analysis results as a report, a user interface including a means for uploading photo data, and a means for using a generative AI model as part of the analysis process. This makes it possible to quickly and accurately identify a target person from photo data, analyze the person's behavioral patterns, and verify the person's identity in cooperation with the identification database.

[1223] "Facial recognition means" refers to a device or program that uses technology or algorithms to detect facial feature points and identify individuals.

[1224] "Clothing feature extraction means" refers to a device or program that uses technology or algorithms to analyze and extract clothing features such as color, pattern, and shape from image data.

[1225] "Means for accessing a video database" refers to a device or program that connects to a database in which video data collected from multiple cameras is stored and acquires the necessary data.

[1226] "Means for searching video data based on identified feature data" refers to devices or programs that use technology or algorithms to search for matching people from videos in a video database using extracted facial and clothing feature data.

[1227] The "means for recording timestamp and location information" refers to a device or program that acquires and records the shooting date and time of the searched video frame and the installation location of the camera.

[1228] "Means for analyzing behavioral patterns" refers to devices or programs that use technologies or algorithms to analyze a target person's behavioral patterns, such as their travel routes, length of stay, and frequently visited places, based on recorded timestamps and location information.

[1229] "Means for verifying identity by matching against an identification database" refers to devices or programs that use technology or algorithms to compare analyzed facial feature information with an identification database and verify the identity of the person based on matching data.

[1230] The "means for generating and presenting the analysis results as a report" refers to a device or program that organizes the results of behavioral pattern analysis and identity verification, and generates a report that can be visually presented to the user.

[1231] The "user interface including a means for uploading photographic data" is an interface that allows a user to input photographic data into the system via a terminal.

[1232] "Means of using a generative AI model as part of the analytical process" refers to devices or programs that utilize generative AI models as part of the analysis, and use technologies or algorithms to achieve more advanced data processing and pattern recognition.

[1233] The present invention is a system that extracts facial recognition and clothing characteristics from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and verifies their identity in cooperation with an identification database. To implement this system, the following hardware and software are used.

[1234] System Overview

[1235] The system consists of the following main components:

[1236] 1. User terminal: An interface where the user inputs photo data and receives the analysis results.

[1237] 2. Server: Responsible for central processing such as data processing, facial recognition, behavior analysis, and database matching.

[1238] 3. Security camera video database: A database that stores video data collected from multiple security cameras.

[1239] 4. Identification Database: A database containing personally identifiable information.

[1240] Hardware and software used

[1241] Face recognition algorithm: Uses OpenCV and dlib libraries.

[1242] Clothing feature extraction algorithm: Uses color analysis using RGB values, texture analysis, and contour detection algorithms.

[1243] Video search algorithm: Facenet and DeepFace are used for face recognition, and Elasticsearch's image search plugin is used for clothing matching.

[1244] Behavioral analysis algorithms: Use custom algorithms to plot movement paths and analyze dwell times.

[1245] Identity database matching algorithm: Uses face identification model to match with identity database.

[1246] Detailed procedure

[1247] The user uploads photo data of the object to be tracked to the device, which then sends the uploaded photo data to the server. The server then uses a facial recognition algorithm to detect the facial area and extract feature points such as the eyes, nose, mouth, facial contours, moles, and scars.

[1248] At the same time, the server analyzes the color, pattern, and shape of the clothing to generate feature data. Based on the analyzed feature data, it accesses a security camera video database to search for matching video data. The server then obtains and records the timestamp and camera position of the video frame of the matching person.

[1249] The system then analyzes the person's behavioral patterns based on timestamps and camera location data to identify their route and location. Finally, the system compares the facial feature data with an identification database to confirm their identity. The server then processes the analysis results and generates a detailed report that is presented to the user's device.

[1250] Specific examples

[1251] For example, when tracking a suspicious person in a public facility, the user uploads photo data to the system. The server extracts facial and clothing features, searches a security camera footage database, and analyzes the suspicious person's behavioral patterns. The server then verifies the person's identity against an identification database and provides a detailed report to the user.

[1252] Prompt Sentence Examples

[1253] "Using this photo, please briefly explain the entire process of identifying the person in question by extracting facial and clothing characteristics, searching a security camera footage database, and analyzing their behavioral patterns to confirm their identity."

[1254] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1255] Step 1: Data entry

[1256] The user uploads photographic data of the individual to be tracked to the device, which includes information on their face and clothing.

[1257] Input: Photo data (face and clothing information)

[1258] Output: Photo data sent to the server

[1259] Specifically, the user selects photo data using the file upload function of the device and clicks the upload button. The device then sends the selected photo data to the server as an HTTP request.

[1260] Step 2: Face recognition and clothing feature extraction (preprocessing)

[1261] The server receives this photo data and extracts the face region using a facial recognition algorithm, for example, using the face detection function of OpenCV or dlib.

[1262] Input: Submitted photo data

[1263] Output: Facial feature points (eyes, nose, mouth position, facial contours, moles, scars, etc.)

[1264] Specifically, the server reads the photo data using an image processing library, detects the facial area, and extracts feature points.

[1265] At the same time, the server extracts the clothing's features by analyzing its color (RGB value), pattern (texture analysis), and shape (contour detection) to generate feature data.

[1266] Input: Submitted photo data

[1267] Output: Clothing feature data (color, pattern, shape)

[1268] Specifically, the server extracts clothing feature data using a color analysis algorithm, a texture analysis algorithm, and a shape analysis algorithm.

[1269] Step 3: Feature data matching and video search

[1270] The server then accesses a security camera video database based on the identified facial and clothing feature data and searches for the relevant video data, using Facenet and DeepFace for facial recognition and Elasticsearch's image search plugin for clothing matching.

[1271] Input: facial feature points, clothing feature data

[1272] Output: Video data of the matching person

[1273] Specifically, the server queries a security camera video database to search for video frames that match the facial features and clothing data.

[1274] Step 4: Record timestamps and camera positions

[1275] For each frame in which a matching person is identified, the server captures and records the timestamp and camera position from the video data.

[1276] Input: Video data of the person to match

[1277] Output: timestamp and camera position information

[1278] Specifically, the server extracts and records the timestamp and camera position from the metadata of the matched video frame.

[1279] Step 5: Behavioral pattern analysis

[1280] The server analyzes the target person's behavioral patterns based on the recorded timestamps and camera location information, plotting their movement paths, analyzing their stay times, and identifying frequently visited locations.

[1281] Input: timestamp, camera location information

[1282] Output: Behavioral pattern analysis results (route, stay time, frequently visited places)

[1283] Specifically, the server analyzes the timestamp and camera position data, plots the target person's movement path, and calculates the length of time they stayed there.

[1284] Step 6: Identity Verification

[1285] After completing the behavioral pattern analysis, the server compares the target person's facial feature data with the identification database using a facial identification model (such as Facenet or DeepFace).

[1286] Input: Facial feature data

[1287] Output: Identity verification result

[1288] Specifically, the server accesses an identification database, matches the facial feature data, and confirms the identity of the matching individual.

[1289] Step 7: Report generation and presentation

[1290] Finally, the server generates a detailed report containing the analysis results and the identification results, including the route of travel, behavioral patterns, timestamps, camera locations, and the identification results. The device then presents this report to the user.

[1291] Input: behavioral pattern analysis results, identity verification results

[1292] Output: Detailed report

[1293] Specifically, the server compiles the movement route, behavioral patterns, timestamps, camera locations, and identity verification results to generate a report, which is then displayed on the device for the user to view.

[1294] Prompt Sentence Examples

[1295] "Using this photo, please briefly explain the entire process of identifying the person in question by extracting facial and clothing characteristics, searching a security camera footage database, and analyzing their behavioral patterns to confirm their identity."

[1296] (Application example 1)

[1297] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1298] Existing systems that use security camera footage and police databases to track and identify suspicious individuals have difficulty informing administrators of suspicious individuals' movements in real time. Furthermore, delays in reporting the results of the analysis mean that appropriate countermeasures cannot be implemented promptly.

[1299] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1300] In this invention, the server includes a facial recognition unit, a clothing feature extraction unit, a unit for accessing a security camera video database, a unit for searching for video based on the identified feature data, a unit for recording timestamps and camera locations, a unit for analyzing behavioral patterns, a unit for verifying identity by checking against a police database, a unit for generating and presenting a report of the analysis results, and a unit for sending data to a smartphone application to notify the administrator. This makes it possible to notify the location and behavioral patterns of suspicious individuals in real time.

[1301] "Facial recognition means" is a technology that extracts facial feature points from an input facial image and identifies an individual using a specific algorithm.

[1302] "Clothing feature extraction means" is a technology that analyzes features such as clothing color, shape, and pattern from an input image and extracts them as data.

[1303] "Means for accessing a security camera video database" refers to technology that accesses a database that stores video data collected from multiple security cameras, and performs searches and retrieval.

[1304] "Means for searching for footage based on identified feature data" refers to a technology that searches security camera footage in a database based on extracted facial and clothing feature data, and identifies the person in question.

[1305] The "means for recording timestamps and camera positions" is a technology for recording the time information of a specified video frame and the location where the camera is installed.

[1306] "Means for analyzing behavioral patterns" refers to technology that analyzes behavioral patterns such as a person's movement route and length of stay based on timestamps and camera position information.

[1307] "Means of confirming identity by comparing with police database" refers to a technology that compares extracted facial feature data with a database held by the police to confirm the identity of the target person.

[1308] "Means for generating and presenting analysis results as a report" refers to a technology that creates a detailed report based on the results of behavioral pattern analysis and identity verification and provides it to the user.

[1309] "Means for sending data to a smartphone application and notifying the user" refers to a technology that sends generated reports and analysis results to a smartphone application in real time and notifies the user.

[1310] The present invention is a system including a facial recognition unit, a clothing feature extraction unit, a security camera video database access unit, a video search unit based on the identified feature data, a timestamp and camera location recording unit, a behavioral pattern analysis unit, a police database comparison unit to confirm identity, a report generating and presenting the analysis results, and a smartphone application that notifies the user of the analysis results. This system is suitable for quickly and accurately tracking and identifying suspicious individuals using security camera video and the police database.

[1311] Face recognition and clothing feature extraction

[1312] The server receives the photo data uploaded by the user to the device and extracts the face area using a face recognition algorithm (e.g., OpenCV or Google Vision API), identifies facial feature points (positions of eyes, nose, mouth, facial contours, and other features), and generates data on clothing features (color, pattern, shape, etc.).

[1313] Video database search and matching

[1314] The server accesses a security camera video database and searches for video data based on the extracted feature data. It then performs facial recognition and clothing feature matching on the security camera footage to identify matching individuals. For each identified frame, the server records the timestamp and camera location, and tracks the target individual's movement path.

[1315] Behavioral pattern analysis and identity verification

[1316] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, it generates information such as the person's route of travel, length of stay, and frequently visited places. The server then organizes the analysis results. The server then compares the target person's facial feature data with the police database to confirm their identity.

[1317] Report generation and notification

[1318] Finally, the server generates a detailed report containing the analysis and identity verification results and sends a notification to the smartphone application, where the user can view the report and take any necessary measures.

[1319] For example, if a public facility manager spots someone behaving suspiciously in a parking lot and uploads a photo of them to a smartphone application, the system will review security camera footage and analyze the suspicious person's movement path. It will then compare the person's identity with a police database and confirm their identity. The manager will receive a real-time notification, enabling them to respond quickly.

[1320] Prompt Sentence Examples

[1321] "Upload a photo of someone behaving suspiciously in your parking lot. We will then extract the suspicious person's facial and clothing characteristics from the photo and search our security camera footage database for matching footage. We will then confirm the person's behavioral patterns and identity, and generate a detailed report to notify you."

[1322] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1323] Step 1:

[1324] A user uploads a photo of a suspicious person using a smartphone application. The input is the photo data of the suspicious person, and the output is the photo data sent to the server.

[1325] Step 2:

[1326] The server analyzes the received photo data and extracts the face region using a face recognition algorithm (e.g., OpenCV or Google Vision API). The input is the photo data, and the output is facial feature points (positions of eyes, nose, mouth, facial contours, and other features).

[1327] Step 3:

[1328] At the same time, the server extracts clothing features from the photo data. Specifically, it analyzes color, pattern, and shape and generates feature data. The input is the photo data, and the output is clothing feature data.

[1329] Step 4:

[1330] The server accesses the security camera video database and searches the video data based on the extracted facial and clothing feature data. The input is the facial and clothing feature data, and the output is the video frame of the matching person.

[1331] Step 5:

[1332] The server records the timestamp and camera position for each frame in which a matching person is identified. The input is the matching video frame, and the output is a record of the time and camera position information.

[1333] Step 6:

[1334] The server analyzes the target person's behavioral patterns based on timestamps and camera location information. The input is time information and camera location information, and the output is behavioral pattern information such as travel route, length of stay, and frequently visited places.

[1335] Step 7:

[1336] The server compares the facial feature data with the police database to confirm the identity of the person. The input is the facial feature data, and the output is the identity confirmation result.

[1337] Step 8:

[1338] The server generates a detailed report including the analysis results and identity verification results, and notifies the smartphone application of the report. The input is behavioral pattern information and identity verification results, and the output is a report sent to the user's device.

[1339] Step 9:

[1340] The user checks the received report through a smartphone application and takes necessary measures. The input is the report from the server, and the output is the countermeasure action taken by the user.

[1341] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1342] The present invention combines a system that extracts facial recognition and clothing features from photographic data, tracks and analyzes the behavior of target individuals using a security camera video database, and verifies their identity in cooperation with a police database with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention will be described in detail below.

[1343] System Overview

[1344] The system consists of the following main components:

[1345] 1. User Device:

[1346] An interface where users can input photo data and receive analysis results.

[1347] An emotion engine that recognizes user emotions and collects emotion data.

[1348] 2. Server:

[1349] It handles core processing such as data processing, facial recognition, behavioral analysis, and database matching.

[1350] Emotional data from the emotion engine is used for analysis to optimize the way information is presented to users.

[1351] 3. Security camera footage database:

[1352] A database that stores video data collected from multiple security cameras.

[1353] 4. Police Database:

[1354] Databases containing personally identifiable information.

[1355] System operation and program processing

[1356] 1. Data Entry

[1357] The user uploads a photo of the individual to be tracked to the device. This photo data includes both the face and clothing. The device sends the photo data to the server and recognizes the user's emotions to generate emotion data. This emotion data is also sent to the server.

[1358] 2. Pretreatment

[1359] The server receives the photo data and first extracts the facial region using a facial recognition algorithm. From this facial region, it identifies specific feature points such as the eyes, nose, and mouth, and generates facial feature data. At the same time, the server analyzes the clothing features (color, pattern, shape, etc.) from the photo data and generates clothing feature data.

[1360] 3. Video search and feature data matching

[1361] The server accesses a security camera video database and searches for video data based on the identified feature data. From the security camera video, it performs facial recognition and clothing feature matching to identify matching individuals.

[1362] 4. Recording timestamps and camera positions

[1363] For each frame in which a matching person is identified, the server records the timestamp and camera position, allowing the path of the person to be traced.

[1364] 5. Behavioral Pattern Analysis

[1365] The server analyzes the target person's behavioral patterns based on the timestamp and camera location information. For example, it generates information such as the person's route, length of stay, and frequently visited places. This analysis also takes into account the user's emotional data.

[1366] 6. Identity Verification

[1367] After the behavioral pattern analysis is completed, the server will compare the target person's facial feature data with the police database, and based on the matching results, the target person's identity will be confirmed.

[1368] 7. Report generation and presentation

[1369] Finally, the server generates a detailed report containing the analysis results, emotion data, and identity verification results. This report includes behavioral patterns, timestamps, camera locations, identity verification results, and information presentation methods adjusted according to the emotion data. The device then presents this report to the user and provides the necessary information.

[1370] Specific examples

[1371] For example, to track a suspicious person in a public facility, a user uploads a photo of the suspicious person to the system. The device receives the photo data and sends it to the server, recognizing the user's emotions to generate emotion data. The server extracts facial and clothing features from the photo and accesses a security camera footage database to detect the person. The server then analyzes the person's behavioral patterns and identifies their route and where they stayed. Finally, it compares the person with a police database to confirm their identity, and a detailed report is provided to the user. This report is based on the user's emotion data and presented in an easy-to-understand format. This enables accurate and fast person tracking.

[1372] In this way, the present invention improves the efficiency and accuracy of crime prevention activities and enables the provision of information that takes into consideration the user's emotions.

[1373] The processing flow will be explained below.

[1374] Step 1:

[1375] The user uploads a photo of the individual to be tracked to the device, which receives the photo data and sends it to the server.

[1376] Step 2:

[1377] The device activates an emotion engine to recognize the user's emotions, collecting emotion data from the user's voice, facial expressions, input data, etc.

[1378] Step 3:

[1379] The device sends the collected emotion data to the server, which receives the photo data and emotion data and begins processing.

[1380] Step 4:

[1381] The server uses a facial recognition algorithm to extract facial regions from the photograph, and then identifies feature points such as the eyes, nose, and mouth from the facial regions to generate facial feature data.

[1382] Step 5:

[1383] The server analyzes the clothing characteristics (color, pattern, shape, etc.) from the photograph data and generates clothing characteristic data.

[1384] Step 6:

[1385] The server accesses the security camera video database and searches the video data based on the acquired feature data. It then performs facial recognition and clothing feature matching on the security camera video to identify matching individuals.

[1386] Step 7:

[1387] The server records the timestamp and camera position for each frame in which a matching person is identified.

[1388] Step 8:

[1389] The server uses the recorded timestamps and camera positions to analyze the target person's movement route, duration of stay, frequently visited places, etc.

[1390] Step 9:

[1391] The server organizes the analysis results and extracts behavioral patterns, summarizing the information in a format that is easy for users to understand.

[1392] Step 10:

[1393] The server uses the facial feature data of the target person to compare it with the police database to confirm their identity, and obtains the matching results.

[1394] Step 11:

[1395] The server integrates the results of behavioral analysis and identity verification into a report, adjusts the report content based on the user's emotional data, and optimizes the information presentation method as needed.

[1396] Step 12:

[1397] The terminal presents the final report to the user, who then checks the report and obtains the necessary information.

[1398] Example 2

[1399] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1400] Conventional security systems are capable of tracking individuals and analyzing their behavior, but they do not provide information that takes into account the user's emotional state, and they are unable to present information in an efficient and easy-to-understand manner. Furthermore, they do not adequately utilize user emotional data to improve the accuracy of behavioral pattern analysis. As a result, tracking results and analytical information are sometimes not practical for users, and there is a need to improve the efficiency and accuracy of security activities.

[1401] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1402] In this invention, the server includes a facial recognition means, a clothing feature extraction means, a means for accessing a security camera video database, a means for searching for video based on the identified feature data, a means for recording a timestamp and camera location, a means for analyzing behavioral patterns, a means for verifying identity by checking against a police database, a means for generating and presenting a report of the analysis results, a means for recognizing a user's emotions using an emotion recognition engine and generating emotion data, and a means for optimizing information presentation by utilizing the generated emotion data in analysis, thereby enabling improved efficiency and accuracy in crime prevention activities.

[1403] The "face recognition means" is a means for detecting an individual's face area from image data, identifying feature points (eyes, nose, mouth, etc.), and extracting facial feature data.

[1404] The "clothing feature extraction means" is a means for analyzing clothing features (color, pattern, shape, etc.) from image data and generating clothing feature data.

[1405] "Means for accessing a security camera video database" refers to a means for connecting to a database that stores video data collected from a large number of security cameras and obtaining the necessary video data from there.

[1406] The "means for searching for video based on identified feature data" refers to a means for searching for video in a security camera video database based on the extracted facial feature data and clothing feature data, and identifying matching individuals.

[1407] The "means for recording timestamps and camera locations" refers to a means for recording the time of capture and the camera location information for each video frame in which a matching person is identified.

[1408] "Means for analyzing behavioral patterns" refers to a method for analyzing a target person's movement route, length of stay, frequently visited places, etc. from recorded timestamps and camera location information to analyze behavioral patterns.

[1409] "Means for confirming identity by checking against a police database" refers to a means for comparing facial feature data after behavioral pattern analysis with a police database and confirming the identity of the target person based on the matching results.

[1410] The "means for generating and presenting analysis results as a report" refers to a means for compiling behavioral analysis results, analysis results based on emotional data, and comparison results with police databases to create a detailed report and provide it to the user.

[1411] The "means for recognizing a user's emotions using an emotion recognition engine and generating emotion data" refers to a means for taking a picture of the user's face, analyzing their emotional state through an emotion recognition algorithm, and generating the results as data.

[1412] "Means for optimizing information presentation by utilizing generated emotional data in analysis" refers to means for adjusting the analysis results and information presentation method based on the acquired emotional data, and providing information in a form that is most easily understandable for the user.

[1413] System Overview

[1414] This system extracts facial recognition and clothing characteristics from personal photo data, and tracks and analyzes the behavior of target individuals using a security camera video database. Furthermore, by linking with a police database to verify identity and combining it with an emotion engine that recognizes user emotions, it achieves efficient and highly accurate crime prevention activities.

[1415] Key Components

[1416] User device:

[1417] It provides an interface where users can input photo data and receive analysis results.

[1418] An emotion recognition engine is used to collect user emotion data and send it to a server.

[1419] server:

[1420] It handles core processing such as data processing, facial recognition, behavioral analysis, and database matching.

[1421] Emotional data from the emotion engine is used for analysis to optimize the way information is presented to users.

[1422] Security camera footage database:

[1423] It is a database that stores video data collected from multiple security cameras.

[1424] Police Database:

[1425] It is a database containing personally identifiable information.

[1426] Specific actions

[1427] 1. Data Entry and Emotion Recognition

[1428] The user uploads a photo of the individual to be tracked to the device. The device receives the photo data and temporarily stores it in local storage. The device's built-in emotion engine then captures the user's face with a camera and generates emotion data in real time using an emotion recognition algorithm, such as TensorFlow. This emotion data is also sent to the server.

[1429] 2. Preprocessing of photo data

[1430] The server preprocesses the received photo data using the OpenCV library. Specifically, it converts the image to grayscale and performs face detection. It then uses the Haar Cascade and Dlib libraries to detect the face area and identify feature points such as the eyes, nose, and mouth to generate facial feature data. At the same time, it analyzes clothing features (color, pattern, shape, etc.) from the photo data and generates clothing feature data. Clothing analysis libraries such as DeepFashion are used for this.

[1431] 3. Video Search and Database Matching

[1432] The server accesses the security camera video database and utilizes SQL queries and NoSQL database technologies (e.g., MongoDB). It searches the video data based on facial and clothing feature data. It uses FaceNet or DeepFace models for facial recognition and similarity search algorithms (e.g., Cosine Similarity) for clothing matching. The server identifies matching individuals from the search results and extracts their video frames.

[1433] 4. Recording timestamps and camera positions

[1434] The server records the timestamp and camera location information for each video frame of the identified person, and stores the person's movement path in a detailed database.The server efficiently indexes the timestamp and location information using Elasticsearch.

[1435] 5. Behavioral Pattern Analysis

[1436] The server uses the collected timestamps and location information to analyze the target person's behavioral patterns using machine learning models (e.g., time series analysis methods and LSTM). It generates detailed information such as the person's route, length of stay, and frequently visited places. The server also incorporates the user's emotional data into the analysis and adjusts the analysis results.

[1437] 6. Identity Verification

[1438] After completing the behavioral pattern analysis, the server compares the facial feature data of the target person with the police database, using REST API or SOAP to access the police database and confirm the target person's identity based on the comparison results.

[1439] 7. Report generation and presentation

[1440] The server compiles the analysis results, emotion data, and identity verification results to generate a detailed report, possibly using Jupyter Notebook or ReportLab. The device receives the generated report and presents it to the user in the form of a dashboard or PDF report.

[1441] Specific examples

[1442] For example, to track a suspicious individual in a public facility, a user uploads a photo of the suspicious individual to the system. The device receives the photo data and sends it to the server, generating the user's emotional data using an emotion engine. The server extracts facial and clothing features from the uploaded photo and searches a security camera footage database to detect the individual. The server then analyzes the identified individual's behavioral patterns, determines their route of travel and where they stayed, and finally verifies their identity by comparing it with a police database. A detailed report is generated and provided to the user via the device. This report is presented in an easy-to-read format, taking into account the user's emotional data.

[1443] Example prompts for generative AI models

[1444] "We are developing a system to track suspicious individuals in public facilities. Please tell us how to build a program that inputs photographic data and searches a security camera footage database using the facial and clothing data obtained. Also, please explain in detail the procedure for analyzing the individual's behavioral patterns obtained in this way and comparing them with the police database to confirm their identity."

[1445] In this way, the system improves the efficiency and accuracy of crime prevention activities while providing information that takes into consideration the user's emotions.

[1446] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1447] Step 1:

[1448] The user uploads a photo of the individual to be tracked to the device. The device receives this photo data and temporarily stores it in local storage. Next, the device's built-in emotion recognition engine captures the user's face with a camera and generates emotion data using an emotion recognition algorithm. TensorFlow is used for this. This emotion data is sent to the server along with the photo data. The input is the individual's photo and the user's facial image, and the output is the photo data and emotion data.

[1449] Step 2:

[1450] The server preprocesses the received photo data using the OpenCV library. Specifically, it converts the image to grayscale and performs face detection. Here, the Haar Cascade and Dlib libraries are used. The face area is detected and feature points such as the eyes, nose, and mouth are identified to generate facial feature data. At the same time, the server uses a clothing analysis library such as DeepFashion to analyze clothing features (color, pattern, shape, etc.) from the photo data and generate clothing feature data. The input is photo data, and the output is facial feature data and clothing feature data.

[1451] Step 3:

[1452] The server accesses the security camera video database and utilizes SQL queries and NoSQL database technology (e.g., MongoDB). It searches security camera footage based on facial feature data and clothing feature data. It uses FaceNet or DeepFace models for facial recognition and a similarity search algorithm (e.g., Cosine Similarity) for clothing matching. It identifies matching people from the search results and extracts their video frames. The input is facial feature data and clothing feature data, and the output is video frames of matching people.

[1453] Step 4:

[1454] The server records the timestamp and camera location information for each video frame of the identified person and stores the person's movement path in a database. Here, we use Elasticsearch to index the timestamp and location information. The input is the video frame of the matching person, and the output is a record of the timestamp and camera location information.

[1455] Step 5:

[1456] The server analyzes the target person's behavioral patterns based on the collected timestamps and location information. This uses machine learning models (e.g., time series analysis methods and LSTM). The behavioral pattern analysis includes information such as travel routes, length of stay, and frequently visited places. In addition, the analysis results are adjusted taking into account the user's emotional data. The inputs are timestamps, camera location information, and emotional data, and the output is the analyzed behavioral patterns.

[1457] Step 6:

[1458] After completing the behavioral pattern analysis, the server compares the facial feature data of the target person with the police database. It accesses the police database using REST API or SOAP. Based on the comparison results, the target person's identity is confirmed. The input is facial feature data, and the output is the identity confirmation result.

[1459] Step 7:

[1460] The server compiles the analysis results, emotion data, and identity verification results to generate a detailed report. Jupyter Notebook or ReportLab can be used for this. The generated report is then presented to the user via their terminal. The information is presented in the form of a dashboard or PDF report. The input is the analysis results, emotion data, and identity verification results, and the output is a detailed report.

[1461] (Application example 2)

[1462] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1463] Conventional security systems were able to track individuals by facial recognition and extracting clothing characteristics, but they were unable to take the user's emotional state into account, making it difficult for them to provide optimal information to users in stressful situations. There was also a need for a system that could track individuals in real time and respond appropriately based on the user's emotions.

[1464] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a clothing feature extraction means, a means for accessing a security camera video database, a means for recording timestamps and camera locations, a means for analyzing behavioral patterns, a means for verifying identity by checking against a police database, a means for incorporating an emotion engine into the user terminal, and a means for utilizing the user's emotion data for analysis and presenting optimized information. This enables accurate tracking and analysis of individuals as well as the provision of information that takes the user's emotions into consideration.

[1465] "Facial recognition means" refers to devices and technology that detect an individual's face from input photographic data, analyze its features, and digitize them.

[1466] "Clothing feature extraction means" refers to a device and technology for analyzing the features of an individual's clothing, such as color, pattern, and shape, from input photographic data and converting them into data.

[1467] "Means for accessing a security camera video database" refers to devices and technologies for accessing video data collected from security cameras and obtaining the necessary information.

[1468] "Means for searching for video based on identified feature data" refers to devices and technologies for searching for matching video within a security camera video database based on feature data obtained through facial recognition or clothing feature extraction.

[1469] "Means for recording timestamps and camera locations" refers to devices and techniques for recording the date and time (timestamp) when the image was recorded and the location of the camera for each frame in which an identified person appears.

[1470] "Means for analyzing behavioral patterns" refers to devices and technology for analyzing behavioral patterns, such as the route a target person took and how long they stayed in a particular location, based on timestamps and camera position data.

[1471] "Means for verifying identity by comparing with police databases" refers to devices and technologies for verifying the identity of a person by comparing the acquired facial feature data with personal identification information held by the police.

[1472] The "means for generating and presenting the analysis results as a report" refers to a device and technology for integrating the above analysis results and generating and presenting them as a report in a format that is easy for the user to understand.

[1473] "Means for incorporating an emotion engine into a user terminal" refers to devices and techniques for incorporating emotion recognition software and hardware into a user terminal.

[1474] "Means for utilizing user emotional data for analysis and presenting optimized information" refers to devices and technologies that analyze the user's emotional state, adjust the way information is presented based on the results, and provide data in a form that is more suitable for the user.

[1475] The present invention is a system that includes a facial recognition unit, a clothing feature extraction unit, a security camera video database access unit, a video search unit based on identified feature data, a timestamp and camera location recorder, a behavioral pattern analysis unit, a police database for identifying the user, a report generating and presenting the analysis results, an emotion engine embedded in the user's device, and a unit for analyzing and presenting optimized information based on the user's emotion data. This system enables accurate tracking of individuals and the provision of information based on the user's emotion.

[1476] Hardware and software used

[1477] Hardware: Smartphones, smart glasses

[1478] Software: OpenCV (image processing library), EmotionRecognition (emotion analysis engine), police database API client

[1479] Data processing and calculation flow

[1480] A user uploads a photo of a specific individual to the device and simultaneously captures the user's emotional state. When the device transmits the photo data to the server, it uses an emotion engine to generate the user's emotional data, which is also transmitted to the server.

[1481] The server uses OpenCV to perform facial recognition based on the received photo data. It extracts the facial area from the photo and digitizes its feature points (the positions of the eyes, nose, mouth, etc.). At the same time, it extracts clothing features from the photo data and generates clothing feature data. This data is used to search the security camera video database and identify footage of matching people.

[1482] The system records the timestamp and camera position for each frame in which the identified person appears, thereby clarifying the person's movement path. The server then performs behavioral analysis to analyze the person's movement patterns and the length of time they stayed in the area. Finally, the results of this analysis are compared with a police database to confirm the person's identity.

[1483] The analysis results are integrated with the user's emotional data to generate a final report, which is presented in an easy-to-understand format for the user, enabling appropriate responses that take into account the user's emotional state.

[1484] Specific example explanation

[1485] For example, if a suspicious person appears at a concert venue, a security officer can upload a photo of the person using their smartphone to the system. The system then sends the photo to a server, which performs facial recognition and clothing feature extraction, then searches the security camera footage database to analyze the person's behavioral patterns.

[1486] If a security officer experiences high stress levels, the system can provide additional support in real time to allow for a rapid response. This process can be facilitated by prompts such as:

[1487] Prompt Sentence Examples

[1488] "Track the person in this photo, verify their identity against police databases, and suggest appropriate actions based on the user's current emotions."

[1489] In this way, the present invention improves the efficiency and accuracy of crime prevention activities and enables the provision of information that takes into consideration the user's emotions.

[1490] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1491] Step 1:

[1492] A user uploads a photo of a specific individual to a smartphone. The input is photo data that clearly shows the face. The output is the photo data that is saved on the device.

[1493] Step 2:

[1494] The device recognizes the user's emotional state from video or image data and generates emotional data using an emotion engine. The input is the user's video or image data, which is analyzed and emotional data (e.g., stress level, tension, etc.) is output.

[1495] Step 3:

[1496] The terminal sends photo data and emotion data to the server. The input at this time is the photo data and emotion data. The output is the data received by the server.

[1497] Step 4:

[1498] The server performs face recognition using the received photo data. The software used is OpenCV. The input is the photo data, the face is detected, and data on specific feature points (eyes, nose, mouth) is output.

[1499] Step 5:

[1500] The server extracts clothing features from the photo data, including information on color, pattern, shape, etc. The input is the photo data, and the output is clothing feature data.

[1501] Step 6:

[1502] The server searches a security camera video database for matching video based on the facial feature data and clothing feature data. The input is the facial feature data and clothing feature data, and the output is the matching video data.

[1503] Step 7:

[1504] The server records the timestamp and camera location of the video in which the matching person appears. The input is the matching video data, and the output is the timestamp and camera location information.

[1505] Step 8:

[1506] The server analyzes the target person's behavioral patterns based on the timestamp and camera position information. For example, the server outputs the person's route and duration of stay. The input is the timestamp and camera position information.

[1507] Step 9:

[1508] The server compares the results of behavioral pattern analysis with the police database to confirm the identity of the person. The input is facial feature data, and the matching results from the police database are output.

[1509] Step 10:

[1510] The server integrates the analysis results with the user's emotional data and generates a final report. The inputs are the behavioral pattern analysis results and the emotional data, and the output is a report.

[1511] Step 11:

[1512] The terminal presents this generated report to the user. The input is the report sent from the server, and the output is to display it in a user-friendly format.

[1513] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1514] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1515] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1516] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1517] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1518] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1519] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1520] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1521] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1522] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1523] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1524] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1525] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1526] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1527] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1528] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1529] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1530] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1531] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1532] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1533] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1534] The following is further disclosed regarding the above embodiment.

[1535] (Claim 1)

[1536] A facial recognition means;

[1537] clothing feature extraction means;

[1538] a means for accessing a security camera footage database;

[1539] A means for searching for a video based on the identified feature data;

[1540] a means for recording a timestamp and camera position;

[1541] a means for analyzing patterns of behavior;

[1542] a means of verifying identity against police databases;

[1543] A means for generating and presenting the analysis results as a report;

[1544] A system including:

[1545] (Claim 2)

[1546] 10. The system of claim 1, further comprising means for obtaining the latest video data from a security camera video database.

[1547] (Claim 3)

[1548] 10. The system of claim 1, further comprising means for performing continuous frame analysis on the acquired video data to track a path of movement of a matching person.

[1549] "Example 1"

[1550] (Claim 1)

[1551] A facial recognition means;

[1552] clothing feature extraction means;

[1553] a means for accessing the video database;

[1554] A means for searching video data based on the identified feature data;

[1555] means for recording timestamps and location information;

[1556] a means for analyzing patterns of behavior;

[1557] a means of verifying identity against an identification database;

[1558] A means for generating and presenting the analysis results as a report;

[1559] a user interface including means for uploading photographic data;

[1560] A means of using generative AI models as part of the analysis process;

[1561] A system including:

[1562] (Claim 2)

[1563] 10. The system of claim 1, further comprising means for obtaining the latest video data from the video database.

[1564] (Claim 3)

[1565] 10. The system of claim 1, further comprising means for performing continuous frame analysis on the acquired video data to track a path of movement of a matching person.

[1566] "Application Example 1"

[1567] (Claim 1)

[1568] A facial recognition means;

[1569] clothing feature extraction means;

[1570] a means for accessing a security camera footage database;

[1571] A means for searching for a video based on the identified feature data;

[1572] a means for recording a timestamp and camera position;

[1573] a means for analyzing patterns of behavior;

[1574] a means of verifying identity against police databases;

[1575] A means for generating and presenting the analysis results as a report;

[1576] A means for transmitting and notifying data to a smartphone application;

[1577] A system including:

[1578] (Claim 2)

[1579] 10. The system of claim 1, further comprising means for obtaining the latest video data from a security camera video database.

[1580] (Claim 3)

[1581] 10. The system of claim 1, further comprising means for performing continuous frame analysis on the acquired video data to track a path of movement of a matching person.

[1582] "Example 2: Combining Emotion Engines"

[1583] (Claim 1)

[1584] A facial recognition means;

[1585] clothing feature extraction means;

[1586] a means for accessing a security camera footage database;

[1587] A means for searching for a video based on the identified feature data;

[1588] a means for recording a timestamp and camera position;

[1589] a means for analyzing patterns of behavior;

[1590] a means of verifying identity against police databases;

[1591] A means for generating and presenting the analysis results as a report;

[1592] means for recognizing a user's emotion using an emotion recognition engine and generating emotion data;

[1593] A means of optimizing information presentation by utilizing the generated emotion data for analysis;

[1594] A system including:

[1595] (Claim 2)

[1596] 10. The system of claim 1, further comprising means for obtaining the latest video data from a security camera video database.

[1597] (Claim 3)

[1598] 10. The system of claim 1, further comprising means for performing continuous frame analysis on the acquired video data to track a path of movement of a matching person.

[1599] "Application example 2 when combining emotion engines"

[1600] (Claim 1)

[1601] A facial recognition means;

[1602] clothing feature extraction means;

[1603] a means for accessing a security camera footage database;

[1604] A means for searching for a video based on the identified feature data;

[1605] a means for recording a timestamp and camera position;

[1606] a means for analyzing patterns of behavior;

[1607] a means of verifying identity against police databases;

[1608] A means for generating and presenting the analysis results as a report;

[1609] A means for incorporating an emotion engine into a user terminal;

[1610] A means of analyzing user emotion data and presenting optimized information;

[1611] A system including:

[1612] (Claim 2)

[1613] 10. The system of claim 1, further comprising means for obtaining the latest video data from a security camera video database.

[1614] (Claim 3)

[1615] 10. The system of claim 1, further comprising means for performing continuous frame analysis on the acquired video data to track a path of movement of a matching person. [Explanation of symbols]

[1616] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A facial recognition means; clothing feature extraction means; a means for accessing a security camera footage database; A means for searching for a video based on the identified feature data; a means for recording a timestamp and camera position; a means for analyzing patterns of behavior; a means of verifying identity against police databases; A means for generating and presenting the analysis results as a report; A system including:

2. 10. The system of claim 1, further comprising means for obtaining the latest video data from a security camera video database.

3. The system of claim 1 , further comprising means for performing a continuous frame analysis on the acquired video data to track a movement path of a matching person.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A