System

The security system accurately distinguishes between family and suspicious individuals using facial and voice recognition, simplifying setup and operation, and enhancing crime prevention efficacy.

JP2026028175APending Publication Date: 2026-02-19SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024130473
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing crime prevention systems are cumbersome to set up and operate, especially for the elderly, and lack accuracy in distinguishing between family members and suspicious individuals, leading to false alarms and inadequate response to genuine threats.

Method used

A security system that registers family members' faces, body shapes, and movements, analyzes video and audio data, and issues alarms or notifications when a suspicious individual is identified, incorporating facial and voice recognition to enhance accuracy.

Benefits of technology

The system provides highly accurate, automatic crime prevention measures that are easy to use, reducing false alarms and enabling quick response to suspicious activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028175000001_ABST
    Figure 2026028175000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: This system includes a means for registering the face, body shape and action of a family, a means for analyzing video and voice data collected by a camera and a microphone, a means for collating the analyzed data with the features of the registered family, a means for emitting warning sound or light when a suspicious person is specified, and a means for summarizing the features of the suspicious person and reporting them to a user and a security company.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, the number of burglaries and thefts has increased, making crime prevention measures increasingly important. However, existing crime prevention systems are time-consuming to set up and operate, making them difficult to operate, especially for the elderly, and limiting their effectiveness. Furthermore, warnings using light and sound alone are insufficient, and there is a need for more reliable methods of identifying and reporting suspicious individuals. Conventional systems have difficulty distinguishing between family members and suspicious individuals, and are prone to false detections and malfunctions. This makes it difficult to achieve the truly necessary crime prevention effect, and there is a need for more reliable crime prevention systems. [Means for solving the problem]

[0005] The present invention provides a security system that accurately distinguishes between family members and suspicious individuals by pre-registering the faces, body shapes, and movements of family members and analyzing video and audio data collected by a camera and microphone. Specifically, the system includes a means for registering family members' characteristics, a means for analyzing video and audio data, a means for comparing the analyzed data with family members' characteristics, a means for emitting an alarm or light when a suspicious individual is identified, and a means for summarizing the suspicious individual's characteristics and notifying the user and the security company. This system aims to eliminate the hassle of setup and operation, realize highly accurate, automatic security measures, and be easy to use, especially for elderly people. Furthermore, by adding an audio data analysis function and registering and comparing the user's voice, the accuracy of suspicious individual identification is improved. Furthermore, a user interface is provided that allows users to set the timing and method of notification for suspicious individual identification, enhancing user convenience.

[0006] "Means for registering family members' faces, body shapes, and movements" refers to a function that allows users to input family members' face photos, heights, body shapes, and movement videos into the system in advance, and then store this information in a database.

[0007] "Means for analyzing video and audio data collected by cameras and microphones" refers to a function that processes video and audio data obtained from cameras and microphones and includes algorithms for facial recognition and voice recognition.

[0008] "Means for matching analyzed data with registered family characteristics" refers to a function that compares the features extracted from collected video and audio data with pre-registered family data to check whether they match.

[0009] "Means of emitting a warning sound or light when a suspicious person is identified" refers to a function that warns the system of a suspicious person by sounding a warning sound or flashing a warning light when the system detects a suspicious person.

[0010] "Means of summarizing the characteristics of suspicious individuals and notifying the user and security company" refers to a function that briefly summarizes the characteristics of identified suspicious individuals (e.g., gender, clothing, body shape, movement patterns, etc.) and sends this to the user's smartphone or the security company.

[0011] The "voice data analysis means" is a function that analyzes voice data collected by a microphone and extracts voice features.

[0012] "Means for registering and comparing user voice and speech" refers to a function that allows the system to register the user's voice and speech characteristics in advance and compare them with collected speech data.

[0013] A "user interface that allows you to set the timing and method of notification" is an operation screen or setting function that allows users to customize the timing (e.g., immediate notification, regular reports, etc.) and method (e.g., voice notification, email notification, etc.) of notification for identifying suspicious individuals. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] The present invention provides a security system that can accurately distinguish between family members and suspicious individuals by registering the faces, body shapes, and movements of family members in advance and analyzing video and audio data collected by a camera and microphone. Specific embodiments are described below.

[0036] 1. System Configuration

[0037] The system mainly includes the following components:

[0038] User device: A smartphone or PC operated by the user.

[0039] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[0040] Server: A central management server responsible for collecting, analyzing, and storing data.

[0041] 2. Initial Setup

[0042] User

[0043] Create an account as a new user and log in to the system.

[0044] Photos of family members and related parties' faces, body shapes, video of their movements, and voice data are uploaded and sent to the server.

[0045] server

[0046] A database is constructed based on the received data, and the characteristics of family members and related parties are registered.

[0047] 3. Daily operations

[0048] a. Data collection

[0049] Terminal

[0050] The camera and microphone collect video and audio in real time and send it to the server.

[0051] b. Data analysis and feature matching

[0052] server

[0053] The collected video and audio data is analyzed in real time to extract features of family members and suspicious individuals.

[0054] The extracted features are compared with data registered in a database to identify suspicious individuals.

[0055] c. Alarms / Notifications

[0056] server

[0057] If a suspicious person is identified, their characteristics are summarized and notified to the user and the security company.

[0058] The timing and method of notification (immediate notification, periodic reporting, etc.) will be determined according to the system settings.

[0059] Terminal

[0060] It receives notifications from the server and alerts those around it by sounding an alarm or flashing a warning light.

[0061] 4. User Response

[0062] (User)

[0063] Check the notification on your smartphone or PC and check detailed information and video of the suspicious person through the system application.

[0064] If necessary, contact the police or a security company and take appropriate action.

[0065] Program processing explanation

[0066] The program processing for each section will be explained below.

[0067] User Initial Settings

[0068] User

[0069] 1. The user logs in to the security system and uploads photos of the faces, body shapes, and movement videos of family members and other people involved to the application.

[0070] 2. The user records their own or their family's voice into the microphone and sends it to the server.

[0071] server

[0072] 1. The server receives the uploaded data and analyzes it using face recognition models, body shape recognition models, motion recognition models, and voice recognition models.

[0073] 2. Extract features and store them in a database.

[0074] Monitoring and Data Collection

[0075] Terminal

[0076] 1. The monitoring terminal (camera and microphone) monitors the installation location 24 hours a day, collecting video and audio in real time.

[0077] 2. Send the collected data to the server.

[0078] server

[0079] 1. The server analyzes the received data and matches it with the characteristics of registered family members and related parties.

[0080] Identifying and alerting suspicious individuals

[0081] server

[0082] 1. Identify suspicious individuals based on the results of feature matching.

[0083] 2. Summarize the characteristics of the suspicious individual and generate notification content.

[0084] 3. Send notifications to users and security companies.

[0085] Terminal

[0086] 1. Receive notifications from the server and emit a warning sound or flash a warning light.

[0087] User response

[0088] User

[0089] 1. Receive a notification on your smartphone or PC about a suspicious person and check the details.

[0090] 2. If necessary, contact the police or a security company and take appropriate action.

[0091] Specific examples

[0092] For example, when installing this security system in a home, the following specific examples are conceivable.

[0093] 1. Initial Setup

[0094] The user logs in to the application and registers the face photos, height, weight, and walking video of each family member, as well as recording and registering each person's voice.

[0095] 2. Daily operations

[0096] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[0097] The monitoring terminal sends the collected data to a server, which analyzes the data and distinguishes between family members and suspicious individuals.

[0098] 3. Identifying suspicious individuals

[0099] If the server identifies a suspicious person, it summarizes their characteristics (e.g., suspicious behavior or voice) and notifies the user's smartphone.

[0100] The user will check the notification and contact the police if necessary.

[0101] In this way, by using the security system of the present invention, even elderly people can easily operate it, and highly accurate security measures can be realized. Furthermore, in order to improve the accuracy of identifying suspicious individuals, an even safer environment can be provided by including voice data in the analysis.

[0102] The processing flow will be explained below.

[0103] User Initial Settings

[0104] User

[0105] Step 1: The user installs the security system application, creates a new account and logs in.

[0106] Step 2: The user uploads photos of their family and friends, along with their height, body type, and movement videos within the application.

[0107] Step 3: The user uses a microphone to record their own voice or that of their family members and sends it to the server.

[0108] server

[0109] Step 4: The server receives the uploaded face photo, height, body shape, motion video, and audio data.

[0110] Step 5: The server analyzes the received data and extracts features from each data using a face recognition model, a body shape recognition model, a motion recognition model, and a voice recognition model.

[0111] Step 6: Store the features in a database to create profiles of family members and associates.

[0112] Monitoring and Data Collection

[0113] Terminal

[0114] Step 1: The monitoring terminal (camera and microphone) monitors the target area 24 hours a day and collects video and audio in real time.

[0115] Step 2: The device sends the collected video and audio data to the server.

[0116] Data analysis and feature matching

[0117] server

[0118] Step 3: The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[0119] Step 4: The server analyzes the received voice data and extracts voice features using a voice recognition algorithm.

[0120] Step 5: The server matches the extracted features with family and related person data in the database.

[0121] Step 6: The server identifies the suspicious person based on the result of the feature matching.

[0122] Identifying and alerting suspicious individuals

[0123] server

[0124] Step 7: If the server identifies a suspicious person, it summarizes their characteristics (e.g., gender, clothing, body shape, movement patterns, etc.).

[0125] Step 8: The server sends a notification to the user and the security company.

[0126] Step 9: Notifications are sent via message, email, in-app notification, or other methods according to the user's settings.

[0127] Terminal

[0128] Step 10: The device receives the notification from the server and plays an alert or turns on a light.

[0129] User response

[0130] User

[0131] Step 11: The user checks the notification on their smartphone or PC.

[0132] Step 12: The user opens a system application to view the suspicious individual's details and video recording.

[0133] Step 13: If necessary, the user contacts the police or a security company and takes appropriate action.

[0134] This system eliminates the need for setup and operation, and provides highly accurate, automatic crime prevention measures. It is also easy to operate, especially for the elderly, and can provide a safe and secure living environment.

[0135] Example 1

[0136] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0137] Modern security systems often lack the technology to accurately distinguish between family members and related parties and suspicious individuals. This can result in frequent false alarms, while genuine suspicious individuals may be overlooked. Real-time monitoring and notification functions are also inadequate, making it difficult to respond quickly. Furthermore, the inability to analyze multidimensional data, including audio data, reduces the accuracy of security measures. To address these issues, a system is needed that can identify suspicious individuals in real time based on multiple characteristics, such as face, body shape, movement, and voice, and issue appropriate alarms.

[0138] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0139] In this invention, the server includes a means for registering the faces, body shapes, and movements of family members, a means for analyzing the video and audio data collected by the camera and microphone, and a means for comparing the analyzed data with the registered family characteristics, thereby enabling accurate real-time identification of family members, related parties, and suspicious individuals.

[0140] A "family" is a group of people who are related by blood or who live together and are registered in the system.

[0141] "Body shape" refers to the physical appearance characteristics of a user or related person, such as height, weight, and body fat percentage.

[0142] A "motion" is a movement pattern that indicates a specific action or gesture of a person and is registered and recognized by the system.

[0143] "Camera" refers to a photographic device for collecting video data.

[0144] "Microphone" refers to a recording device for collecting audio data.

[0145] "Video data" is digital data containing visual information collected by a camera.

[0146] "Audio data" is digital data containing acoustic information collected by a microphone.

[0147] "Analysis" is the process of examining collected video and audio data using algorithms and models to extract features.

[0148] "Features" are identification information such as facial feature points, body shape data, movement patterns, and voice frequency components.

[0149] A "suspicious person" is a person who is not registered in the system and who the system determines to be suspicious based on an event.

[0150] A "warning sound" is a sound that is emitted when a suspicious person is identified, and is an acoustic signal that alerts those in the vicinity.

[0151] "Light" is the visual warning signal emitted by the warning device.

[0152] "Summarization" is the process of extracting important parts from analyzed data and summarizing them concisely.

[0153] "Notification" is a message that notifies the user or the security company of information about a suspicious person.

[0154] "Push notification" is a real-time notification method that the system proactively sends to the user's device.

[0155] A "monitoring terminal" is a device that collects data and transmits it to a server, and includes a camera and a microphone.

[0156] "Real-time" refers to the near-simultaneous process of collecting and analyzing data.

[0157] The present invention is a system that highly strengthens security functions, and by registering the faces, body shapes, movements, and voices of family members and related parties in advance and analyzing the video and audio data collected by the camera and microphone in real time, it can reliably distinguish between family members and suspicious individuals. Specific embodiments for implementing the present invention will be described below.

[0158] 1. System Configuration

[0159] The system includes the following components:

[0160] User device: A smartphone or PC operated by the user.

[0161] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[0162] Server: A central management server responsible for collecting, analyzing, and storing data.

[0163] 2. Initial Setup

[0164] User

[0165] The user launches the security system application, creates a new account, and logs in to the system. Face photos, body shape information, movement videos, and voice data of family members and related parties are uploaded through the application and sent to the server. Specifically, the user takes a face photo using the smartphone camera and taps the upload button. Body shape information is entered into the text field, and movement videos are taken with the smartphone camera and uploaded by tapping the movement registration button. Audio data is recorded by tapping the record button within the app, and once recording is complete, the user taps the save button to send the data to the server.

[0166] 3. Daily operations

[0167] Monitoring terminal

[0168] The monitoring terminal (camera and microphone) monitors the installation location 24 hours a day, collecting video and audio data in real time. The collected video and audio data is then sent to the server in real time.

[0169] server

[0170] The server receives real-time video and audio data sent from the monitoring terminal. Generative AI models such as face recognition, body shape recognition, movement recognition, and voice recognition are used for analysis. These models are used to analyze the collected data and extract characteristic information. This characteristic information is then compared with the characteristics of family members and related parties registered in a database to identify suspicious individuals.

[0171] 4. Identifying suspicious individuals and issuing alerts

[0172] server

[0173] The server compares the analyzed data in real time with the characteristic data stored in the database. Based on the comparison results, if a person not registered in the system or suspicious behavior or voice patterns are detected, the person is identified as a suspicious person. If a suspicious person is identified, the characteristics, time of detection, and location are summarized and a notification is generated. This notification is sent to the user's smartphone as a push notification.

[0174] Terminal

[0175] The device receives a notification from the server and alerts those around it by sounding an alarm or flashing a warning light.

[0176] 5. User Response

[0177] User

[0178] Users receive a notification on their smartphone or PC that a suspicious person has been identified and can check the details of the suspicious person. Through the application, users can check the suspicious person's video and characteristic information, and if necessary, contact the police or security company and take appropriate action. Specifically, they can call the police by pressing the emergency contact button within the app.

[0179] Prompt Sentence Examples

[0180] Here is an example of a prompt for a generative AI model:

[0181] "Explain a security system that can reliably distinguish between family members and suspicious individuals by registering their faces, body shapes, and movements in advance and analyzing the video and audio data collected by the camera and microphone. Describe in detail the components of this system, the initial setup procedure, the data collection and analysis methods, the process for identifying suspicious individuals, and the notifications received by the user. Also, explain with concrete examples."

[0182] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0183] Step 1:

[0184] Data registration by users

[0185] Input: Smartphone or PC operated by the user, facial photo, body shape data, movement video, and voice data.

[0186] Specific operation: The user launches the security system app, creates a new account, and logs in. They then use their smartphone camera to take photos of their family members or other people involved and tap the upload button. Next, they enter their body shape data, such as height and weight, into the text fields. They then take a video of their movements with the smartphone camera and press the action registration button to upload it. Finally, they use the microphone to record the voices of their family members or other people involved, and once the recording is complete, they press the save button to send it to the server.

[0187] Output: Facial photo data, body shape data, motion video data, and voice data are sent to the server.

[0188] Step 2:

[0189] Data analysis by server

[0190] Input: Face photo, body shape data, motion video, and audio data sent by the user.

[0191] Specific operation: The server receives the uploaded data and begins analysis. A facial recognition model is used to extract feature points from the facial photo, and a body shape recognition model is used to analyze height and weight data. The motion video is analyzed using a motion recognition model to identify unique motion patterns. A voice recognition model is used to analyze the voice data and extract voice features.

[0192] Data processing and data calculation: Analyze each data using a generative AI model and extract feature information.

[0193] Output: The extracted feature information is registered in a database.

[0194] Step 3:

[0195] Real-time data collection via terminal

[0196] Input: Camera and microphone installed on the monitoring terminal.

[0197] How it works: The monitoring device collects video and audio 24 hours a day and converts them into the appropriate format. For example, a camera installed at the entrance records continuously, and a microphone collects the surrounding audio. This data is then sent to the server in real time.

[0198] Output: The video and audio data collected in real time is sent to the server.

[0199] Step 4:

[0200] Real-time analysis and verification by the server

[0201] Input: Real-time video and audio data transmitted from the monitoring terminal.

[0202] Specific operation: The server analyzes the received data in real time. The collected video and audio data is analyzed using a generative AI model to extract feature information.

[0203] Data processing and calculation: The characteristic information obtained through real-time analysis is compared with the characteristics of family members and related parties registered in advance.

[0204] Output: Matching results are generated.

[0205] Step 5:

[0206] Identifying suspicious individuals by the server

[0207] Input: Real-time analysis and matching results.

[0208] Specific operation: Based on the matching results, the server identifies unregistered individuals and suspicious behavior and voices. It summarizes the characteristics of the suspicious individual and the data at the time of detection, and generates a notification.

[0209] Data processing and data calculation: Based on the suspicious person detection algorithm, the characteristics of suspicious persons are stored in a database and summary information is generated.

[0210] Output: A summary of the suspicious person notification is generated.

[0211] Step 6:

[0212] Server sends notifications

[0213] Input: Summary information of the suspect.

[0214] Specific operation: The server generates a notification based on the summary information and sends it to the user or security company as a push notification. The notification includes the specific time, location, and characteristics of the suspicious person.

[0215] Output: A suspicious person notification is sent to the user's smartphone and the security company.

[0216] Step 7:

[0217] Alarm operation by terminal

[0218] Input: Suspicious person notification from the server.

[0219] Specific operation: The device receives a notification from the server and alerts those around it by sounding an alarm or flashing a warning light.

[0220] Output: An alarm is sounded to the surrounding area.

[0221] Step 8:

[0222] User confirmation and response

[0223] Input: Suspicious person notification sent to the user's smartphone or PC.

[0224] Specific operation: The user receives a notification and checks the details of the suspicious person through the application, checks the video and characteristics of the suspicious person, and if necessary, presses the emergency contact button to contact the police or a security company.

[0225] Output: The user requests the police or a security company to take appropriate action.

[0226] (Application example 1)

[0227] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0228] Strengthening safety management and security is an essential issue in modern manufacturing sites. However, it is difficult to completely prevent the intrusion of suspicious individuals into a work environment, and advanced identification technology is required. Furthermore, conventional security systems have difficulty in accurately identifying individuals using motion and voice data, and are unable to cope with complex work environments. Therefore, there is a need for an effective security system specialized for manufacturing sites.

[0229] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0230] In this invention, the server includes means for registering the faces, body shapes, and movements of workers at the manufacturing site, means for analyzing video and audio data collected by the camera and microphone, means for comparing the analyzed data with the characteristics of registered workers, means for emitting an alarm sound or light when a suspicious person is identified, and means for summarizing the characteristics of the suspicious person and notifying the manager and security organization. This makes it possible to strengthen security at the manufacturing site in real time, immediately identify suspicious people, and take appropriate measures.

[0231] A "manufacturing site" is a place where products are produced or assembled, and generally refers to a factory or workshop.

[0232] A "worker" is an individual who performs a specific task or work, and in a manufacturing site, refers to an employee engaged in production or assembly work.

[0233] "Facial recognition" is the process of identifying an individual's face from images captured by a camera and comparing it with pre-registered facial data.

[0234] "Body shape recognition" is the process of analyzing an individual's body shape and posture from video data and comparing it with pre-registered data.

[0235] "Motion recognition" is a technology that analyzes video data acquired from a camera to identify and distinguish individual movements and behaviors.

[0236] "Voice recognition" is a technology that analyzes voice data acquired by a microphone and identifies an individual's voice.

[0237] A "suspicious person" refers to a person who is not registered at the manufacturing site or who behaves inappropriately.

[0238] An "alarm sound" is an audio signal emitted by the system to alert the user to the intrusion of a suspicious person or an abnormality.

[0239] The "light" refers to the optical signal emitted by the system to alert the user to the intrusion of a suspicious person or an abnormality.

[0240] "Signature summarization" is the process of generating a concise summary of a person's appearance and behavioral characteristics when they are identified as a suspicious person.

[0241] "Manager" refers to the person in charge of the operation and safety management of the manufacturing site.

[0242] The "security organization" is a specialized organization responsible for security at the manufacturing site and dealing with suspicious individuals.

[0243] The present invention relates to a system for enhancing security at a manufacturing site. Specific embodiments of the present invention will be described below.

[0244] System Configuration

[0245] The system mainly includes the following components:

[0246] User device: A smartphone or PC operated by an administrator.

[0247] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[0248] Server: A central management server responsible for collecting, analyzing, and storing data.

[0249] Initial Setup

[0250] User

[0251] 1. The administrator logs in to the system and uploads the worker's face photo, body shape, movement video, and voice data.

[0252] 2. The administrator records each person's voice and sends it to the server.

[0253] server

[0254] 1. The server receives the uploaded data and analyzes it using face recognition, body shape recognition, movement recognition, and voice recognition models.

[0255] 2. Extract features and store them in a database.

[0256] Daily operations

[0257] a. Data collection

[0258] Terminal

[0259] 1. Monitoring terminals (cameras and microphones) installed within the factory monitor the plant 24 hours a day, collecting video and audio in real time and sending it to a server.

[0260] b. Data analysis and feature matching

[0261] server

[0262] 1. Analyze collected video and audio data in real time to extract features of workers and suspicious individuals.

[0263] 2. The extracted features are compared with data registered in a database to identify suspicious individuals.

[0264] c. Alarms / Notifications

[0265] server

[0266] 1. If a suspicious individual is identified, summarize their characteristics and notify management and security organizations.

[0267] 2. Determine the timing and method of notification (immediate notification, periodic reporting, etc.) according to the system settings.

[0268] Terminal

[0269] 1. Receives a notification from the server and activates an alarm sound or warning light to alert those in the vicinity.

[0270] User response

[0271] User

[0272] 1. The administrator checks the notification on a smartphone or PC and checks the suspicious person's detailed information and video through the system application.

[0273] 2. If necessary, contact security organizations and take appropriate action.

[0274] Program processing explanation

[0275] server

[0276] Hardware: A server with a powerful GPU.

[0277] Software: Facial recognition model (OpenCV, Dlib), body shape recognition model (BodyPix), movement recognition model (OpenPose), speech recognition model (Google Speech-to-Text API), notification system (Firebase, Push notifications), cloud storage (AWS S3), real-time streaming (WebRTC).

[0278] Processing: Analyzes the received data, stores the features in a database, identifies suspicious individuals, generates notifications, and sends them to administrators and security companies.

[0279] Specific examples

[0280] For example, when installing this security system in a factory, the following specific example can be considered.

[0281] 1. Initial Setup

[0282] The administrator logs into the application and registers the face photos, height, weight, and walking video of all workers, as well as recording and registering each worker's voice.

[0283] 2. Daily operations

[0284] Cameras and microphones are installed at entrances and exits and important areas within the factory, allowing for 24-hour monitoring.

[0285] The monitoring terminal sends the collected data to a server, which analyzes the data and identifies workers and suspicious individuals.

[0286] 3. Identifying suspicious individuals

[0287] If the server identifies a suspicious person, it summarizes their characteristics (such as suspicious behavior or voice) and notifies the administrator's smartphone.

[0288] The administrator will review the notification and contact security or police if necessary.

[0289] Prompt Sentence Examples

[0290] Below are some examples of prompt sentences.

[0291] Implement a security system that analyzes video and audio data collected by cameras and microphones in a factory and compares it with the faces, body shapes, movements, and voices of pre-registered workers to identify suspicious individuals.Specific methods include using facial recognition, body shape recognition, movement recognition, and voice recognition models, and combining alarm and notification systems to build an application that provides managers and security organizations with information about suspicious individuals in real time.

[0292] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0293] Step 1:

[0294] The user logs into an application for administrators and uploads the worker's face photo, body shape data (height, weight), walking video, and voice data. A smartphone or PC is used as input to collect various data (images, videos, and audio files). This input data is then sent to the server.

[0295] Step 2:

[0296] The server receives the uploaded data and analyzes it using face recognition, body recognition, movement recognition, and voice recognition models. Specifically, it analyzes face data using a face recognition model (OpenCV, Dlib), body data using a body recognition model (BodyPix), movement data using a movement recognition model (OpenPose), and voice data using a voice recognition model (Google Speech-to-Text API). Features are extracted and stored in a database.

[0297] Step 3:

[0298] Terminals (monitoring terminals) are installed in the factory and monitor it 24 hours a day using cameras and microphones. The input data collected is real-time video and audio, and this data is sent to the server.

[0299] Step 4:

[0300] The server receives the video and audio data transmitted in real time and analyzes the collected data using each model. Specifically, the video and audio data is analyzed using a face recognition model, a body shape recognition model, a motion recognition model, and a voice recognition model to extract features. These features are then compared and verified with data registered in an existing database. As a result, the worker or suspicious individual is identified.

[0301] Step 5:

[0302] When the server identifies a suspicious individual, it summarizes their characteristics, including suspicious behavior and voice, and generates a warning notification based on this.

[0303] Step 6:

[0304] Server-generated alert notifications are sent to administrators and security organizations. The timing and method of notifications are determined by the system configuration, and can be immediate or periodic.

[0305] Step 7:

[0306] The device receives the notification from the server and activates the alarm device, specifically emitting an alarm sound or flashing a warning light.

[0307] Step 8:

[0308] The user (administrator) checks the notification on their smartphone or PC. The notification contains detailed information and video of the suspicious individual. The notification content is input data, and the administrator checks the situation as output. If necessary, they contact the security organization and take appropriate action.

[0309] Through the above processing steps, this system strengthens the security of the manufacturing site in real time, making it possible to immediately identify suspicious individuals and take appropriate measures.

[0310] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0311] The present invention relates to a security system that accurately distinguishes between family members and suspicious individuals by registering their faces, body shapes, and movements in advance and analyzing video and audio data collected using a camera and microphone. Furthermore, by combining it with an emotion engine, the system aims to recognize the user's emotions and automatically provide appropriate countermeasures. Specific embodiments are described in detail below.

[0312] 1. System Configuration

[0313] The system consists of the following components:

[0314] User device: A smartphone or PC operated by the user.

[0315] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[0316] Server: A central server that collects, analyzes, stores, and recognizes data.

[0317] 2. Initial Setup

[0318] User

[0319] 1. The user installs the security system application, creates a new account and logs in.

[0320] 2. Upload facial photos, body shapes, video of movements, and voice data of family members and related parties to the application and send them to the server.

[0321] server

[0322] 1. Create a database based on the received data and register the characteristics of family members and related parties.

[0323] 3. Daily operations

[0324] a. Data collection

[0325] Terminal

[0326] 1. The monitoring terminal (camera and microphone) monitors the monitored area 24 hours a day and collects video and audio in real time.

[0327] 2. Send the collected data to the server.

[0328] b. Data analysis and feature matching

[0329] server

[0330] 1. Analyze the received video data and extract facial features using a facial recognition algorithm.

[0331] 2. Analyze the received voice data and extract voice features using a voice recognition algorithm.

[0332] 3. The extracted features are compared with family and related person data in the database to identify suspicious individuals.

[0333] 4. Emotion recognition and countermeasures

[0334] a. Emotion recognition

[0335] server

[0336] 1. During the process of analyzing voice data, the emotion engine is used to determine the user's emotions.

[0337] 2. Emotion recognition algorithms analyze the tone, pitch, rhythm, and speed of the voice to determine whether the user is in an emotional state (e.g., fear or panic).

[0338] b. Emotion-based measures

[0339] server

[0340] 1. Adjust the intensity and method of warning sounds and lights based on emotion recognition results. For example, if the user is feeling fear, a stronger warning will be issued.

[0341] 2. Based on the emotion recognition results, the system proposes and implements appropriate countermeasures. For example, if the user feels fear, the system automatically notifies a security company.

[0342] 3. If necessary, notify the user's friends and family and ask for their assistance.

[0343] 5. User Response

[0344] User

[0345] 1. Check the notification on your smartphone or PC and check the suspicious person's details and recorded video through the system application.

[0346] 2. If necessary, contact a security company or the police and take appropriate action.

[0347] Program processing explanation

[0348] The processing of the program of the present invention will be specifically described below.

[0349] User Initial Settings

[0350] User

[0351] 1. The user logs in to the security system and uploads photos of the faces, body shapes, and movement videos of family members and related parties to the application.

[0352] 2. The user uses a microphone to record their own voice or that of their family members and sends it to the server.

[0353] server

[0354] 1. The server receives the uploaded data and analyzes it using facial recognition, body shape recognition, movement recognition, and voice recognition algorithms.

[0355] 2. Extract features and store them in a database.

[0356] Data collection and analysis

[0357] Terminal

[0358] 1. The monitoring terminal collects video and audio in real time and sends them to the server.

[0359] server

[0360] 1. The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[0361] 2. Analyze the voice data and extract voice features using a voice recognition algorithm.

[0362] 3. The server compares the features and identifies suspicious individuals.

[0363] Emotion recognition and countermeasures

[0364] server

[0365] 1. Use an emotion engine to recognize user emotions while analyzing voice data.

[0366] 2. Based on the emotion recognition results, set and implement appropriate warnings and notification methods.

[0367] 3. Automatically notify security companies or contacts for assistance if necessary.

[0368] Specific examples

[0369] For example, when installing this security system in a home, the following specific examples are conceivable.

[0370] 1. Initial Setup

[0371] A user logs into a security system application and configures the cameras and microphones in their home.

[0372] Register photos of each family member's face, height, body shape, and video of their movements. Also, record and register each person's voice.

[0373] 2. Daily operations

[0374] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[0375] The monitoring terminal sends the collected data to a server, which analyzes it.

[0376] 3. Identifying suspicious individuals

[0377] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[0378] The user checks the notification and confirms the details of the suspicious person.

[0379] 4. Emotion Recognition and Countermeasures

[0380] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[0381] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[0382] In this way, the crime prevention system of the present invention enables highly accurate and automatic crime prevention measures. Furthermore, the emotion recognition function allows for quick and appropriate responses even when the user feels fear, providing a safer and more secure living environment.

[0383] The processing flow will be explained below.

[0384] User Initial Settings

[0385] User

[0386] Step 1: The user installs the security system application on their smartphone or PC, creates a new account and logs in.

[0387] Step 2: The user takes and uploads photos of family members and related people, along with their height, body shape, and movement videos, to the application.

[0388] Step 3: The user uses the microphone to record their own voice or that of a family member and sends it to the server via the application.

[0389] server

[0390] Step 4: The server receives the uploaded face photo, height, body shape, motion video, and audio data.

[0391] Step 5: The server analyzes the received data and extracts features from each data using algorithms for face recognition, body shape recognition, movement recognition, and voice recognition.

[0392] Step 6: The server stores the extracted features in a database and creates profiles of family members and related parties.

[0393] Data Collection and Monitoring

[0394] Terminal

[0395] Step 1: The monitoring terminal (camera and microphone) monitors the monitored area 24 hours a day and collects video and audio in real time.

[0396] Step 2: The device sends the collected video and audio data to the server.

[0397] Data analysis and feature matching

[0398] server

[0399] Step 3: The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[0400] Step 4: The server analyzes the received voice data and extracts voice features using a voice recognition algorithm.

[0401] Step 5: The server compares the extracted features with family and related person data in the database and performs identification.

[0402] Emotion recognition and countermeasures

[0403] server

[0404] Step 6: The server uses the emotion engine while analyzing the voice data to determine the user's emotion by analyzing the tone, pitch, rhythm, speed, etc. of the voice.

[0405] Step 7: The server adjusts the intensity and method of the warning sound and light based on the emotion recognition result. For example, if the user feels fear, it will emit a strong warning sound.

[0406] Step 8: The server sends notifications to the user's friends and family to ask for help, if necessary.

[0407] Step 9: Based on the user's emotion recognition results, a function to automatically notify a security company is executed.

[0408] Identifying and alerting suspicious individuals

[0409] server

[0410] Step 10: If the server identifies a suspicious person, it summarizes their characteristics (gender, clothing, body shape, movement patterns, etc.).

[0411] Step 11: The server sends a notification to the user and the security company.

[0412] Step 12: Notifications are sent via message, email, in-app notification, or other methods according to the user's settings.

[0413] Terminal

[0414] Step 13: The device receives the notification from the server and issues an alert or turns on a light.

[0415] User response

[0416] User

[0417] Step 14: The user checks the notification on their smartphone or PC and checks the suspicious person's details and recorded video through the system application.

[0418] Step 15: If necessary, the user contacts a security company or the police and takes appropriate action.

[0419] As a concrete example, when installing this security system at home, the process is as follows:

[0420] 1. Initial Setup

[0421] The user logs in to the application and registers face photos, body shapes, motion videos, and voices of all family members.

[0422] 2. Daily operations

[0423] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[0424] The collected data is sent to a server and analyzed in real time.

[0425] 3. Identifying suspicious individuals

[0426] If the server identifies a suspicious person, it notifies the user of their characteristics.

[0427] The user checks the notification and confirms the details of the suspicious person.

[0428] 4. Emotion Recognition and Countermeasures

[0429] The server analyzes the voice characteristics and if the user is feeling fearful, it will emit a strong warning sound.

[0430] If necessary, notify a security company or acquaintances and request assistance.

[0431] As a result, the security system of the present invention eliminates the hassle of setup and operation and provides highly accurate, automatic security measures. It can also be easily used by elderly people, helping to create a safe and secure living environment.

[0432] Example 2

[0433] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0434] The main function of existing security systems is to identify suspicious individuals by registering the faces, body shapes, and movements of family members and related parties and analyzing the data collected by cameras and microphones. However, these systems lack the ability to recognize the user's emotional state and take appropriate action based on that. This makes it difficult to respond quickly and appropriately when the user becomes frightened or panics.

[0435] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for registering personal recognition data of family members, means for analyzing video and audio data collected by the image capture device and the audio capture device, means for comparing the analyzed data with registered family characteristics, means for emitting an alarm sound or light when a suspicious person is identified, means for summarizing the characteristics of the suspicious person and notifying the user and a security agency, means for analyzing the user's emotional state using voice recognition, and means for adjusting the method of warning and notification based on the emotional state. This enables a quick and appropriate response even when the user is feeling fear or panic.

[0436] "Family member personal recognition data" refers to data including facial photographs, body shapes, motion videos, and voice features of family members.

[0437] "Image capture device" refers to a camera or other video capture device that collects video data of a monitored area.

[0438] "Audio capture device" refers to a microphone or other audio capture device that collects audio data from a monitored area.

[0439] "Means of analysis" refers to algorithms and software for processing collected video and audio data and extracting meaningful features from it.

[0440] The "matching means" refers to the technology and method for comparing the analyzed features with pre-registered personal recognition data of family members and determining the degree of match.

[0441] "Means for emitting warning sounds or lights" refers to a device or system that emits sounds or lights to warn a user when a suspicious person is identified.

[0442] The "means for summarizing characteristics and notifying the user and security agency" refers to the technology and method for summarizing the identification information of a suspicious person and transmitting that information to the user's terminal and the security agency.

[0443] "Means for analyzing a user's emotional state using voice recognition" refers to algorithms or software that analyze a user's voice data to determine their emotional state (e.g., fear, anger, calm, etc.).

[0444] "Means for adjusting warning and notification methods based on emotional state" refers to techniques and methods for changing the intensity and method of warning sounds and notifications depending on the analyzed emotional state of the user.

[0445] This invention relates to a crime prevention system that can accurately distinguish between family members and suspicious individuals by registering face, body shape, and movement data of family members in advance and analyzing video and audio data collected using an image capture device and an audio capture device. Furthermore, by combining it with an emotion engine, the system aims to recognize the user's emotional state and automatically provide appropriate countermeasures.

[0446] System Configuration

[0447] The system consists of the following components:

[0448] User device: A smartphone or PC operated by the user.

[0449] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[0450] Server: A central server that collects, analyzes, stores, and recognizes data.

[0451] Initial Setup

[0452] User

[0453] 1. The user installs the security system application on their smartphone or PC, creates a new account and logs in.

[0454] 2. Within the application, take or select and upload photos of your family members' faces, body shapes and movement videos. Also, use the microphone to record the voices of all family members and send them to the server.

[0455] server

[0456] 1. The server analyzes the received facial photos, body shape, motion video, and audio data and extracts each feature. TensorFlow is used for facial recognition, Google Cloud Speech-to-Text API for voice recognition, and OpenCV for body shape and motion recognition.

[0457] 2. Save the extracted features in a database (e.g., MySQL).

[0458] Daily operations

[0459] Terminal

[0460] 1. Surveillance devices (e.g. security cameras and microphones) monitor the area 24 hours a day.

[0461] 2. The monitoring device sends the collected video and audio data to the server in real time using AWS Lambda.

[0462] server

[0463] 1. The server analyzes the received video data using a facial recognition algorithm (e.g., DeepFace) and extracts facial features.

[0464] 2. The voice data is analyzed using a voice recognition algorithm (e.g., DeepSpeech) to extract voice features.

[0465] 3. The features are matched with existing data in the database to identify suspicious individuals. This matching is performed using SQL queries.

[0466] Emotion recognition and countermeasures

[0467] server

[0468] 1. The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions based on the voice data collected in real time.

[0469] 2. Emotion recognition algorithms are used to analyze the tone, pitch, rhythm, and speed of the voice to determine the user's emotional state (e.g., fear or panic).

[0470] 3. Depending on the emotional state, take action, such as "set alarm sound to maximum level" or "increase lighting."

[0471] 4. If necessary, use the Twilio API to notify security companies or friends.

[0472] Specific examples

[0473] For example, when installing this security system in a home, the following specific examples are possible:

[0474] 1. Initial Setup

[0475] A user logs into a security system application (e.g., FamilySafe) and configures their home camera and microphone.

[0476] Register photos of each family member's face, body shape, and movement videos, and also record and register each person's voice.

[0477] 2. Daily operations

[0478] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[0479] The monitoring terminal sends the collected data to a server, which analyzes it.

[0480] 3. Identifying suspicious individuals

[0481] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[0482] The user checks the notification and checks the suspicious person's details and recorded video.

[0483] 4. Emotion Recognition and Countermeasures

[0484] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[0485] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[0486] Prompt Sentence Examples

[0487] Here are some example prompts to input to a generative AI model:

[0488] "Please explain a system that identifies suspicious faces and voices and analyzes user emotions from video and audio data collected by home surveillance cameras."

[0489] The crime prevention system of this invention enables highly accurate and automatic crime prevention measures. Furthermore, the emotion recognition function allows for quick and appropriate responses even when the user feels fear, providing a safer and more secure living environment.

[0490] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0491] Step 1: User Data Entry

[0492] User

[0493] Users install the security system application on their smartphone or PC, create a new account, and log in. Within the application, they take or select and upload photos of their family members' faces, body shapes, and movement videos. They also use a microphone to record the voices of all family members and send them to the server.

[0494] Input: Face photo, body shape, motion video, audio recording

[0495] Output: Personal identification data sent to the server

[0496] Step 2: Data analysis and registration by the server

[0497] server

[0498] The server analyzes the received facial photos, body shape, motion video, and audio data, extracting each feature. This process uses TensorFlow and the Google Cloud Speech-to-Text API, and image data is analyzed using OpenCV. For example, it performs tasks such as "extracting facial feature points," "measuring height," and "identifying motion patterns." The extracted features are then stored in a database (e.g., MySQL).

[0499] Input: Personal identification data submitted by the user

[0500] Output: Analyzed feature data registered in the database

[0501] Step 3: Collect data from the device

[0502] Terminal

[0503] The monitoring device (e.g., security camera and microphone) monitors the area 24 hours a day. The monitoring device transmits the collected video and audio data to the server in real time. This is done using AWS Lambda. For example, the camera frame rate is set to 30 fps (30 frames per second), and the microphone is highly sensitive.

[0504] Input: Video and audio data of the monitored area

[0505] Output: Video and audio data sent to the server in real time

[0506] Step 4: Data analysis and collation by the server

[0507] server

[0508] The server analyzes the received video data using a facial recognition algorithm (e.g., DeepFace) to extract facial features. It also analyzes the audio data using a speech recognition algorithm (e.g., DeepSpeech) to extract voice features. These features are then compared with existing data in a database to identify suspicious individuals. For example, it performs tasks such as "calculating the degree of match between facial features and those in the database" and "analyzing the tone and pitch of the voice." Matching is performed using SQL queries.

[0509] Input: Real-time video and audio data from a monitoring terminal

[0510] Output: Identified suspicious person information

[0511] Step 5: Emotion recognition by the server

[0512] server

[0513] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions from the voice data collected in real time. It uses an emotion recognition algorithm to analyze the tone, pitch, rhythm, and speed of the voice to determine the user's emotional state (e.g., fear or panic). For example, it performs "tone analysis of voice data" and "pitch fluctuation analysis."

[0514] Input: Audio data from the monitored area

[0515] Output: Parsed user's emotional state

[0516] Step 6: Server takes action

[0517] server

[0518] Based on the emotion recognition results, the intensity and method of warning sounds and lighting are adjusted. For example, if it detects fear in the user, it can set the warning sound to maximum level or turn on a powerful flashlight. It also uses the Twilio API to send notifications to security companies and acquaintances. For example, if the user feels panicked, it can automatically notify the security company or send an SMS to an acquaintance.

[0519] Input: Analyzed user's emotional state, suspicious person information

[0520] Output: Adjusted alert and notification methods, notifications sent

[0521] (Application example 2)

[0522] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0523] In recent years, there has been a demand for improved home security, but existing security systems are limited to detecting suspicious individuals and are not able to adequately address the fear and anxiety felt by users. Another issue is that when a suspicious individual is detected, the system does not respond quickly, and it takes time to ensure the user's peace of mind. Furthermore, users are unable to flexibly configure the system to suit their needs, which is inconvenient.

[0524] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0525] In this invention, the server includes means for registering the faces, body shapes, and movements of family members, means for analyzing video and audio data collected by the monitoring terminal, means for comparing the analyzed data with the registered family members' characteristics, means for emitting an alarm sound or light when a suspicious person is identified, means for summarizing the suspicious person's characteristics and notifying the user and the security company, means for recognizing the user's emotions from the voice, and means for issuing a strong alarm or automatically reporting to the security company when the user feels fear. This enables quick detection and response of suspicious people as well as security measures that take the user's emotions into consideration.

[0526] "Means for registering family members' faces, body shapes, and movements" refers to a device or method for registering and saving photographs of the faces, body shapes, and movements of family members and related parties as digital data.

[0527] "Means for analyzing video and audio data collected by a camera and a microphone" refers to a device or method that analyzes video and audio data collected by a camera or a microphone and extracts specific features from them.

[0528] "Means for matching analyzed data with registered family characteristics" refers to a device or method that compares the data obtained by analysis with pre-registered family data to determine whether there is a match.

[0529] "Means for emitting a warning sound or light when a suspicious person is identified" refers to a device or method that issues a warning by sounding a warning sound or flashing a light when the system identifies a suspicious person.

[0530] "Means for summarizing the characteristics of a suspicious person and notifying the user and security company" refers to a device or method that, when a suspicious person is identified, summarizes the characteristics and sends a notification to the user's device or security company.

[0531] The "means for recognizing a user's emotion from voice" refers to a device or method for analyzing a user's voice data and identifying the emotion.

[0532] "Means for issuing a strong warning or automatically notifying a security company if the user feels fear" refers to a device or method that analyzes the emotions from the user's voice and, if it determines that the user feels fear, issues a strong warning sound and light or automatically notifies a security company.

[0533] As an embodiment of the present invention, a home security system will be described in detail. This security system ensures safety and security at home by detecting suspicious individuals and responding quickly, as well as providing appropriate measures based on the user's emotions. The specific configuration and operation of the system will be described below.

[0534] 1. System Configuration

[0535] The system consists of the following components:

[0536] User device: A smartphone or PC operated by the user.

[0537] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[0538] Server: A central server that collects, analyzes, stores, and recognizes emotions.

[0539] 2. Initial Setup

[0540] User terminal

[0541] 1. The user installs the security system application, creates a new account and logs in.

[0542] 2. Upload facial photos, body shapes, video of movements, and voice data of family members and related parties to the application and send them to the server.

[0543] server

[0544] 1. Create a database based on the received data and register the characteristics of family members and related parties.

[0545] 2. The registered feature data is analyzed using a recognition algorithm and used for subsequent matching.

[0546] 3. Daily operations

[0547] a. Data collection

[0548] Monitoring terminal

[0549] 1. The monitoring terminal monitors the area 24 hours a day and collects video and audio in real time.

[0550] 2. Send the collected data to the server.

[0551] b. Data analysis and feature matching

[0552] server

[0553] 1. Analyze the received video data using "OpenCV" and the "face_recognition" library to extract facial features.

[0554] 2. The received voice data is analyzed using voice recognition software to extract voice features.

[0555] 3. The extracted features are compared with family and related person data in the database to identify suspicious individuals.

[0556] 4. Emotion recognition and countermeasures

[0557] a. Emotion recognition

[0558] server

[0559] 1. During the process of analyzing the voice data, the "EmotionRecognizer" emotion engine is used to determine the user's emotions.

[0560] 2. Emotion recognition algorithms analyze the tone, pitch, rhythm, and speed of the voice to determine whether the user is in an emotional state (e.g., fear or panic).

[0561] b. Emotion-based measures

[0562] server

[0563] 1. Adjust the intensity and method of warning sounds and lights based on emotion recognition results. For example, if the user is feeling fear, a stronger warning will be issued.

[0564] 2. Based on the emotion recognition results, appropriate countermeasures are proposed and implemented. For example, if the user feels fear, a security company is automatically notified.

[0565] 3. If necessary, notify the user's friends and family and ask for their assistance.

[0566] Specific examples

[0567] For example, when installing this security system in a home, the following specific examples are conceivable.

[0568] 1. Initial Setup

[0569] A user logs into a security system application and configures the cameras and microphones in their home.

[0570] Register photos of each family member's face, height, body shape, and video of their movements. Also, record and register each person's voice.

[0571] 2. Daily operations

[0572] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[0573] The monitoring terminal sends the collected data to a server, which analyzes it.

[0574] 3. Identifying suspicious individuals

[0575] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[0576] The user checks the notification and confirms the details of the suspicious person.

[0577] 4. Emotion Recognition and Countermeasures

[0578] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[0579] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[0580] Specific prompt examples

[0581] "Design a smart home security system that automatically alerts a security company if it detects fear or anxiety in the user. It uses cameras and microphones to collect data in real time, performs facial and emotion recognition, and notifies the security company in a specified manner. Please provide example code for implementation."

[0582] In this way, by implementing the security system of the present invention, it becomes possible to detect suspicious individuals quickly and with high accuracy and to take measures based on the user's emotions, thereby effectively improving the safety and security of the home.

[0583] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0584] Step 1:

[0585] Initial Setup

[0586] Input: The user uploads facial photos, body shapes, movement videos, and voice data of family members and related parties to the security system app.

[0587] Data processing / calculation: The server analyzes the received data and extracts features using algorithms for face recognition, body shape recognition, movement recognition, and voice recognition.

[0588] Output: Save these features in a database.

[0589] Specific operation: The user uploads family information using a smartphone application. The server receives the data, extracts features using various recognition algorithms, and saves them.

[0590] Step 2:

[0591] Real-time data collection

[0592] Input: A monitoring terminal collects video and audio data in real time.

[0593] Data processing / calculation: Video data is captured using a camera, and audio data is collected using a microphone. The collected data is sent to a server.

[0594] Output: The server receives the video and audio data.

[0595] Specific operation: The surveillance cameras and microphones are always on, and the collected data is periodically uploaded to the server.

[0596] Step 3:

[0597] Data analysis and collation

[0598] Input: Video and audio data received by the server from the surveillance terminal.

[0599] Data processing / calculation: The server analyzes the facial features of the received video data using "OpenCV" and the "face_recognition" library. The audio data is analyzed using voice recognition software to extract voice features.

[0600] Output: The features are matched with family data in the database to identify suspicious individuals.

[0601] How it works: The server analyzes the data in real time and compares it with registered family members and associates. If the data does not match, the person is identified as a suspicious person.

[0602] Step 4:

[0603] emotion recognition

[0604] Input: The server analyzes the audio data in real time and recognizes the user's voice.

[0605] Data processing / calculation: Use the "EmotionRecognizer" emotion engine to extract emotional data (tone, pitch, rhythm, speed, etc.) from audio data.

[0606] Output: Obtain the emotion recognition result and determine the type of emotion (e.g., fear or panic).

[0607] Specific operation: The server analyzes the user's voice using an emotion engine to determine the user's emotional state.

[0608] Step 5:

[0609] Measures implemented

[0610] Input: Emotion recognition results and suspicious person identification results.

[0611] Data processing / calculation: Based on the emotion recognition results, the intensity of the warning sound and light is adjusted. If the user feels fear, the system will automatically notify a security company.

[0612] Output: Emit appropriate alarm sounds and lights, and notify the security company.

[0613] Specific actions: The server determines the action to take based on the emotion recognition results, and if necessary, notifies a security company. For example, if the user feels fear, the server will emit a loud alarm and automatically notify a security company.

[0614] Step 6:

[0615] Notification and confirmation

[0616] Input: Suspicious person identification results and response results.

[0617] Data processing / calculation: The server summarizes the characteristics of the suspicious individual and the results of countermeasures, and notifies the user and the security company.

[0618] Output: Send a notification to your smartphone or security company's device.

[0619] How it works: Users receive a notification on their smartphone and can check the details of the suspicious individual. Security companies also receive a notification so they can take appropriate action.

[0620] The above are the specific processing steps in the embodiment of the invention, which enable quick and highly accurate detection of suspicious individuals and flexible countermeasures based on the user's emotions.

[0621] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0622] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0623] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0624] [Second embodiment]

[0625] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0626] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0627] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0628] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0629] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0630] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0631] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0632] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0633] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0634] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0635] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0636] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0637] The present invention provides a security system that can accurately distinguish between family members and suspicious individuals by registering the faces, body shapes, and movements of family members in advance and analyzing video and audio data collected by a camera and microphone. Specific embodiments are described below.

[0638] 1. System Configuration

[0639] The system mainly includes the following components:

[0640] User device: A smartphone or PC operated by the user.

[0641] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[0642] Server: A central management server responsible for collecting, analyzing, and storing data.

[0643] 2. Initial Setup

[0644] User

[0645] Create an account as a new user and log in to the system.

[0646] Photos of family members and related parties' faces, body shapes, video of their movements, and voice data are uploaded and sent to the server.

[0647] server

[0648] A database is constructed based on the received data, and the characteristics of family members and related parties are registered.

[0649] 3. Daily operations

[0650] a. Data collection

[0651] Terminal

[0652] The camera and microphone collect video and audio in real time and send it to the server.

[0653] b. Data analysis and feature matching

[0654] server

[0655] The collected video and audio data is analyzed in real time to extract features of family members and suspicious individuals.

[0656] The extracted features are compared with data registered in a database to identify suspicious individuals.

[0657] c. Alarms / Notifications

[0658] server

[0659] If a suspicious person is identified, their characteristics are summarized and notified to the user and the security company.

[0660] The timing and method of notification (immediate notification, periodic reporting, etc.) will be determined according to the system settings.

[0661] Terminal

[0662] It receives notifications from the server and alerts those around it by sounding an alarm or flashing a warning light.

[0663] 4. User Response

[0664] (User)

[0665] Check the notification on your smartphone or PC and check detailed information and video of the suspicious person through the system application.

[0666] If necessary, contact the police or a security company and take appropriate action.

[0667] Program processing explanation

[0668] The program processing for each section will be explained below.

[0669] User Initial Settings

[0670] User

[0671] 1. The user logs in to the security system and uploads photos of the faces, body shapes, and movement videos of family members and other people involved to the application.

[0672] 2. The user records their own or their family's voice into the microphone and sends it to the server.

[0673] server

[0674] 1. The server receives the uploaded data and analyzes it using face recognition models, body shape recognition models, motion recognition models, and voice recognition models.

[0675] 2. Extract features and store them in a database.

[0676] Monitoring and Data Collection

[0677] Terminal

[0678] 1. The monitoring terminal (camera and microphone) monitors the installation location 24 hours a day, collecting video and audio in real time.

[0679] 2. Send the collected data to the server.

[0680] server

[0681] 1. The server analyzes the received data and matches it with the characteristics of registered family members and related parties.

[0682] Identifying and alerting suspicious individuals

[0683] server

[0684] 1. Identify suspicious individuals based on the results of feature matching.

[0685] 2. Summarize the characteristics of the suspicious individual and generate notification content.

[0686] 3. Send notifications to users and security companies.

[0687] Terminal

[0688] 1. Receive notifications from the server and emit a warning sound or flash a warning light.

[0689] User response

[0690] User

[0691] 1. Receive a notification on your smartphone or PC about a suspicious person and check the details.

[0692] 2. If necessary, contact the police or a security company and take appropriate action.

[0693] Specific examples

[0694] For example, when installing this security system in a home, the following specific examples are conceivable.

[0695] 1. Initial Setup

[0696] The user logs in to the application and registers the face photos, height, weight, and walking video of each family member, as well as recording and registering each person's voice.

[0697] 2. Daily operations

[0698] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[0699] The monitoring terminal sends the collected data to a server, which analyzes the data and distinguishes between family members and suspicious individuals.

[0700] 3. Identifying suspicious individuals

[0701] If the server identifies a suspicious person, it summarizes their characteristics (e.g., suspicious behavior or voice) and notifies the user's smartphone.

[0702] The user will check the notification and contact the police if necessary.

[0703] In this way, by using the security system of the present invention, even elderly people can easily operate it, and highly accurate security measures can be realized. Furthermore, in order to improve the accuracy of identifying suspicious individuals, an even safer environment can be provided by including voice data in the analysis.

[0704] The processing flow will be explained below.

[0705] User Initial Settings

[0706] User

[0707] Step 1: The user installs the security system application, creates a new account and logs in.

[0708] Step 2: The user uploads photos of their family and friends, along with their height, body type, and movement videos within the application.

[0709] Step 3: The user uses a microphone to record their own voice or that of their family members and sends it to the server.

[0710] server

[0711] Step 4: The server receives the uploaded face photo, height, body shape, motion video, and audio data.

[0712] Step 5: The server analyzes the received data and extracts features from each data using a face recognition model, a body shape recognition model, a motion recognition model, and a voice recognition model.

[0713] Step 6: Store the features in a database to create profiles of family members and associates.

[0714] Monitoring and Data Collection

[0715] Terminal

[0716] Step 1: The monitoring terminal (camera and microphone) monitors the target area 24 hours a day and collects video and audio in real time.

[0717] Step 2: The device sends the collected video and audio data to the server.

[0718] Data analysis and feature matching

[0719] server

[0720] Step 3: The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[0721] Step 4: The server analyzes the received voice data and extracts voice features using a voice recognition algorithm.

[0722] Step 5: The server matches the extracted features with family and related person data in the database.

[0723] Step 6: The server identifies the suspicious person based on the result of the feature matching.

[0724] Identifying and alerting suspicious individuals

[0725] server

[0726] Step 7: If the server identifies a suspicious person, it summarizes their characteristics (e.g., gender, clothing, body shape, movement patterns, etc.).

[0727] Step 8: The server sends a notification to the user and the security company.

[0728] Step 9: Notifications are sent via message, email, in-app notification, or other methods according to the user's settings.

[0729] Terminal

[0730] Step 10: The device receives the notification from the server and plays an alert or turns on a light.

[0731] User response

[0732] User

[0733] Step 11: The user checks the notification on their smartphone or PC.

[0734] Step 12: The user opens a system application to view the suspicious individual's details and video recording.

[0735] Step 13: If necessary, the user contacts the police or a security company and takes appropriate action.

[0736] This system eliminates the need for setup and operation, and provides highly accurate, automatic crime prevention measures. It is also easy to operate, especially for the elderly, and can provide a safe and secure living environment.

[0737] Example 1

[0738] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0739] Modern security systems often lack the technology to accurately distinguish between family members and related parties and suspicious individuals. This can result in frequent false alarms, while genuine suspicious individuals may be overlooked. Real-time monitoring and notification functions are also inadequate, making it difficult to respond quickly. Furthermore, the inability to analyze multidimensional data, including audio data, reduces the accuracy of security measures. To address these issues, a system is needed that can identify suspicious individuals in real time based on multiple characteristics, such as face, body shape, movement, and voice, and issue appropriate alarms.

[0740] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0741] In this invention, the server includes a means for registering the faces, body shapes, and movements of family members, a means for analyzing the video and audio data collected by the camera and microphone, and a means for comparing the analyzed data with the registered family characteristics, thereby enabling accurate real-time identification of family members, related parties, and suspicious individuals.

[0742] A "family" is a group of people who are related by blood or who live together and are registered in the system.

[0743] "Body shape" refers to the physical appearance characteristics of a user or related person, such as height, weight, and body fat percentage.

[0744] A "motion" is a movement pattern that indicates a specific action or gesture of a person and is registered and recognized by the system.

[0745] "Camera" refers to a photographic device for collecting video data.

[0746] "Microphone" refers to a recording device for collecting audio data.

[0747] "Video data" is digital data containing visual information collected by a camera.

[0748] "Audio data" is digital data containing acoustic information collected by a microphone.

[0749] "Analysis" is the process of examining collected video and audio data using algorithms and models to extract features.

[0750] "Features" are identification information such as facial feature points, body shape data, movement patterns, and voice frequency components.

[0751] A "suspicious person" is a person who is not registered in the system and who the system determines to be suspicious based on an event.

[0752] A "warning sound" is a sound that is emitted when a suspicious person is identified, and is an acoustic signal that alerts those in the vicinity.

[0753] "Light" is the visual warning signal emitted by the warning device.

[0754] "Summarization" is the process of extracting important parts from analyzed data and summarizing them concisely.

[0755] "Notification" is a message that notifies the user or the security company of information about a suspicious person.

[0756] "Push notification" is a real-time notification method that the system proactively sends to the user's device.

[0757] A "monitoring terminal" is a device that collects data and transmits it to a server, and includes a camera and a microphone.

[0758] "Real-time" refers to the near-simultaneous process of collecting and analyzing data.

[0759] The present invention is a system that highly strengthens security functions, and by registering the faces, body shapes, movements, and voices of family members and related parties in advance and analyzing the video and audio data collected by the camera and microphone in real time, it can reliably distinguish between family members and suspicious individuals. Specific embodiments for implementing the present invention will be described below.

[0760] 1. System Configuration

[0761] The system includes the following components:

[0762] User device: A smartphone or PC operated by the user.

[0763] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[0764] Server: A central management server responsible for collecting, analyzing, and storing data.

[0765] 2. Initial Setup

[0766] User

[0767] The user launches the security system application, creates a new account, and logs in to the system. Face photos, body shape information, movement videos, and voice data of family members and related parties are uploaded through the application and sent to the server. Specifically, the user takes a face photo using the smartphone camera and taps the upload button. Body shape information is entered into the text field, and movement videos are taken with the smartphone camera and uploaded by tapping the movement registration button. Audio data is recorded by tapping the record button within the app, and once recording is complete, the user taps the save button to send the data to the server.

[0768] 3. Daily operations

[0769] Monitoring terminal

[0770] The monitoring terminal (camera and microphone) monitors the installation location 24 hours a day, collecting video and audio data in real time. The collected video and audio data is then sent to the server in real time.

[0771] server

[0772] The server receives real-time video and audio data sent from the monitoring terminal. Generative AI models such as face recognition, body shape recognition, movement recognition, and voice recognition are used for analysis. These models are used to analyze the collected data and extract characteristic information. This characteristic information is then compared with the characteristics of family members and related parties registered in a database to identify suspicious individuals.

[0773] 4. Identifying suspicious individuals and issuing alerts

[0774] server

[0775] The server compares the analyzed data in real time with the characteristic data stored in the database. Based on the comparison results, if a person not registered in the system or suspicious behavior or voice patterns are detected, the person is identified as a suspicious person. If a suspicious person is identified, the characteristics, time of detection, and location are summarized and a notification is generated. This notification is sent to the user's smartphone as a push notification.

[0776] Terminal

[0777] The device receives a notification from the server and alerts those around it by sounding an alarm or flashing a warning light.

[0778] 5. User Response

[0779] User

[0780] Users receive a notification on their smartphone or PC that a suspicious person has been identified and can check the details of the suspicious person. Through the application, users can check the suspicious person's video and characteristic information, and if necessary, contact the police or security company and take appropriate action. Specifically, they can call the police by pressing the emergency contact button within the app.

[0781] Prompt Sentence Examples

[0782] Here is an example of a prompt for a generative AI model:

[0783] "Explain a security system that can reliably distinguish between family members and suspicious individuals by registering their faces, body shapes, and movements in advance and analyzing the video and audio data collected by the camera and microphone. Describe in detail the components of this system, the initial setup procedure, the data collection and analysis methods, the process for identifying suspicious individuals, and the notifications received by the user. Also, explain with concrete examples."

[0784] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0785] Step 1:

[0786] Data registration by users

[0787] Input: Smartphone or PC operated by the user, facial photo, body shape data, movement video, and voice data.

[0788] Specific operation: The user launches the security system app, creates a new account, and logs in. They then use their smartphone camera to take photos of their family members or other people involved and tap the upload button. Next, they enter their body shape data, such as height and weight, into the text fields. They then take a video of their movements with the smartphone camera and press the action registration button to upload it. Finally, they use the microphone to record the voices of their family members or other people involved, and once the recording is complete, they press the save button to send it to the server.

[0789] Output: Facial photo data, body shape data, motion video data, and voice data are sent to the server.

[0790] Step 2:

[0791] Data analysis by server

[0792] Input: Face photo, body shape data, motion video, and audio data sent by the user.

[0793] Specific operation: The server receives the uploaded data and begins analysis. A facial recognition model is used to extract feature points from the facial photo, and a body shape recognition model is used to analyze height and weight data. The motion video is analyzed using a motion recognition model to identify unique motion patterns. A voice recognition model is used to analyze the voice data and extract voice features.

[0794] Data processing and data calculation: Analyze each data using a generative AI model and extract feature information.

[0795] Output: The extracted feature information is registered in a database.

[0796] Step 3:

[0797] Real-time data collection via terminal

[0798] Input: Camera and microphone installed on the monitoring terminal.

[0799] How it works: The monitoring device collects video and audio 24 hours a day and converts them into the appropriate format. For example, a camera installed at the entrance records continuously, and a microphone collects the surrounding audio. This data is then sent to the server in real time.

[0800] Output: The video and audio data collected in real time is sent to the server.

[0801] Step 4:

[0802] Real-time analysis and verification by the server

[0803] Input: Real-time video and audio data transmitted from the monitoring terminal.

[0804] Specific operation: The server analyzes the received data in real time. The collected video and audio data is analyzed using a generative AI model to extract feature information.

[0805] Data processing and calculation: The characteristic information obtained through real-time analysis is compared with the characteristics of family members and related parties registered in advance.

[0806] Output: Matching results are generated.

[0807] Step 5:

[0808] Identifying suspicious individuals by the server

[0809] Input: Real-time analysis and matching results.

[0810] Specific operation: Based on the matching results, the server identifies unregistered individuals and suspicious behavior and voices. It summarizes the characteristics of the suspicious individual and the data at the time of detection, and generates a notification.

[0811] Data processing and data calculation: Based on the suspicious person detection algorithm, the characteristics of suspicious persons are stored in a database and summary information is generated.

[0812] Output: A summary of the suspicious person notification is generated.

[0813] Step 6:

[0814] Server sends notifications

[0815] Input: Summary information of the suspect.

[0816] Specific operation: The server generates a notification based on the summary information and sends it to the user or security company as a push notification. The notification includes the specific time, location, and characteristics of the suspicious person.

[0817] Output: A suspicious person notification is sent to the user's smartphone and the security company.

[0818] Step 7:

[0819] Alarm operation by terminal

[0820] Input: Suspicious person notification from the server.

[0821] Specific operation: The device receives a notification from the server and alerts those around it by sounding an alarm or flashing a warning light.

[0822] Output: An alarm is sounded to the surrounding area.

[0823] Step 8:

[0824] User confirmation and response

[0825] Input: Suspicious person notification sent to the user's smartphone or PC.

[0826] Specific operation: The user receives a notification and checks the details of the suspicious person through the application, checks the video and characteristics of the suspicious person, and if necessary, presses the emergency contact button to contact the police or a security company.

[0827] Output: The user requests the police or a security company to take appropriate action.

[0828] (Application example 1)

[0829] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0830] Strengthening safety management and security is an essential issue in modern manufacturing sites. However, it is difficult to completely prevent the intrusion of suspicious individuals into a work environment, and advanced identification technology is required. Furthermore, conventional security systems have difficulty in accurately identifying individuals using motion and voice data, and are unable to cope with complex work environments. Therefore, there is a need for an effective security system specialized for manufacturing sites.

[0831] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0832] In this invention, the server includes means for registering the faces, body shapes, and movements of workers at the manufacturing site, means for analyzing video and audio data collected by the camera and microphone, means for comparing the analyzed data with the characteristics of registered workers, means for emitting an alarm sound or light when a suspicious person is identified, and means for summarizing the characteristics of the suspicious person and notifying the manager and security organization. This makes it possible to strengthen security at the manufacturing site in real time, immediately identify suspicious people, and take appropriate measures.

[0833] A "manufacturing site" is a place where products are produced or assembled, and generally refers to a factory or workshop.

[0834] A "worker" is an individual who performs a specific task or work, and in a manufacturing site, refers to an employee engaged in production or assembly work.

[0835] "Facial recognition" is the process of identifying an individual's face from images captured by a camera and comparing it with pre-registered facial data.

[0836] "Body shape recognition" is the process of analyzing an individual's body shape and posture from video data and comparing it with pre-registered data.

[0837] "Motion recognition" is a technology that analyzes video data acquired from a camera to identify and distinguish individual movements and behaviors.

[0838] "Voice recognition" is a technology that analyzes voice data acquired by a microphone and identifies an individual's voice.

[0839] A "suspicious person" refers to a person who is not registered at the manufacturing site or who behaves inappropriately.

[0840] An "alarm sound" is an audio signal emitted by the system to alert the user to the intrusion of a suspicious person or an abnormality.

[0841] The "light" refers to the optical signal emitted by the system to alert the user to the intrusion of a suspicious person or an abnormality.

[0842] "Signature summarization" is the process of generating a concise summary of a person's appearance and behavioral characteristics when they are identified as a suspicious person.

[0843] "Manager" refers to the person in charge of the operation and safety management of the manufacturing site.

[0844] The "security organization" is a specialized organization responsible for security at the manufacturing site and dealing with suspicious individuals.

[0845] The present invention relates to a system for enhancing security at a manufacturing site. Specific embodiments of the present invention will be described below.

[0846] System Configuration

[0847] The system mainly includes the following components:

[0848] User device: A smartphone or PC operated by an administrator.

[0849] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[0850] Server: A central management server responsible for collecting, analyzing, and storing data.

[0851] Initial Setup

[0852] User

[0853] 1. The administrator logs in to the system and uploads the worker's face photo, body shape, movement video, and voice data.

[0854] 2. The administrator records each person's voice and sends it to the server.

[0855] server

[0856] 1. The server receives the uploaded data and analyzes it using face recognition, body shape recognition, movement recognition, and voice recognition models.

[0857] 2. Extract features and store them in a database.

[0858] Daily operations

[0859] a. Data collection

[0860] Terminal

[0861] 1. Monitoring terminals (cameras and microphones) installed within the factory monitor the plant 24 hours a day, collecting video and audio in real time and sending it to a server.

[0862] b. Data analysis and feature matching

[0863] server

[0864] 1. Analyze collected video and audio data in real time to extract features of workers and suspicious individuals.

[0865] 2. The extracted features are compared with data registered in a database to identify suspicious individuals.

[0866] c. Alarms / Notifications

[0867] server

[0868] 1. If a suspicious individual is identified, summarize their characteristics and notify management and security organizations.

[0869] 2. Determine the timing and method of notification (immediate notification, periodic reporting, etc.) according to the system settings.

[0870] Terminal

[0871] 1. Receives a notification from the server and activates an alarm sound or warning light to alert those in the vicinity.

[0872] User response

[0873] User

[0874] 1. The administrator checks the notification on a smartphone or PC and checks the suspicious person's detailed information and video through the system application.

[0875] 2. If necessary, contact security organizations and take appropriate action.

[0876] Program processing explanation

[0877] server

[0878] Hardware: A server with a powerful GPU.

[0879] Software: Facial recognition model (OpenCV, Dlib), body shape recognition model (BodyPix), movement recognition model (OpenPose), speech recognition model (Google Speech-to-Text API), notification system (Firebase, Push notifications), cloud storage (AWS S3), real-time streaming (WebRTC).

[0880] Processing: Analyzes the received data, stores the features in a database, identifies suspicious individuals, generates notifications, and sends them to administrators and security companies.

[0881] Specific examples

[0882] For example, when installing this security system in a factory, the following specific example can be considered.

[0883] 1. Initial Setup

[0884] The administrator logs into the application and registers the face photos, height, weight, and walking video of all workers, as well as recording and registering each worker's voice.

[0885] 2. Daily operations

[0886] Cameras and microphones are installed at entrances and exits and important areas within the factory, allowing for 24-hour monitoring.

[0887] The monitoring terminal sends the collected data to a server, which analyzes the data and identifies workers and suspicious individuals.

[0888] 3. Identifying suspicious individuals

[0889] If the server identifies a suspicious person, it summarizes their characteristics (such as suspicious behavior or voice) and notifies the administrator's smartphone.

[0890] The administrator will review the notification and contact security or police if necessary.

[0891] Prompt Sentence Examples

[0892] Below are some examples of prompt sentences.

[0893] Implement a security system that analyzes video and audio data collected by cameras and microphones in a factory and compares it with the faces, body shapes, movements, and voices of pre-registered workers to identify suspicious individuals.Specific methods include using facial recognition, body shape recognition, movement recognition, and voice recognition models, and combining alarm and notification systems to build an application that provides managers and security organizations with information about suspicious individuals in real time.

[0894] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0895] Step 1:

[0896] The user logs into an application for administrators and uploads the worker's face photo, body shape data (height, weight), walking video, and voice data. A smartphone or PC is used as input to collect various data (images, videos, and audio files). This input data is then sent to the server.

[0897] Step 2:

[0898] The server receives the uploaded data and analyzes it using face recognition, body recognition, movement recognition, and voice recognition models. Specifically, it analyzes face data using a face recognition model (OpenCV, Dlib), body data using a body recognition model (BodyPix), movement data using a movement recognition model (OpenPose), and voice data using a voice recognition model (Google Speech-to-Text API). Features are extracted and stored in a database.

[0899] Step 3:

[0900] Terminals (monitoring terminals) are installed in the factory and monitor it 24 hours a day using cameras and microphones. The input data collected is real-time video and audio, and this data is sent to the server.

[0901] Step 4:

[0902] The server receives the video and audio data transmitted in real time and analyzes the collected data using each model. Specifically, the video and audio data is analyzed using a face recognition model, a body shape recognition model, a motion recognition model, and a voice recognition model to extract features. These features are then compared and verified with data registered in an existing database. As a result, the worker or suspicious individual is identified.

[0903] Step 5:

[0904] When the server identifies a suspicious individual, it summarizes their characteristics, including suspicious behavior and voice, and generates a warning notification based on this.

[0905] Step 6:

[0906] Server-generated alert notifications are sent to administrators and security organizations. The timing and method of notifications are determined by the system configuration, and can be immediate or periodic.

[0907] Step 7:

[0908] The device receives the notification from the server and activates the alarm device, specifically emitting an alarm sound or flashing a warning light.

[0909] Step 8:

[0910] The user (administrator) checks the notification on their smartphone or PC. The notification contains detailed information and video of the suspicious individual. The notification content is input data, and the administrator checks the situation as output. If necessary, they contact the security organization and take appropriate action.

[0911] Through the above processing steps, this system strengthens the security of the manufacturing site in real time, making it possible to immediately identify suspicious individuals and take appropriate measures.

[0912] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0913] The present invention relates to a security system that accurately distinguishes between family members and suspicious individuals by registering their faces, body shapes, and movements in advance and analyzing video and audio data collected using a camera and microphone. Furthermore, by combining it with an emotion engine, the system aims to recognize the user's emotions and automatically provide appropriate countermeasures. Specific embodiments are described in detail below.

[0914] 1. System Configuration

[0915] The system consists of the following components:

[0916] User device: A smartphone or PC operated by the user.

[0917] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[0918] Server: A central server that collects, analyzes, stores, and recognizes data.

[0919] 2. Initial Setup

[0920] User

[0921] 1. The user installs the security system application, creates a new account and logs in.

[0922] 2. Upload facial photos, body shapes, video of movements, and voice data of family members and related parties to the application and send them to the server.

[0923] server

[0924] 1. Create a database based on the received data and register the characteristics of family members and related parties.

[0925] 3. Daily operations

[0926] a. Data collection

[0927] Terminal

[0928] 1. The monitoring terminal (camera and microphone) monitors the monitored area 24 hours a day and collects video and audio in real time.

[0929] 2. Send the collected data to the server.

[0930] b. Data analysis and feature matching

[0931] server

[0932] 1. Analyze the received video data and extract facial features using a facial recognition algorithm.

[0933] 2. Analyze the received voice data and extract voice features using a voice recognition algorithm.

[0934] 3. The extracted features are compared with family and related person data in the database to identify suspicious individuals.

[0935] 4. Emotion recognition and countermeasures

[0936] a. Emotion recognition

[0937] server

[0938] 1. During the process of analyzing voice data, the emotion engine is used to determine the user's emotions.

[0939] 2. Emotion recognition algorithms analyze the tone, pitch, rhythm, and speed of the voice to determine whether the user is in an emotional state (e.g., fear or panic).

[0940] b. Emotion-based measures

[0941] server

[0942] 1. Adjust the intensity and method of warning sounds and lights based on emotion recognition results. For example, if the user is feeling fear, a stronger warning will be issued.

[0943] 2. Based on the emotion recognition results, the system proposes and implements appropriate countermeasures. For example, if the user feels fear, the system automatically notifies a security company.

[0944] 3. If necessary, notify the user's friends and family and ask for their assistance.

[0945] 5. User Response

[0946] User

[0947] 1. Check the notification on your smartphone or PC and check the suspicious person's details and recorded video through the system application.

[0948] 2. If necessary, contact a security company or the police and take appropriate action.

[0949] Program processing explanation

[0950] The processing of the program of the present invention will be specifically described below.

[0951] User Initial Settings

[0952] User

[0953] 1. The user logs in to the security system and uploads photos of the faces, body shapes, and movement videos of family members and related parties to the application.

[0954] 2. The user uses a microphone to record their own voice or that of their family members and sends it to the server.

[0955] server

[0956] 1. The server receives the uploaded data and analyzes it using facial recognition, body shape recognition, movement recognition, and voice recognition algorithms.

[0957] 2. Extract features and store them in a database.

[0958] Data collection and analysis

[0959] Terminal

[0960] 1. The monitoring terminal collects video and audio in real time and sends them to the server.

[0961] server

[0962] 1. The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[0963] 2. Analyze the voice data and extract voice features using a voice recognition algorithm.

[0964] 3. The server compares the features and identifies suspicious individuals.

[0965] Emotion recognition and countermeasures

[0966] server

[0967] 1. Use an emotion engine to recognize user emotions while analyzing voice data.

[0968] 2. Based on the emotion recognition results, set and implement appropriate warnings and notification methods.

[0969] 3. Automatically notify security companies or contacts for assistance if necessary.

[0970] Specific examples

[0971] For example, when installing this security system in a home, the following specific examples are conceivable.

[0972] 1. Initial Setup

[0973] A user logs into a security system application and configures the cameras and microphones in their home.

[0974] Register photos of each family member's face, height, body shape, and video of their movements. Also, record and register each person's voice.

[0975] 2. Daily operations

[0976] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[0977] The monitoring terminal sends the collected data to a server, which analyzes it.

[0978] 3. Identifying suspicious individuals

[0979] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[0980] The user checks the notification and confirms the details of the suspicious person.

[0981] 4. Emotion Recognition and Countermeasures

[0982] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[0983] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[0984] In this way, the crime prevention system of the present invention enables highly accurate and automatic crime prevention measures. Furthermore, the emotion recognition function allows for quick and appropriate responses even when the user feels fear, providing a safer and more secure living environment.

[0985] The processing flow will be explained below.

[0986] User Initial Settings

[0987] User

[0988] Step 1: The user installs the security system application on their smartphone or PC, creates a new account and logs in.

[0989] Step 2: The user takes and uploads photos of family members and related people, along with their height, body shape, and movement videos, to the application.

[0990] Step 3: The user uses the microphone to record their own voice or that of a family member and sends it to the server via the application.

[0991] server

[0992] Step 4: The server receives the uploaded face photo, height, body shape, motion video, and audio data.

[0993] Step 5: The server analyzes the received data and extracts features from each data using algorithms for face recognition, body shape recognition, movement recognition, and voice recognition.

[0994] Step 6: The server stores the extracted features in a database and creates profiles of family members and related parties.

[0995] Data Collection and Monitoring

[0996] Terminal

[0997] Step 1: The monitoring terminal (camera and microphone) monitors the monitored area 24 hours a day and collects video and audio in real time.

[0998] Step 2: The device sends the collected video and audio data to the server.

[0999] Data analysis and feature matching

[1000] server

[1001] Step 3: The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[1002] Step 4: The server analyzes the received voice data and extracts voice features using a voice recognition algorithm.

[1003] Step 5: The server compares the extracted features with family and related person data in the database and performs identification.

[1004] Emotion recognition and countermeasures

[1005] server

[1006] Step 6: The server uses the emotion engine while analyzing the voice data to determine the user's emotion by analyzing the tone, pitch, rhythm, speed, etc. of the voice.

[1007] Step 7: The server adjusts the intensity and method of the warning sound and light based on the emotion recognition result. For example, if the user feels fear, it will emit a strong warning sound.

[1008] Step 8: The server sends notifications to the user's friends and family to ask for help, if necessary.

[1009] Step 9: Based on the user's emotion recognition results, a function to automatically notify a security company is executed.

[1010] Identifying and alerting suspicious individuals

[1011] server

[1012] Step 10: If the server identifies a suspicious person, it summarizes their characteristics (gender, clothing, body shape, movement patterns, etc.).

[1013] Step 11: The server sends a notification to the user and the security company.

[1014] Step 12: Notifications are sent via message, email, in-app notification, or other methods according to the user's settings.

[1015] Terminal

[1016] Step 13: The device receives the notification from the server and issues an alert or turns on a light.

[1017] User response

[1018] User

[1019] Step 14: The user checks the notification on their smartphone or PC and checks the suspicious person's details and recorded video through the system application.

[1020] Step 15: If necessary, the user contacts a security company or the police and takes appropriate action.

[1021] As a concrete example, when installing this security system at home, the process is as follows:

[1022] 1. Initial Setup

[1023] The user logs in to the application and registers face photos, body shapes, motion videos, and voices of all family members.

[1024] 2. Daily operations

[1025] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[1026] The collected data is sent to a server and analyzed in real time.

[1027] 3. Identifying suspicious individuals

[1028] If the server identifies a suspicious person, it notifies the user of their characteristics.

[1029] The user checks the notification and confirms the details of the suspicious person.

[1030] 4. Emotion Recognition and Countermeasures

[1031] The server analyzes the voice characteristics and if the user is feeling fearful, it will emit a strong warning sound.

[1032] If necessary, notify a security company or acquaintances and request assistance.

[1033] As a result, the security system of the present invention eliminates the hassle of setup and operation and provides highly accurate, automatic security measures. It can also be easily used by elderly people, helping to create a safe and secure living environment.

[1034] Example 2

[1035] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1036] The main function of existing security systems is to identify suspicious individuals by registering the faces, body shapes, and movements of family members and related parties and analyzing the data collected by cameras and microphones. However, these systems lack the ability to recognize the user's emotional state and take appropriate action based on that. This makes it difficult to respond quickly and appropriately when the user becomes frightened or panics.

[1037] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for registering personal recognition data of family members, means for analyzing video and audio data collected by the image capture device and the audio capture device, means for comparing the analyzed data with registered family characteristics, means for emitting an alarm sound or light when a suspicious person is identified, means for summarizing the characteristics of the suspicious person and notifying the user and a security agency, means for analyzing the user's emotional state using voice recognition, and means for adjusting the method of warning and notification based on the emotional state. This enables a quick and appropriate response even when the user is feeling fear or panic.

[1038] "Family member personal recognition data" refers to data including facial photographs, body shapes, motion videos, and voice features of family members.

[1039] "Image capture device" refers to a camera or other video capture device that collects video data of a monitored area.

[1040] "Audio capture device" refers to a microphone or other audio capture device that collects audio data from a monitored area.

[1041] "Means of analysis" refers to algorithms and software for processing collected video and audio data and extracting meaningful features from it.

[1042] The "matching means" refers to the technology and method for comparing the analyzed features with pre-registered personal recognition data of family members and determining the degree of match.

[1043] "Means for emitting warning sounds or lights" refers to a device or system that emits sounds or lights to warn a user when a suspicious person is identified.

[1044] The "means for summarizing characteristics and notifying the user and security agency" refers to the technology and method for summarizing the identification information of a suspicious person and transmitting that information to the user's terminal and the security agency.

[1045] "Means for analyzing a user's emotional state using voice recognition" refers to algorithms or software that analyze a user's voice data to determine their emotional state (e.g., fear, anger, calm, etc.).

[1046] "Means for adjusting warning and notification methods based on emotional state" refers to techniques and methods for changing the intensity and method of warning sounds and notifications depending on the analyzed emotional state of the user.

[1047] This invention relates to a crime prevention system that can accurately distinguish between family members and suspicious individuals by registering face, body shape, and movement data of family members in advance and analyzing video and audio data collected using an image capture device and an audio capture device. Furthermore, by combining it with an emotion engine, the system aims to recognize the user's emotional state and automatically provide appropriate countermeasures.

[1048] System Configuration

[1049] The system consists of the following components:

[1050] User device: A smartphone or PC operated by the user.

[1051] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[1052] Server: A central server that collects, analyzes, stores, and recognizes data.

[1053] Initial Setup

[1054] User

[1055] 1. The user installs the security system application on their smartphone or PC, creates a new account and logs in.

[1056] 2. Within the application, take or select and upload photos of your family members' faces, body shapes and movement videos. Also, use the microphone to record the voices of all family members and send them to the server.

[1057] server

[1058] 1. The server analyzes the received facial photos, body shape, motion video, and audio data and extracts each feature. TensorFlow is used for facial recognition, Google Cloud Speech-to-Text API for voice recognition, and OpenCV for body shape and motion recognition.

[1059] 2. Save the extracted features in a database (e.g., MySQL).

[1060] Daily operations

[1061] Terminal

[1062] 1. Surveillance devices (e.g. security cameras and microphones) monitor the area 24 hours a day.

[1063] 2. The monitoring device sends the collected video and audio data to the server in real time using AWS Lambda.

[1064] server

[1065] 1. The server analyzes the received video data using a facial recognition algorithm (e.g., DeepFace) and extracts facial features.

[1066] 2. The voice data is analyzed using a voice recognition algorithm (e.g., DeepSpeech) to extract voice features.

[1067] 3. The features are matched with existing data in the database to identify suspicious individuals. This matching is performed using SQL queries.

[1068] Emotion recognition and countermeasures

[1069] server

[1070] 1. The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions based on the voice data collected in real time.

[1071] 2. Emotion recognition algorithms are used to analyze the tone, pitch, rhythm, and speed of the voice to determine the user's emotional state (e.g., fear or panic).

[1072] 3. Depending on the emotional state, take action, such as "set alarm sound to maximum level" or "increase lighting."

[1073] 4. If necessary, use the Twilio API to notify security companies or friends.

[1074] Specific examples

[1075] For example, when installing this security system in a home, the following specific examples are possible:

[1076] 1. Initial Setup

[1077] A user logs into a security system application (e.g., FamilySafe) and configures their home camera and microphone.

[1078] Register photos of each family member's face, body shape, and movement videos, and also record and register each person's voice.

[1079] 2. Daily operations

[1080] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[1081] The monitoring terminal sends the collected data to a server, which analyzes it.

[1082] 3. Identifying suspicious individuals

[1083] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[1084] The user checks the notification and checks the suspicious person's details and recorded video.

[1085] 4. Emotion Recognition and Countermeasures

[1086] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[1087] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[1088] Prompt Sentence Examples

[1089] Here are some example prompts to input to a generative AI model:

[1090] "Please explain a system that identifies suspicious faces and voices and analyzes user emotions from video and audio data collected by home surveillance cameras."

[1091] The crime prevention system of this invention enables highly accurate and automatic crime prevention measures. Furthermore, the emotion recognition function allows for quick and appropriate responses even when the user feels fear, providing a safer and more secure living environment.

[1092] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1093] Step 1: User Data Entry

[1094] User

[1095] Users install the security system application on their smartphone or PC, create a new account, and log in. Within the application, they take or select and upload photos of their family members' faces, body shapes, and movement videos. They also use a microphone to record the voices of all family members and send them to the server.

[1096] Input: Face photo, body shape, motion video, audio recording

[1097] Output: Personal identification data sent to the server

[1098] Step 2: Data analysis and registration by the server

[1099] server

[1100] The server analyzes the received facial photos, body shape, motion video, and audio data, extracting each feature. This process uses TensorFlow and the Google Cloud Speech-to-Text API, and image data is analyzed using OpenCV. For example, it performs tasks such as "extracting facial feature points," "measuring height," and "identifying motion patterns." The extracted features are then stored in a database (e.g., MySQL).

[1101] Input: Personal identification data submitted by the user

[1102] Output: Analyzed feature data registered in the database

[1103] Step 3: Collect data from the device

[1104] Terminal

[1105] The monitoring device (e.g., security camera and microphone) monitors the area 24 hours a day. The monitoring device transmits the collected video and audio data to the server in real time. This is done using AWS Lambda. For example, the camera frame rate is set to 30 fps (30 frames per second), and the microphone is highly sensitive.

[1106] Input: Video and audio data of the monitored area

[1107] Output: Video and audio data sent to the server in real time

[1108] Step 4: Data analysis and collation by the server

[1109] server

[1110] The server analyzes the received video data using a facial recognition algorithm (e.g., DeepFace) to extract facial features. It also analyzes the audio data using a speech recognition algorithm (e.g., DeepSpeech) to extract voice features. These features are then compared with existing data in a database to identify suspicious individuals. For example, it performs tasks such as "calculating the degree of match between facial features and those in the database" and "analyzing the tone and pitch of the voice." Matching is performed using SQL queries.

[1111] Input: Real-time video and audio data from a monitoring terminal

[1112] Output: Identified suspicious person information

[1113] Step 5: Emotion recognition by the server

[1114] server

[1115] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions from the voice data collected in real time. It uses an emotion recognition algorithm to analyze the tone, pitch, rhythm, and speed of the voice to determine the user's emotional state (e.g., fear or panic). For example, it performs "tone analysis of voice data" and "pitch fluctuation analysis."

[1116] Input: Audio data from the monitored area

[1117] Output: Parsed user's emotional state

[1118] Step 6: Server takes action

[1119] server

[1120] Based on the emotion recognition results, the intensity and method of warning sounds and lighting are adjusted. For example, if it detects fear in the user, it can set the warning sound to maximum level or turn on a powerful flashlight. It also uses the Twilio API to send notifications to security companies and acquaintances. For example, if the user feels panicked, it can automatically notify the security company or send an SMS to an acquaintance.

[1121] Input: Analyzed user's emotional state, suspicious person information

[1122] Output: Adjusted alert and notification methods, notifications sent

[1123] (Application example 2)

[1124] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1125] In recent years, there has been a demand for improved home security, but existing security systems are limited to detecting suspicious individuals and are not able to adequately address the fear and anxiety felt by users. Another issue is that when a suspicious individual is detected, the system does not respond quickly, and it takes time to ensure the user's peace of mind. Furthermore, users are unable to flexibly configure the system to suit their needs, which is inconvenient.

[1126] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1127] In this invention, the server includes means for registering the faces, body shapes, and movements of family members, means for analyzing video and audio data collected by the monitoring terminal, means for comparing the analyzed data with the registered family members' characteristics, means for emitting an alarm sound or light when a suspicious person is identified, means for summarizing the suspicious person's characteristics and notifying the user and the security company, means for recognizing the user's emotions from the voice, and means for issuing a strong alarm or automatically reporting to the security company when the user feels fear. This enables quick detection and response of suspicious people as well as security measures that take the user's emotions into consideration.

[1128] "Means for registering family members' faces, body shapes, and movements" refers to a device or method for registering and saving photographs of the faces, body shapes, and movements of family members and related parties as digital data.

[1129] "Means for analyzing video and audio data collected by a camera and a microphone" refers to a device or method that analyzes video and audio data collected by a camera or a microphone and extracts specific features from them.

[1130] "Means for matching analyzed data with registered family characteristics" refers to a device or method that compares the data obtained by analysis with pre-registered family data to determine whether there is a match.

[1131] "Means for emitting a warning sound or light when a suspicious person is identified" refers to a device or method that issues a warning by sounding a warning sound or flashing a light when the system identifies a suspicious person.

[1132] "Means for summarizing the characteristics of a suspicious person and notifying the user and security company" refers to a device or method that, when a suspicious person is identified, summarizes the characteristics and sends a notification to the user's device or security company.

[1133] The "means for recognizing a user's emotion from voice" refers to a device or method for analyzing a user's voice data and identifying the emotion.

[1134] "Means for issuing a strong warning or automatically notifying a security company if the user feels fear" refers to a device or method that analyzes the emotions from the user's voice and, if it determines that the user feels fear, issues a strong warning sound and light or automatically notifies a security company.

[1135] As an embodiment of the present invention, a home security system will be described in detail. This security system ensures safety and security at home by detecting suspicious individuals and responding quickly, as well as providing appropriate measures based on the user's emotions. The specific configuration and operation of the system will be described below.

[1136] 1. System Configuration

[1137] The system consists of the following components:

[1138] User device: A smartphone or PC operated by the user.

[1139] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[1140] Server: A central server that collects, analyzes, stores, and recognizes emotions.

[1141] 2. Initial Setup

[1142] User terminal

[1143] 1. The user installs the security system application, creates a new account and logs in.

[1144] 2. Upload facial photos, body shapes, video of movements, and voice data of family members and related parties to the application and send them to the server.

[1145] server

[1146] 1. Create a database based on the received data and register the characteristics of family members and related parties.

[1147] 2. The registered feature data is analyzed using a recognition algorithm and used for subsequent matching.

[1148] 3. Daily operations

[1149] a. Data collection

[1150] Monitoring terminal

[1151] 1. The monitoring terminal monitors the area 24 hours a day and collects video and audio in real time.

[1152] 2. Send the collected data to the server.

[1153] b. Data analysis and feature matching

[1154] server

[1155] 1. Analyze the received video data using "OpenCV" and the "face_recognition" library to extract facial features.

[1156] 2. The received voice data is analyzed using voice recognition software to extract voice features.

[1157] 3. The extracted features are compared with family and related person data in the database to identify suspicious individuals.

[1158] 4. Emotion recognition and countermeasures

[1159] a. Emotion recognition

[1160] server

[1161] 1. During the process of analyzing the voice data, the "EmotionRecognizer" emotion engine is used to determine the user's emotions.

[1162] 2. Emotion recognition algorithms analyze the tone, pitch, rhythm, and speed of the voice to determine whether the user is in an emotional state (e.g., fear or panic).

[1163] b. Emotion-based measures

[1164] server

[1165] 1. Adjust the intensity and method of warning sounds and lights based on emotion recognition results. For example, if the user is feeling fear, a stronger warning will be issued.

[1166] 2. Based on the emotion recognition results, appropriate countermeasures are proposed and implemented. For example, if the user feels fear, a security company is automatically notified.

[1167] 3. If necessary, notify the user's friends and family and ask for their assistance.

[1168] Specific examples

[1169] For example, when installing this security system in a home, the following specific examples are conceivable.

[1170] 1. Initial Setup

[1171] A user logs into a security system application and configures the cameras and microphones in their home.

[1172] Register photos of each family member's face, height, body shape, and video of their movements. Also, record and register each person's voice.

[1173] 2. Daily operations

[1174] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[1175] The monitoring terminal sends the collected data to a server, which analyzes it.

[1176] 3. Identifying suspicious individuals

[1177] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[1178] The user checks the notification and confirms the details of the suspicious person.

[1179] 4. Emotion Recognition and Countermeasures

[1180] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[1181] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[1182] Specific prompt examples

[1183] "Design a smart home security system that automatically alerts a security company if it detects fear or anxiety in the user. It uses cameras and microphones to collect data in real time, performs facial and emotion recognition, and notifies the security company in a specified manner. Please provide example code for implementation."

[1184] In this way, by implementing the security system of the present invention, it becomes possible to detect suspicious individuals quickly and with high accuracy and to take measures based on the user's emotions, thereby effectively improving the safety and security of the home.

[1185] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1186] Step 1:

[1187] Initial Setup

[1188] Input: The user uploads facial photos, body shapes, movement videos, and voice data of family members and related parties to the security system app.

[1189] Data processing / calculation: The server analyzes the received data and extracts features using algorithms for face recognition, body shape recognition, movement recognition, and voice recognition.

[1190] Output: Save these features in a database.

[1191] Specific operation: The user uploads family information using a smartphone application. The server receives the data, extracts features using various recognition algorithms, and saves them.

[1192] Step 2:

[1193] Real-time data collection

[1194] Input: A monitoring terminal collects video and audio data in real time.

[1195] Data processing / calculation: Video data is captured using a camera, and audio data is collected using a microphone. The collected data is sent to a server.

[1196] Output: The server receives the video and audio data.

[1197] Specific operation: The surveillance cameras and microphones are always on, and the collected data is periodically uploaded to the server.

[1198] Step 3:

[1199] Data analysis and collation

[1200] Input: Video and audio data received by the server from the surveillance terminal.

[1201] Data processing / calculation: The server analyzes the facial features of the received video data using "OpenCV" and the "face_recognition" library. The audio data is analyzed using voice recognition software to extract voice features.

[1202] Output: The features are matched with family data in the database to identify suspicious individuals.

[1203] How it works: The server analyzes the data in real time and compares it with registered family members and associates. If the data does not match, the person is identified as a suspicious person.

[1204] Step 4:

[1205] emotion recognition

[1206] Input: The server analyzes the audio data in real time and recognizes the user's voice.

[1207] Data processing / calculation: Use the "EmotionRecognizer" emotion engine to extract emotional data (tone, pitch, rhythm, speed, etc.) from audio data.

[1208] Output: Obtain the emotion recognition result and determine the type of emotion (e.g., fear or panic).

[1209] Specific operation: The server analyzes the user's voice using an emotion engine to determine the user's emotional state.

[1210] Step 5:

[1211] Measures implemented

[1212] Input: Emotion recognition results and suspicious person identification results.

[1213] Data processing / calculation: Based on the emotion recognition results, the intensity of the warning sound and light is adjusted. If the user feels fear, the system will automatically notify a security company.

[1214] Output: Emit appropriate alarm sounds and lights, and notify the security company.

[1215] Specific actions: The server determines the action to take based on the emotion recognition results, and if necessary, notifies a security company. For example, if the user feels fear, the server will emit a loud alarm and automatically notify a security company.

[1216] Step 6:

[1217] Notification and confirmation

[1218] Input: Suspicious person identification results and response results.

[1219] Data processing / calculation: The server summarizes the characteristics of the suspicious individual and the results of countermeasures, and notifies the user and the security company.

[1220] Output: Send a notification to your smartphone or security company's device.

[1221] How it works: Users receive a notification on their smartphone and can check the details of the suspicious individual. Security companies also receive a notification so they can take appropriate action.

[1222] The above are the specific processing steps in the embodiment of the invention, which enable quick and highly accurate detection of suspicious individuals and flexible countermeasures based on the user's emotions.

[1223] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1224] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1225] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1226] [Third embodiment]

[1227] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1228] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1229] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1230] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1231] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1232] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1233] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1234] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1235] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1236] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1237] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1238] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1239] The present invention provides a security system that can accurately distinguish between family members and suspicious individuals by registering the faces, body shapes, and movements of family members in advance and analyzing video and audio data collected by a camera and microphone. Specific embodiments are described below.

[1240] 1. System Configuration

[1241] The system mainly includes the following components:

[1242] User device: A smartphone or PC operated by the user.

[1243] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[1244] Server: A central management server responsible for collecting, analyzing, and storing data.

[1245] 2. Initial Setup

[1246] User

[1247] Create an account as a new user and log in to the system.

[1248] Photos of family members and related parties' faces, body shapes, video of their movements, and voice data are uploaded and sent to the server.

[1249] server

[1250] A database is constructed based on the received data, and the characteristics of family members and related parties are registered.

[1251] 3. Daily operations

[1252] a. Data collection

[1253] Terminal

[1254] The camera and microphone collect video and audio in real time and send it to the server.

[1255] b. Data analysis and feature matching

[1256] server

[1257] The collected video and audio data is analyzed in real time to extract features of family members and suspicious individuals.

[1258] The extracted features are compared with data registered in a database to identify suspicious individuals.

[1259] c. Alarms / Notifications

[1260] server

[1261] If a suspicious person is identified, their characteristics are summarized and notified to the user and the security company.

[1262] The timing and method of notification (immediate notification, periodic reporting, etc.) will be determined according to the system settings.

[1263] Terminal

[1264] It receives notifications from the server and alerts those around it by sounding an alarm or flashing a warning light.

[1265] 4. User Response

[1266] (User)

[1267] Check the notification on your smartphone or PC and check detailed information and video of the suspicious person through the system application.

[1268] If necessary, contact the police or a security company and take appropriate action.

[1269] Program processing explanation

[1270] The program processing for each section will be explained below.

[1271] User Initial Settings

[1272] User

[1273] 1. The user logs in to the security system and uploads photos of the faces, body shapes, and movement videos of family members and other people involved to the application.

[1274] 2. The user records their own or their family's voice into the microphone and sends it to the server.

[1275] server

[1276] 1. The server receives the uploaded data and analyzes it using face recognition models, body shape recognition models, motion recognition models, and voice recognition models.

[1277] 2. Extract features and store them in a database.

[1278] Monitoring and Data Collection

[1279] Terminal

[1280] 1. The monitoring terminal (camera and microphone) monitors the installation location 24 hours a day, collecting video and audio in real time.

[1281] 2. Send the collected data to the server.

[1282] server

[1283] 1. The server analyzes the received data and matches it with the characteristics of registered family members and related parties.

[1284] Identifying and alerting suspicious individuals

[1285] server

[1286] 1. Identify suspicious individuals based on the results of feature matching.

[1287] 2. Summarize the characteristics of the suspicious individual and generate notification content.

[1288] 3. Send notifications to users and security companies.

[1289] Terminal

[1290] 1. Receive notifications from the server and emit a warning sound or flash a warning light.

[1291] User response

[1292] User

[1293] 1. Receive a notification on your smartphone or PC about a suspicious person and check the details.

[1294] 2. If necessary, contact the police or a security company and take appropriate action.

[1295] Specific examples

[1296] For example, when installing this security system in a home, the following specific examples are conceivable.

[1297] 1. Initial Setup

[1298] The user logs in to the application and registers the face photos, height, weight, and walking video of each family member, as well as recording and registering each person's voice.

[1299] 2. Daily operations

[1300] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[1301] The monitoring terminal sends the collected data to a server, which analyzes the data and distinguishes between family members and suspicious individuals.

[1302] 3. Identifying suspicious individuals

[1303] If the server identifies a suspicious person, it summarizes their characteristics (e.g., suspicious behavior or voice) and notifies the user's smartphone.

[1304] The user will check the notification and contact the police if necessary.

[1305] In this way, by using the security system of the present invention, even elderly people can easily operate it, and highly accurate security measures can be realized. Furthermore, in order to improve the accuracy of identifying suspicious individuals, an even safer environment can be provided by including voice data in the analysis.

[1306] The processing flow will be explained below.

[1307] User Initial Settings

[1308] User

[1309] Step 1: The user installs the security system application, creates a new account and logs in.

[1310] Step 2: The user uploads photos of their family and friends, along with their height, body type, and movement videos within the application.

[1311] Step 3: The user uses a microphone to record their own voice or that of their family members and sends it to the server.

[1312] server

[1313] Step 4: The server receives the uploaded face photo, height, body shape, motion video, and audio data.

[1314] Step 5: The server analyzes the received data and extracts features from each data using a face recognition model, a body shape recognition model, a motion recognition model, and a voice recognition model.

[1315] Step 6: Store the features in a database to create profiles of family members and associates.

[1316] Monitoring and Data Collection

[1317] Terminal

[1318] Step 1: The monitoring terminal (camera and microphone) monitors the target area 24 hours a day and collects video and audio in real time.

[1319] Step 2: The device sends the collected video and audio data to the server.

[1320] Data analysis and feature matching

[1321] server

[1322] Step 3: The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[1323] Step 4: The server analyzes the received voice data and extracts voice features using a voice recognition algorithm.

[1324] Step 5: The server matches the extracted features with family and related person data in the database.

[1325] Step 6: The server identifies the suspicious person based on the result of the feature matching.

[1326] Identifying and alerting suspicious individuals

[1327] server

[1328] Step 7: If the server identifies a suspicious person, it summarizes their characteristics (e.g., gender, clothing, body shape, movement patterns, etc.).

[1329] Step 8: The server sends a notification to the user and the security company.

[1330] Step 9: Notifications are sent via message, email, in-app notification, or other methods according to the user's settings.

[1331] Terminal

[1332] Step 10: The device receives the notification from the server and plays an alert or turns on a light.

[1333] User response

[1334] User

[1335] Step 11: The user checks the notification on their smartphone or PC.

[1336] Step 12: The user opens a system application to view the suspicious individual's details and video recording.

[1337] Step 13: If necessary, the user contacts the police or a security company and takes appropriate action.

[1338] This system eliminates the need for setup and operation, and provides highly accurate, automatic crime prevention measures. It is also easy to operate, especially for the elderly, and can provide a safe and secure living environment.

[1339] Example 1

[1340] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1341] Modern security systems often lack the technology to accurately distinguish between family members and related parties and suspicious individuals. This can result in frequent false alarms, while genuine suspicious individuals may be overlooked. Real-time monitoring and notification functions are also inadequate, making it difficult to respond quickly. Furthermore, the inability to analyze multidimensional data, including audio data, reduces the accuracy of security measures. To address these issues, a system is needed that can identify suspicious individuals in real time based on multiple characteristics, such as face, body shape, movement, and voice, and issue appropriate alarms.

[1342] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1343] In this invention, the server includes a means for registering the faces, body shapes, and movements of family members, a means for analyzing the video and audio data collected by the camera and microphone, and a means for comparing the analyzed data with the registered family characteristics, thereby enabling accurate real-time identification of family members, related parties, and suspicious individuals.

[1344] A "family" is a group of people who are related by blood or who live together and are registered in the system.

[1345] "Body shape" refers to the physical appearance characteristics of a user or related person, such as height, weight, and body fat percentage.

[1346] A "motion" is a movement pattern that indicates a specific action or gesture of a person and is registered and recognized by the system.

[1347] "Camera" refers to a photographic device for collecting video data.

[1348] "Microphone" refers to a recording device for collecting audio data.

[1349] "Video data" is digital data containing visual information collected by a camera.

[1350] "Audio data" is digital data containing acoustic information collected by a microphone.

[1351] "Analysis" is the process of examining collected video and audio data using algorithms and models to extract features.

[1352] "Features" are identification information such as facial feature points, body shape data, movement patterns, and voice frequency components.

[1353] A "suspicious person" is a person who is not registered in the system and who the system determines to be suspicious based on an event.

[1354] A "warning sound" is a sound that is emitted when a suspicious person is identified, and is an acoustic signal that alerts those in the vicinity.

[1355] "Light" is the visual warning signal emitted by the warning device.

[1356] "Summarization" is the process of extracting important parts from analyzed data and summarizing them concisely.

[1357] "Notification" is a message that notifies the user or the security company of information about a suspicious person.

[1358] "Push notification" is a real-time notification method that the system proactively sends to the user's device.

[1359] A "monitoring terminal" is a device that collects data and transmits it to a server, and includes a camera and a microphone.

[1360] "Real-time" refers to the near-simultaneous process of collecting and analyzing data.

[1361] The present invention is a system that highly strengthens security functions, and by registering the faces, body shapes, movements, and voices of family members and related parties in advance and analyzing the video and audio data collected by the camera and microphone in real time, it can reliably distinguish between family members and suspicious individuals. Specific embodiments for implementing the present invention will be described below.

[1362] 1. System Configuration

[1363] The system includes the following components:

[1364] User device: A smartphone or PC operated by the user.

[1365] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[1366] Server: A central management server responsible for collecting, analyzing, and storing data.

[1367] 2. Initial Setup

[1368] User

[1369] The user launches the security system application, creates a new account, and logs in to the system. Face photos, body shape information, movement videos, and voice data of family members and related parties are uploaded through the application and sent to the server. Specifically, the user takes a face photo using the smartphone camera and taps the upload button. Body shape information is entered into the text field, and movement videos are taken with the smartphone camera and uploaded by tapping the movement registration button. Audio data is recorded by tapping the record button within the app, and once recording is complete, the user taps the save button to send the data to the server.

[1370] 3. Daily operations

[1371] Monitoring terminal

[1372] The monitoring terminal (camera and microphone) monitors the installation location 24 hours a day, collecting video and audio data in real time. The collected video and audio data is then sent to the server in real time.

[1373] server

[1374] The server receives real-time video and audio data sent from the monitoring terminal. Generative AI models such as face recognition, body shape recognition, movement recognition, and voice recognition are used for analysis. These models are used to analyze the collected data and extract characteristic information. This characteristic information is then compared with the characteristics of family members and related parties registered in a database to identify suspicious individuals.

[1375] 4. Identifying suspicious individuals and issuing alerts

[1376] server

[1377] The server compares the analyzed data in real time with the characteristic data stored in the database. Based on the comparison results, if a person not registered in the system or suspicious behavior or voice patterns are detected, the person is identified as a suspicious person. If a suspicious person is identified, the characteristics, time of detection, and location are summarized and a notification is generated. This notification is sent to the user's smartphone as a push notification.

[1378] Terminal

[1379] The device receives a notification from the server and alerts those around it by sounding an alarm or flashing a warning light.

[1380] 5. User Response

[1381] User

[1382] Users receive a notification on their smartphone or PC that a suspicious person has been identified and can check the details of the suspicious person. Through the application, users can check the suspicious person's video and characteristic information, and if necessary, contact the police or security company and take appropriate action. Specifically, they can call the police by pressing the emergency contact button within the app.

[1383] Prompt Sentence Examples

[1384] Here is an example of a prompt for a generative AI model:

[1385] "Explain a security system that can reliably distinguish between family members and suspicious individuals by registering their faces, body shapes, and movements in advance and analyzing the video and audio data collected by the camera and microphone. Describe in detail the components of this system, the initial setup procedure, the data collection and analysis methods, the process for identifying suspicious individuals, and the notifications received by the user. Also, explain with concrete examples."

[1386] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1387] Step 1:

[1388] Data registration by users

[1389] Input: Smartphone or PC operated by the user, facial photo, body shape data, movement video, and voice data.

[1390] Specific operation: The user launches the security system app, creates a new account, and logs in. They then use their smartphone camera to take photos of their family members or other people involved and tap the upload button. Next, they enter their body shape data, such as height and weight, into the text fields. They then take a video of their movements with the smartphone camera and press the action registration button to upload it. Finally, they use the microphone to record the voices of their family members or other people involved, and once the recording is complete, they press the save button to send it to the server.

[1391] Output: Facial photo data, body shape data, motion video data, and voice data are sent to the server.

[1392] Step 2:

[1393] Data analysis by server

[1394] Input: Face photo, body shape data, motion video, and audio data sent by the user.

[1395] Specific operation: The server receives the uploaded data and begins analysis. A facial recognition model is used to extract feature points from the facial photo, and a body shape recognition model is used to analyze height and weight data. The motion video is analyzed using a motion recognition model to identify unique motion patterns. A voice recognition model is used to analyze the voice data and extract voice features.

[1396] Data processing and data calculation: Analyze each data using a generative AI model and extract feature information.

[1397] Output: The extracted feature information is registered in a database.

[1398] Step 3:

[1399] Real-time data collection via terminal

[1400] Input: Camera and microphone installed on the monitoring terminal.

[1401] How it works: The monitoring device collects video and audio 24 hours a day and converts them into the appropriate format. For example, a camera installed at the entrance records continuously, and a microphone collects the surrounding audio. This data is then sent to the server in real time.

[1402] Output: The video and audio data collected in real time is sent to the server.

[1403] Step 4:

[1404] Real-time analysis and verification by the server

[1405] Input: Real-time video and audio data transmitted from the monitoring terminal.

[1406] Specific operation: The server analyzes the received data in real time. The collected video and audio data is analyzed using a generative AI model to extract feature information.

[1407] Data processing and calculation: The characteristic information obtained through real-time analysis is compared with the characteristics of family members and related parties registered in advance.

[1408] Output: Matching results are generated.

[1409] Step 5:

[1410] Identifying suspicious individuals by the server

[1411] Input: Real-time analysis and matching results.

[1412] Specific operation: Based on the matching results, the server identifies unregistered individuals and suspicious behavior and voices. It summarizes the characteristics of the suspicious individual and the data at the time of detection, and generates a notification.

[1413] Data processing and data calculation: Based on the suspicious person detection algorithm, the characteristics of suspicious persons are stored in a database and summary information is generated.

[1414] Output: A summary of the suspicious person notification is generated.

[1415] Step 6:

[1416] Server sends notifications

[1417] Input: Summary information of the suspect.

[1418] Specific operation: The server generates a notification based on the summary information and sends it to the user or security company as a push notification. The notification includes the specific time, location, and characteristics of the suspicious person.

[1419] Output: A suspicious person notification is sent to the user's smartphone and the security company.

[1420] Step 7:

[1421] Alarm operation by terminal

[1422] Input: Suspicious person notification from the server.

[1423] Specific operation: The device receives a notification from the server and alerts those around it by sounding an alarm or flashing a warning light.

[1424] Output: An alarm is sounded to the surrounding area.

[1425] Step 8:

[1426] User confirmation and response

[1427] Input: Suspicious person notification sent to the user's smartphone or PC.

[1428] Specific operation: The user receives a notification and checks the details of the suspicious person through the application, checks the video and characteristics of the suspicious person, and if necessary, presses the emergency contact button to contact the police or a security company.

[1429] Output: The user requests the police or a security company to take appropriate action.

[1430] (Application example 1)

[1431] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1432] Strengthening safety management and security is an essential issue in modern manufacturing sites. However, it is difficult to completely prevent the intrusion of suspicious individuals into a work environment, and advanced identification technology is required. Furthermore, conventional security systems have difficulty in accurately identifying individuals using motion and voice data, and are unable to cope with complex work environments. Therefore, there is a need for an effective security system specialized for manufacturing sites.

[1433] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1434] In this invention, the server includes means for registering the faces, body shapes, and movements of workers at the manufacturing site, means for analyzing video and audio data collected by the camera and microphone, means for comparing the analyzed data with the characteristics of registered workers, means for emitting an alarm sound or light when a suspicious person is identified, and means for summarizing the characteristics of the suspicious person and notifying the manager and security organization. This makes it possible to strengthen security at the manufacturing site in real time, immediately identify suspicious people, and take appropriate measures.

[1435] A "manufacturing site" is a place where products are produced or assembled, and generally refers to a factory or workshop.

[1436] A "worker" is an individual who performs a specific task or work, and in a manufacturing site, refers to an employee engaged in production or assembly work.

[1437] "Facial recognition" is the process of identifying an individual's face from images captured by a camera and comparing it with pre-registered facial data.

[1438] "Body shape recognition" is the process of analyzing an individual's body shape and posture from video data and comparing it with pre-registered data.

[1439] "Motion recognition" is a technology that analyzes video data acquired from a camera to identify and distinguish individual movements and behaviors.

[1440] "Voice recognition" is a technology that analyzes voice data acquired by a microphone and identifies an individual's voice.

[1441] A "suspicious person" refers to a person who is not registered at the manufacturing site or who behaves inappropriately.

[1442] An "alarm sound" is an audio signal emitted by the system to alert the user to the intrusion of a suspicious person or an abnormality.

[1443] The "light" refers to the optical signal emitted by the system to alert the user to the intrusion of a suspicious person or an abnormality.

[1444] "Signature summarization" is the process of generating a concise summary of a person's appearance and behavioral characteristics when they are identified as a suspicious person.

[1445] "Manager" refers to the person in charge of the operation and safety management of the manufacturing site.

[1446] The "security organization" is a specialized organization responsible for security at the manufacturing site and dealing with suspicious individuals.

[1447] The present invention relates to a system for enhancing security at a manufacturing site. Specific embodiments of the present invention will be described below.

[1448] System Configuration

[1449] The system mainly includes the following components:

[1450] User device: A smartphone or PC operated by an administrator.

[1451] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[1452] Server: A central management server responsible for collecting, analyzing, and storing data.

[1453] Initial Setup

[1454] User

[1455] 1. The administrator logs in to the system and uploads the worker's face photo, body shape, movement video, and voice data.

[1456] 2. The administrator records each person's voice and sends it to the server.

[1457] server

[1458] 1. The server receives the uploaded data and analyzes it using face recognition, body shape recognition, movement recognition, and voice recognition models.

[1459] 2. Extract features and store them in a database.

[1460] Daily operations

[1461] a. Data collection

[1462] Terminal

[1463] 1. Monitoring terminals (cameras and microphones) installed within the factory monitor the plant 24 hours a day, collecting video and audio in real time and sending it to a server.

[1464] b. Data analysis and feature matching

[1465] server

[1466] 1. Analyze collected video and audio data in real time to extract features of workers and suspicious individuals.

[1467] 2. The extracted features are compared with data registered in a database to identify suspicious individuals.

[1468] c. Alarms / Notifications

[1469] server

[1470] 1. If a suspicious individual is identified, summarize their characteristics and notify management and security organizations.

[1471] 2. Determine the timing and method of notification (immediate notification, periodic reporting, etc.) according to the system settings.

[1472] Terminal

[1473] 1. Receives a notification from the server and activates an alarm sound or warning light to alert those in the vicinity.

[1474] User response

[1475] User

[1476] 1. The administrator checks the notification on a smartphone or PC and checks the suspicious person's detailed information and video through the system application.

[1477] 2. If necessary, contact security organizations and take appropriate action.

[1478] Program processing explanation

[1479] server

[1480] Hardware: A server with a powerful GPU.

[1481] Software: Facial recognition model (OpenCV, Dlib), body shape recognition model (BodyPix), movement recognition model (OpenPose), speech recognition model (Google Speech-to-Text API), notification system (Firebase, Push notifications), cloud storage (AWS S3), real-time streaming (WebRTC).

[1482] Processing: Analyzes the received data, stores the features in a database, identifies suspicious individuals, generates notifications, and sends them to administrators and security companies.

[1483] Specific examples

[1484] For example, when installing this security system in a factory, the following specific example can be considered.

[1485] 1. Initial Setup

[1486] The administrator logs into the application and registers the face photos, height, weight, and walking video of all workers, as well as recording and registering each worker's voice.

[1487] 2. Daily operations

[1488] Cameras and microphones are installed at entrances and exits and important areas within the factory, allowing for 24-hour monitoring.

[1489] The monitoring terminal sends the collected data to a server, which analyzes the data and identifies workers and suspicious individuals.

[1490] 3. Identifying suspicious individuals

[1491] If the server identifies a suspicious person, it summarizes their characteristics (such as suspicious behavior or voice) and notifies the administrator's smartphone.

[1492] The administrator will review the notification and contact security or police if necessary.

[1493] Prompt Sentence Examples

[1494] Below are some examples of prompt sentences.

[1495] Implement a security system that analyzes video and audio data collected by cameras and microphones in a factory and compares it with the faces, body shapes, movements, and voices of pre-registered workers to identify suspicious individuals.Specific methods include using facial recognition, body shape recognition, movement recognition, and voice recognition models, and combining alarm and notification systems to build an application that provides managers and security organizations with information about suspicious individuals in real time.

[1496] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1497] Step 1:

[1498] The user logs into an application for administrators and uploads the worker's face photo, body shape data (height, weight), walking video, and voice data. A smartphone or PC is used as input to collect various data (images, videos, and audio files). This input data is then sent to the server.

[1499] Step 2:

[1500] The server receives the uploaded data and analyzes it using face recognition, body recognition, movement recognition, and voice recognition models. Specifically, it analyzes face data using a face recognition model (OpenCV, Dlib), body data using a body recognition model (BodyPix), movement data using a movement recognition model (OpenPose), and voice data using a voice recognition model (Google Speech-to-Text API). Features are extracted and stored in a database.

[1501] Step 3:

[1502] Terminals (monitoring terminals) are installed in the factory and monitor it 24 hours a day using cameras and microphones. The input data collected is real-time video and audio, and this data is sent to the server.

[1503] Step 4:

[1504] The server receives the video and audio data transmitted in real time and analyzes the collected data using each model. Specifically, the video and audio data is analyzed using a face recognition model, a body shape recognition model, a motion recognition model, and a voice recognition model to extract features. These features are then compared and verified with data registered in an existing database. As a result, the worker or suspicious individual is identified.

[1505] Step 5:

[1506] When the server identifies a suspicious individual, it summarizes their characteristics, including suspicious behavior and voice, and generates a warning notification based on this.

[1507] Step 6:

[1508] Server-generated alert notifications are sent to administrators and security organizations. The timing and method of notifications are determined by the system configuration, and can be immediate or periodic.

[1509] Step 7:

[1510] The device receives the notification from the server and activates the alarm device, specifically emitting an alarm sound or flashing a warning light.

[1511] Step 8:

[1512] The user (administrator) checks the notification on their smartphone or PC. The notification contains detailed information and video of the suspicious individual. The notification content is input data, and the administrator checks the situation as output. If necessary, they contact the security organization and take appropriate action.

[1513] Through the above processing steps, this system strengthens the security of the manufacturing site in real time, making it possible to immediately identify suspicious individuals and take appropriate measures.

[1514] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1515] The present invention relates to a security system that accurately distinguishes between family members and suspicious individuals by registering their faces, body shapes, and movements in advance and analyzing video and audio data collected using a camera and microphone. Furthermore, by combining it with an emotion engine, the system aims to recognize the user's emotions and automatically provide appropriate countermeasures. Specific embodiments are described in detail below.

[1516] 1. System Configuration

[1517] The system consists of the following components:

[1518] User device: A smartphone or PC operated by the user.

[1519] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[1520] Server: A central server that collects, analyzes, stores, and recognizes data.

[1521] 2. Initial Setup

[1522] User

[1523] 1. The user installs the security system application, creates a new account and logs in.

[1524] 2. Upload facial photos, body shapes, video of movements, and voice data of family members and related parties to the application and send them to the server.

[1525] server

[1526] 1. Create a database based on the received data and register the characteristics of family members and related parties.

[1527] 3. Daily operations

[1528] a. Data collection

[1529] Terminal

[1530] 1. The monitoring terminal (camera and microphone) monitors the monitored area 24 hours a day and collects video and audio in real time.

[1531] 2. Send the collected data to the server.

[1532] b. Data analysis and feature matching

[1533] server

[1534] 1. Analyze the received video data and extract facial features using a facial recognition algorithm.

[1535] 2. Analyze the received voice data and extract voice features using a voice recognition algorithm.

[1536] 3. The extracted features are compared with family and related person data in the database to identify suspicious individuals.

[1537] 4. Emotion recognition and countermeasures

[1538] a. Emotion recognition

[1539] server

[1540] 1. During the process of analyzing voice data, the emotion engine is used to determine the user's emotions.

[1541] 2. Emotion recognition algorithms analyze the tone, pitch, rhythm, and speed of the voice to determine whether the user is in an emotional state (e.g., fear or panic).

[1542] b. Emotion-based measures

[1543] server

[1544] 1. Adjust the intensity and method of warning sounds and lights based on emotion recognition results. For example, if the user is feeling fear, a stronger warning will be issued.

[1545] 2. Based on the emotion recognition results, the system proposes and implements appropriate countermeasures. For example, if the user feels fear, the system automatically notifies a security company.

[1546] 3. If necessary, notify the user's friends and family and ask for their assistance.

[1547] 5. User Response

[1548] User

[1549] 1. Check the notification on your smartphone or PC and check the suspicious person's details and recorded video through the system application.

[1550] 2. If necessary, contact a security company or the police and take appropriate action.

[1551] Program processing explanation

[1552] The processing of the program of the present invention will be specifically described below.

[1553] User Initial Settings

[1554] User

[1555] 1. The user logs in to the security system and uploads photos of the faces, body shapes, and movement videos of family members and related parties to the application.

[1556] 2. The user uses a microphone to record their own voice or that of their family members and sends it to the server.

[1557] server

[1558] 1. The server receives the uploaded data and analyzes it using facial recognition, body shape recognition, movement recognition, and voice recognition algorithms.

[1559] 2. Extract features and store them in a database.

[1560] Data collection and analysis

[1561] Terminal

[1562] 1. The monitoring terminal collects video and audio in real time and sends them to the server.

[1563] server

[1564] 1. The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[1565] 2. Analyze the voice data and extract voice features using a voice recognition algorithm.

[1566] 3. The server compares the features and identifies suspicious individuals.

[1567] Emotion recognition and countermeasures

[1568] server

[1569] 1. Use an emotion engine to recognize user emotions while analyzing voice data.

[1570] 2. Based on the emotion recognition results, set and implement appropriate warnings and notification methods.

[1571] 3. Automatically notify security companies or contacts for assistance if necessary.

[1572] Specific examples

[1573] For example, when installing this security system in a home, the following specific examples are conceivable.

[1574] 1. Initial Setup

[1575] A user logs into a security system application and configures the cameras and microphones in their home.

[1576] Register photos of each family member's face, height, body shape, and video of their movements. Also, record and register each person's voice.

[1577] 2. Daily operations

[1578] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[1579] The monitoring terminal sends the collected data to a server, which analyzes it.

[1580] 3. Identifying suspicious individuals

[1581] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[1582] The user checks the notification and confirms the details of the suspicious person.

[1583] 4. Emotion Recognition and Countermeasures

[1584] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[1585] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[1586] In this way, the crime prevention system of the present invention enables highly accurate and automatic crime prevention measures. Furthermore, the emotion recognition function allows for quick and appropriate responses even when the user feels fear, providing a safer and more secure living environment.

[1587] The processing flow will be explained below.

[1588] User Initial Settings

[1589] User

[1590] Step 1: The user installs the security system application on their smartphone or PC, creates a new account and logs in.

[1591] Step 2: The user takes and uploads photos of family members and related people, along with their height, body shape, and movement videos, to the application.

[1592] Step 3: The user uses the microphone to record their own voice or that of a family member and sends it to the server via the application.

[1593] server

[1594] Step 4: The server receives the uploaded face photo, height, body shape, motion video, and audio data.

[1595] Step 5: The server analyzes the received data and extracts features from each data using algorithms for face recognition, body shape recognition, movement recognition, and voice recognition.

[1596] Step 6: The server stores the extracted features in a database and creates profiles of family members and related parties.

[1597] Data Collection and Monitoring

[1598] Terminal

[1599] Step 1: The monitoring terminal (camera and microphone) monitors the monitored area 24 hours a day and collects video and audio in real time.

[1600] Step 2: The device sends the collected video and audio data to the server.

[1601] Data analysis and feature matching

[1602] server

[1603] Step 3: The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[1604] Step 4: The server analyzes the received voice data and extracts voice features using a voice recognition algorithm.

[1605] Step 5: The server compares the extracted features with family and related person data in the database and performs identification.

[1606] Emotion recognition and countermeasures

[1607] server

[1608] Step 6: The server uses the emotion engine while analyzing the voice data to determine the user's emotion by analyzing the tone, pitch, rhythm, speed, etc. of the voice.

[1609] Step 7: The server adjusts the intensity and method of the warning sound and light based on the emotion recognition result. For example, if the user feels fear, it will emit a strong warning sound.

[1610] Step 8: The server sends notifications to the user's friends and family to ask for help, if necessary.

[1611] Step 9: Based on the user's emotion recognition results, a function to automatically notify a security company is executed.

[1612] Identifying and alerting suspicious individuals

[1613] server

[1614] Step 10: If the server identifies a suspicious person, it summarizes their characteristics (gender, clothing, body shape, movement patterns, etc.).

[1615] Step 11: The server sends a notification to the user and the security company.

[1616] Step 12: Notifications are sent via message, email, in-app notification, or other methods according to the user's settings.

[1617] Terminal

[1618] Step 13: The device receives the notification from the server and issues an alert or turns on a light.

[1619] User response

[1620] User

[1621] Step 14: The user checks the notification on their smartphone or PC and checks the suspicious person's details and recorded video through the system application.

[1622] Step 15: If necessary, the user contacts a security company or the police and takes appropriate action.

[1623] As a concrete example, when installing this security system at home, the process is as follows:

[1624] 1. Initial Setup

[1625] The user logs in to the application and registers face photos, body shapes, motion videos, and voices of all family members.

[1626] 2. Daily operations

[1627] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[1628] The collected data is sent to a server and analyzed in real time.

[1629] 3. Identifying suspicious individuals

[1630] If the server identifies a suspicious person, it notifies the user of their characteristics.

[1631] The user checks the notification and confirms the details of the suspicious person.

[1632] 4. Emotion Recognition and Countermeasures

[1633] The server analyzes the voice characteristics and if the user is feeling fearful, it will emit a strong warning sound.

[1634] If necessary, notify a security company or acquaintances and request assistance.

[1635] As a result, the security system of the present invention eliminates the hassle of setup and operation and provides highly accurate, automatic security measures. It can also be easily used by elderly people, helping to create a safe and secure living environment.

[1636] Example 2

[1637] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1638] The main function of existing security systems is to identify suspicious individuals by registering the faces, body shapes, and movements of family members and related parties and analyzing the data collected by cameras and microphones. However, these systems lack the ability to recognize the user's emotional state and take appropriate action based on that. This makes it difficult to respond quickly and appropriately when the user becomes frightened or panics.

[1639] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for registering personal recognition data of family members, means for analyzing video and audio data collected by the image capture device and the audio capture device, means for comparing the analyzed data with registered family characteristics, means for emitting an alarm sound or light when a suspicious person is identified, means for summarizing the characteristics of the suspicious person and notifying the user and a security agency, means for analyzing the user's emotional state using voice recognition, and means for adjusting the method of warning and notification based on the emotional state. This enables a quick and appropriate response even when the user is feeling fear or panic.

[1640] "Family member personal recognition data" refers to data including facial photographs, body shapes, motion videos, and voice features of family members.

[1641] "Image capture device" refers to a camera or other video capture device that collects video data of a monitored area.

[1642] "Audio capture device" refers to a microphone or other audio capture device that collects audio data from a monitored area.

[1643] "Means of analysis" refers to algorithms and software for processing collected video and audio data and extracting meaningful features from it.

[1644] The "matching means" refers to the technology and method for comparing the analyzed features with pre-registered personal recognition data of family members and determining the degree of match.

[1645] "Means for emitting warning sounds or lights" refers to a device or system that emits sounds or lights to warn a user when a suspicious person is identified.

[1646] The "means for summarizing characteristics and notifying the user and security agency" refers to the technology and method for summarizing the identification information of a suspicious person and transmitting that information to the user's terminal and the security agency.

[1647] "Means for analyzing a user's emotional state using voice recognition" refers to algorithms or software that analyze a user's voice data to determine their emotional state (e.g., fear, anger, calm, etc.).

[1648] "Means for adjusting warning and notification methods based on emotional state" refers to techniques and methods for changing the intensity and method of warning sounds and notifications depending on the analyzed emotional state of the user.

[1649] This invention relates to a crime prevention system that can accurately distinguish between family members and suspicious individuals by registering face, body shape, and movement data of family members in advance and analyzing video and audio data collected using an image capture device and an audio capture device. Furthermore, by combining it with an emotion engine, the system aims to recognize the user's emotional state and automatically provide appropriate countermeasures.

[1650] System Configuration

[1651] The system consists of the following components:

[1652] User device: A smartphone or PC operated by the user.

[1653] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[1654] Server: A central server that collects, analyzes, stores, and recognizes data.

[1655] Initial Setup

[1656] User

[1657] 1. The user installs the security system application on their smartphone or PC, creates a new account and logs in.

[1658] 2. Within the application, take or select and upload photos of your family members' faces, body shapes and movement videos. Also, use the microphone to record the voices of all family members and send them to the server.

[1659] server

[1660] 1. The server analyzes the received facial photos, body shape, motion video, and audio data and extracts each feature. TensorFlow is used for facial recognition, Google Cloud Speech-to-Text API for voice recognition, and OpenCV for body shape and motion recognition.

[1661] 2. Save the extracted features in a database (e.g., MySQL).

[1662] Daily operations

[1663] Terminal

[1664] 1. Surveillance devices (e.g. security cameras and microphones) monitor the area 24 hours a day.

[1665] 2. The monitoring device sends the collected video and audio data to the server in real time using AWS Lambda.

[1666] server

[1667] 1. The server analyzes the received video data using a facial recognition algorithm (e.g., DeepFace) and extracts facial features.

[1668] 2. The voice data is analyzed using a voice recognition algorithm (e.g., DeepSpeech) to extract voice features.

[1669] 3. The features are matched with existing data in the database to identify suspicious individuals. This matching is performed using SQL queries.

[1670] Emotion recognition and countermeasures

[1671] server

[1672] 1. The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions based on the voice data collected in real time.

[1673] 2. Emotion recognition algorithms are used to analyze the tone, pitch, rhythm, and speed of the voice to determine the user's emotional state (e.g., fear or panic).

[1674] 3. Depending on the emotional state, take action, such as "set alarm sound to maximum level" or "increase lighting."

[1675] 4. If necessary, use the Twilio API to notify security companies or friends.

[1676] Specific examples

[1677] For example, when installing this security system in a home, the following specific examples are possible:

[1678] 1. Initial Setup

[1679] A user logs into a security system application (e.g., FamilySafe) and configures their home camera and microphone.

[1680] Register photos of each family member's face, body shape, and movement videos, and also record and register each person's voice.

[1681] 2. Daily operations

[1682] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[1683] The monitoring terminal sends the collected data to a server, which analyzes it.

[1684] 3. Identifying suspicious individuals

[1685] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[1686] The user checks the notification and checks the suspicious person's details and recorded video.

[1687] 4. Emotion Recognition and Countermeasures

[1688] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[1689] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[1690] Prompt Sentence Examples

[1691] Here are some example prompts to input to a generative AI model:

[1692] "Please explain a system that identifies suspicious faces and voices and analyzes user emotions from video and audio data collected by home surveillance cameras."

[1693] The crime prevention system of this invention enables highly accurate and automatic crime prevention measures. Furthermore, the emotion recognition function allows for quick and appropriate responses even when the user feels fear, providing a safer and more secure living environment.

[1694] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1695] Step 1: User Data Entry

[1696] User

[1697] Users install the security system application on their smartphone or PC, create a new account, and log in. Within the application, they take or select and upload photos of their family members' faces, body shapes, and movement videos. They also use a microphone to record the voices of all family members and send them to the server.

[1698] Input: Face photo, body shape, motion video, audio recording

[1699] Output: Personal identification data sent to the server

[1700] Step 2: Data analysis and registration by the server

[1701] server

[1702] The server analyzes the received facial photos, body shape, motion video, and audio data, extracting each feature. This process uses TensorFlow and the Google Cloud Speech-to-Text API, and image data is analyzed using OpenCV. For example, it performs tasks such as "extracting facial feature points," "measuring height," and "identifying motion patterns." The extracted features are then stored in a database (e.g., MySQL).

[1703] Input: Personal identification data submitted by the user

[1704] Output: Analyzed feature data registered in the database

[1705] Step 3: Collect data from the device

[1706] Terminal

[1707] The monitoring device (e.g., security camera and microphone) monitors the area 24 hours a day. The monitoring device transmits the collected video and audio data to the server in real time. This is done using AWS Lambda. For example, the camera frame rate is set to 30 fps (30 frames per second), and the microphone is highly sensitive.

[1708] Input: Video and audio data of the monitored area

[1709] Output: Video and audio data sent to the server in real time

[1710] Step 4: Data analysis and collation by the server

[1711] server

[1712] The server analyzes the received video data using a facial recognition algorithm (e.g., DeepFace) to extract facial features. It also analyzes the audio data using a speech recognition algorithm (e.g., DeepSpeech) to extract voice features. These features are then compared with existing data in a database to identify suspicious individuals. For example, it performs tasks such as "calculating the degree of match between facial features and those in the database" and "analyzing the tone and pitch of the voice." Matching is performed using SQL queries.

[1713] Input: Real-time video and audio data from a monitoring terminal

[1714] Output: Identified suspicious person information

[1715] Step 5: Emotion recognition by the server

[1716] server

[1717] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions from the voice data collected in real time. It uses an emotion recognition algorithm to analyze the tone, pitch, rhythm, and speed of the voice to determine the user's emotional state (e.g., fear or panic). For example, it performs "tone analysis of voice data" and "pitch fluctuation analysis."

[1718] Input: Audio data from the monitored area

[1719] Output: Parsed user's emotional state

[1720] Step 6: Server takes action

[1721] server

[1722] Based on the emotion recognition results, the intensity and method of warning sounds and lighting are adjusted. For example, if it detects fear in the user, it can set the warning sound to maximum level or turn on a powerful flashlight. It also uses the Twilio API to send notifications to security companies and acquaintances. For example, if the user feels panicked, it can automatically notify the security company or send an SMS to an acquaintance.

[1723] Input: Analyzed user's emotional state, suspicious person information

[1724] Output: Adjusted alert and notification methods, notifications sent

[1725] (Application example 2)

[1726] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1727] In recent years, there has been a demand for improved home security, but existing security systems are limited to detecting suspicious individuals and are not able to adequately address the fear and anxiety felt by users. Another issue is that when a suspicious individual is detected, the system does not respond quickly, and it takes time to ensure the user's peace of mind. Furthermore, users are unable to flexibly configure the system to suit their needs, which is inconvenient.

[1728] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1729] In this invention, the server includes means for registering the faces, body shapes, and movements of family members, means for analyzing video and audio data collected by the monitoring terminal, means for comparing the analyzed data with the registered family members' characteristics, means for emitting an alarm sound or light when a suspicious person is identified, means for summarizing the suspicious person's characteristics and notifying the user and the security company, means for recognizing the user's emotions from the voice, and means for issuing a strong alarm or automatically reporting to the security company when the user feels fear. This enables quick detection and response of suspicious people as well as security measures that take the user's emotions into consideration.

[1730] "Means for registering family members' faces, body shapes, and movements" refers to a device or method for registering and saving photographs of the faces, body shapes, and movements of family members and related parties as digital data.

[1731] "Means for analyzing video and audio data collected by a camera and a microphone" refers to a device or method that analyzes video and audio data collected by a camera or a microphone and extracts specific features from them.

[1732] "Means for matching analyzed data with registered family characteristics" refers to a device or method that compares the data obtained by analysis with pre-registered family data to determine whether there is a match.

[1733] "Means for emitting a warning sound or light when a suspicious person is identified" refers to a device or method that issues a warning by sounding a warning sound or flashing a light when the system identifies a suspicious person.

[1734] "Means for summarizing the characteristics of a suspicious person and notifying the user and security company" refers to a device or method that, when a suspicious person is identified, summarizes the characteristics and sends a notification to the user's device or security company.

[1735] The "means for recognizing a user's emotion from voice" refers to a device or method for analyzing a user's voice data and identifying the emotion.

[1736] "Means for issuing a strong warning or automatically notifying a security company if the user feels fear" refers to a device or method that analyzes the emotions from the user's voice and, if it determines that the user feels fear, issues a strong warning sound and light or automatically notifies a security company.

[1737] As an embodiment of the present invention, a home security system will be described in detail. This security system ensures safety and security at home by detecting suspicious individuals and responding quickly, as well as providing appropriate measures based on the user's emotions. The specific configuration and operation of the system will be described below.

[1738] 1. System Configuration

[1739] The system consists of the following components:

[1740] User device: A smartphone or PC operated by the user.

[1741] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[1742] Server: A central server that collects, analyzes, stores, and recognizes emotions.

[1743] 2. Initial Setup

[1744] User terminal

[1745] 1. The user installs the security system application, creates a new account and logs in.

[1746] 2. Upload facial photos, body shapes, video of movements, and voice data of family members and related parties to the application and send them to the server.

[1747] server

[1748] 1. Create a database based on the received data and register the characteristics of family members and related parties.

[1749] 2. The registered feature data is analyzed using a recognition algorithm and used for subsequent matching.

[1750] 3. Daily operations

[1751] a. Data collection

[1752] Monitoring terminal

[1753] 1. The monitoring terminal monitors the area 24 hours a day and collects video and audio in real time.

[1754] 2. Send the collected data to the server.

[1755] b. Data analysis and feature matching

[1756] server

[1757] 1. Analyze the received video data using "OpenCV" and the "face_recognition" library to extract facial features.

[1758] 2. The received voice data is analyzed using voice recognition software to extract voice features.

[1759] 3. The extracted features are compared with family and related person data in the database to identify suspicious individuals.

[1760] 4. Emotion recognition and countermeasures

[1761] a. Emotion recognition

[1762] server

[1763] 1. During the process of analyzing the voice data, the "EmotionRecognizer" emotion engine is used to determine the user's emotions.

[1764] 2. Emotion recognition algorithms analyze the tone, pitch, rhythm, and speed of the voice to determine whether the user is in an emotional state (e.g., fear or panic).

[1765] b. Emotion-based measures

[1766] server

[1767] 1. Adjust the intensity and method of warning sounds and lights based on emotion recognition results. For example, if the user is feeling fear, a stronger warning will be issued.

[1768] 2. Based on the emotion recognition results, appropriate countermeasures are proposed and implemented. For example, if the user feels fear, a security company is automatically notified.

[1769] 3. If necessary, notify the user's friends and family and ask for their assistance.

[1770] Specific examples

[1771] For example, when installing this security system in a home, the following specific examples are conceivable.

[1772] 1. Initial Setup

[1773] A user logs into a security system application and configures the cameras and microphones in their home.

[1774] Register photos of each family member's face, height, body shape, and video of their movements. Also, record and register each person's voice.

[1775] 2. Daily operations

[1776] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[1777] The monitoring terminal sends the collected data to a server, which analyzes it.

[1778] 3. Identifying suspicious individuals

[1779] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[1780] The user checks the notification and confirms the details of the suspicious person.

[1781] 4. Emotion Recognition and Countermeasures

[1782] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[1783] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[1784] Specific prompt examples

[1785] "Design a smart home security system that automatically alerts a security company if it detects fear or anxiety in the user. It uses cameras and microphones to collect data in real time, performs facial and emotion recognition, and notifies the security company in a specified manner. Please provide example code for implementation."

[1786] In this way, by implementing the security system of the present invention, it becomes possible to detect suspicious individuals quickly and with high accuracy and to take measures based on the user's emotions, thereby effectively improving the safety and security of the home.

[1787] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1788] Step 1:

[1789] Initial Setup

[1790] Input: The user uploads facial photos, body shapes, movement videos, and voice data of family members and related parties to the security system app.

[1791] Data processing / calculation: The server analyzes the received data and extracts features using algorithms for face recognition, body shape recognition, movement recognition, and voice recognition.

[1792] Output: Save these features in a database.

[1793] Specific operation: The user uploads family information using a smartphone application. The server receives the data, extracts features using various recognition algorithms, and saves them.

[1794] Step 2:

[1795] Real-time data collection

[1796] Input: A monitoring terminal collects video and audio data in real time.

[1797] Data processing / calculation: Video data is captured using a camera, and audio data is collected using a microphone. The collected data is sent to a server.

[1798] Output: The server receives the video and audio data.

[1799] Specific operation: The surveillance cameras and microphones are always on, and the collected data is periodically uploaded to the server.

[1800] Step 3:

[1801] Data analysis and collation

[1802] Input: Video and audio data received by the server from the surveillance terminal.

[1803] Data processing / calculation: The server analyzes the facial features of the received video data using "OpenCV" and the "face_recognition" library. The audio data is analyzed using voice recognition software to extract voice features.

[1804] Output: The features are matched with family data in the database to identify suspicious individuals.

[1805] How it works: The server analyzes the data in real time and compares it with registered family members and associates. If the data does not match, the person is identified as a suspicious person.

[1806] Step 4:

[1807] emotion recognition

[1808] Input: The server analyzes the audio data in real time and recognizes the user's voice.

[1809] Data processing / calculation: Use the "EmotionRecognizer" emotion engine to extract emotional data (tone, pitch, rhythm, speed, etc.) from audio data.

[1810] Output: Obtain the emotion recognition result and determine the type of emotion (e.g., fear or panic).

[1811] Specific operation: The server analyzes the user's voice using an emotion engine to determine the user's emotional state.

[1812] Step 5:

[1813] Measures implemented

[1814] Input: Emotion recognition results and suspicious person identification results.

[1815] Data processing / calculation: Based on the emotion recognition results, the intensity of the warning sound and light is adjusted. If the user feels fear, the system will automatically notify a security company.

[1816] Output: Emit appropriate alarm sounds and lights, and notify the security company.

[1817] Specific actions: The server determines the action to take based on the emotion recognition results, and if necessary, notifies a security company. For example, if the user feels fear, the server will emit a loud alarm and automatically notify a security company.

[1818] Step 6:

[1819] Notification and confirmation

[1820] Input: Suspicious person identification results and response results.

[1821] Data processing / calculation: The server summarizes the characteristics of the suspicious individual and the results of countermeasures, and notifies the user and the security company.

[1822] Output: Send a notification to your smartphone or security company's device.

[1823] How it works: Users receive a notification on their smartphone and can check the details of the suspicious individual. Security companies also receive a notification so they can take appropriate action.

[1824] The above are the specific processing steps in the embodiment of the invention, which enable quick and highly accurate detection of suspicious individuals and flexible countermeasures based on the user's emotions.

[1825] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1826] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1827] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1828] [Fourth embodiment]

[1829] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1830] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1831] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1832] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1833] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1834] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1835] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1836] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1837] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1838] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1839] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1840] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1841] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1842] The present invention provides a security system that can accurately distinguish between family members and suspicious individuals by registering the faces, body shapes, and movements of family members in advance and analyzing video and audio data collected by a camera and microphone. Specific embodiments are described below.

[1843] 1. System Configuration

[1844] The system mainly includes the following components:

[1845] User device: A smartphone or PC operated by the user.

[1846] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[1847] Server: A central management server responsible for collecting, analyzing, and storing data.

[1848] 2. Initial Setup

[1849] User

[1850] Create an account as a new user and log in to the system.

[1851] Photos of family members and related parties' faces, body shapes, video of their movements, and voice data are uploaded and sent to the server.

[1852] server

[1853] A database is constructed based on the received data, and the characteristics of family members and related parties are registered.

[1854] 3. Daily operations

[1855] a. Data collection

[1856] Terminal

[1857] The camera and microphone collect video and audio in real time and send it to the server.

[1858] b. Data analysis and feature matching

[1859] server

[1860] The collected video and audio data is analyzed in real time to extract features of family members and suspicious individuals.

[1861] The extracted features are compared with data registered in a database to identify suspicious individuals.

[1862] c. Alarms / Notifications

[1863] server

[1864] If a suspicious person is identified, their characteristics are summarized and notified to the user and the security company.

[1865] The timing and method of notification (immediate notification, periodic reporting, etc.) will be determined according to the system settings.

[1866] Terminal

[1867] It receives notifications from the server and alerts those around it by sounding an alarm or flashing a warning light.

[1868] 4. User Response

[1869] (User)

[1870] Check the notification on your smartphone or PC and check detailed information and video of the suspicious person through the system application.

[1871] If necessary, contact the police or a security company and take appropriate action.

[1872] Program processing explanation

[1873] The program processing for each section will be explained below.

[1874] User Initial Settings

[1875] User

[1876] 1. The user logs in to the security system and uploads photos of the faces, body shapes, and movement videos of family members and other people involved to the application.

[1877] 2. The user records their own or their family's voice into the microphone and sends it to the server.

[1878] server

[1879] 1. The server receives the uploaded data and analyzes it using face recognition models, body shape recognition models, motion recognition models, and voice recognition models.

[1880] 2. Extract features and store them in a database.

[1881] Monitoring and Data Collection

[1882] Terminal

[1883] 1. The monitoring terminal (camera and microphone) monitors the installation location 24 hours a day, collecting video and audio in real time.

[1884] 2. Send the collected data to the server.

[1885] server

[1886] 1. The server analyzes the received data and matches it with the characteristics of registered family members and related parties.

[1887] Identifying and alerting suspicious individuals

[1888] server

[1889] 1. Identify suspicious individuals based on the results of feature matching.

[1890] 2. Summarize the characteristics of the suspicious individual and generate notification content.

[1891] 3. Send notifications to users and security companies.

[1892] Terminal

[1893] 1. Receive notifications from the server and emit a warning sound or flash a warning light.

[1894] User response

[1895] User

[1896] 1. Receive a notification on your smartphone or PC about a suspicious person and check the details.

[1897] 2. If necessary, contact the police or a security company and take appropriate action.

[1898] Specific examples

[1899] For example, when installing this security system in a home, the following specific examples are conceivable.

[1900] 1. Initial Setup

[1901] The user logs in to the application and registers the face photos, height, weight, and walking video of each family member, as well as recording and registering each person's voice.

[1902] 2. Daily operations

[1903] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[1904] The monitoring terminal sends the collected data to a server, which analyzes the data and distinguishes between family members and suspicious individuals.

[1905] 3. Identifying suspicious individuals

[1906] If the server identifies a suspicious person, it summarizes their characteristics (e.g., suspicious behavior or voice) and notifies the user's smartphone.

[1907] The user will check the notification and contact the police if necessary.

[1908] In this way, by using the security system of the present invention, even elderly people can easily operate it, and highly accurate security measures can be realized. Furthermore, in order to improve the accuracy of identifying suspicious individuals, an even safer environment can be provided by including voice data in the analysis.

[1909] The processing flow will be explained below.

[1910] User Initial Settings

[1911] User

[1912] Step 1: The user installs the security system application, creates a new account and logs in.

[1913] Step 2: The user uploads photos of their family and friends, along with their height, body type, and movement videos within the application.

[1914] Step 3: The user uses a microphone to record their own voice or that of their family members and sends it to the server.

[1915] server

[1916] Step 4: The server receives the uploaded face photo, height, body shape, motion video, and audio data.

[1917] Step 5: The server analyzes the received data and extracts features from each data using a face recognition model, a body shape recognition model, a motion recognition model, and a voice recognition model.

[1918] Step 6: Store the features in a database to create profiles of family members and associates.

[1919] Monitoring and Data Collection

[1920] Terminal

[1921] Step 1: The monitoring terminal (camera and microphone) monitors the target area 24 hours a day and collects video and audio in real time.

[1922] Step 2: The device sends the collected video and audio data to the server.

[1923] Data analysis and feature matching

[1924] server

[1925] Step 3: The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[1926] Step 4: The server analyzes the received voice data and extracts voice features using a voice recognition algorithm.

[1927] Step 5: The server matches the extracted features with family and related person data in the database.

[1928] Step 6: The server identifies the suspicious person based on the result of the feature matching.

[1929] Identifying and alerting suspicious individuals

[1930] server

[1931] Step 7: If the server identifies a suspicious person, it summarizes their characteristics (e.g., gender, clothing, body shape, movement patterns, etc.).

[1932] Step 8: The server sends a notification to the user and the security company.

[1933] Step 9: Notifications are sent via message, email, in-app notification, or other methods according to the user's settings.

[1934] Terminal

[1935] Step 10: The device receives the notification from the server and plays an alert or turns on a light.

[1936] User response

[1937] User

[1938] Step 11: The user checks the notification on their smartphone or PC.

[1939] Step 12: The user opens a system application to view the suspicious individual's details and video recording.

[1940] Step 13: If necessary, the user contacts the police or a security company and takes appropriate action.

[1941] This system eliminates the need for setup and operation, and provides highly accurate, automatic crime prevention measures. It is also easy to operate, especially for the elderly, and can provide a safe and secure living environment.

[1942] Example 1

[1943] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1944] Modern security systems often lack the technology to accurately distinguish between family members and related parties and suspicious individuals. This can result in frequent false alarms, while genuine suspicious individuals may be overlooked. Real-time monitoring and notification functions are also inadequate, making it difficult to respond quickly. Furthermore, the inability to analyze multidimensional data, including audio data, reduces the accuracy of security measures. To address these issues, a system is needed that can identify suspicious individuals in real time based on multiple characteristics, such as face, body shape, movement, and voice, and issue appropriate alarms.

[1945] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1946] In this invention, the server includes a means for registering the faces, body shapes, and movements of family members, a means for analyzing the video and audio data collected by the camera and microphone, and a means for comparing the analyzed data with the registered family characteristics, thereby enabling accurate real-time identification of family members, related parties, and suspicious individuals.

[1947] A "family" is a group of people who are related by blood or who live together and are registered in the system.

[1948] "Body shape" refers to the physical appearance characteristics of a user or related person, such as height, weight, and body fat percentage.

[1949] A "motion" is a movement pattern that indicates a specific action or gesture of a person and is registered and recognized by the system.

[1950] "Camera" refers to a photographic device for collecting video data.

[1951] "Microphone" refers to a recording device for collecting audio data.

[1952] "Video data" is digital data containing visual information collected by a camera.

[1953] "Audio data" is digital data containing acoustic information collected by a microphone.

[1954] "Analysis" is the process of examining collected video and audio data using algorithms and models to extract features.

[1955] "Features" are identification information such as facial feature points, body shape data, movement patterns, and voice frequency components.

[1956] A "suspicious person" is a person who is not registered in the system and who the system determines to be suspicious based on an event.

[1957] A "warning sound" is a sound that is emitted when a suspicious person is identified, and is an acoustic signal that alerts those in the vicinity.

[1958] "Light" is the visual warning signal emitted by the warning device.

[1959] "Summarization" is the process of extracting important parts from analyzed data and summarizing them concisely.

[1960] "Notification" is a message that notifies the user or the security company of information about a suspicious person.

[1961] "Push notification" is a real-time notification method that the system proactively sends to the user's device.

[1962] A "monitoring terminal" is a device that collects data and transmits it to a server, and includes a camera and a microphone.

[1963] "Real-time" refers to the near-simultaneous process of collecting and analyzing data.

[1964] The present invention is a system that highly strengthens security functions, and by registering the faces, body shapes, movements, and voices of family members and related parties in advance and analyzing the video and audio data collected by the camera and microphone in real time, it can reliably distinguish between family members and suspicious individuals. Specific embodiments for implementing the present invention will be described below.

[1965] 1. System Configuration

[1966] The system includes the following components:

[1967] User device: A smartphone or PC operated by the user.

[1968] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[1969] Server: A central management server responsible for collecting, analyzing, and storing data.

[1970] 2. Initial Setup

[1971] User

[1972] The user launches the security system application, creates a new account, and logs in to the system. Face photos, body shape information, movement videos, and voice data of family members and related parties are uploaded through the application and sent to the server. Specifically, the user takes a face photo using the smartphone camera and taps the upload button. Body shape information is entered into the text field, and movement videos are taken with the smartphone camera and uploaded by tapping the movement registration button. Audio data is recorded by tapping the record button within the app, and once recording is complete, the user taps the save button to send the data to the server.

[1973] 3. Daily operations

[1974] Monitoring terminal

[1975] The monitoring terminal (camera and microphone) monitors the installation location 24 hours a day, collecting video and audio data in real time. The collected video and audio data is then sent to the server in real time.

[1976] server

[1977] The server receives real-time video and audio data sent from the monitoring terminal. Generative AI models such as face recognition, body shape recognition, movement recognition, and voice recognition are used for analysis. These models are used to analyze the collected data and extract characteristic information. This characteristic information is then compared with the characteristics of family members and related parties registered in a database to identify suspicious individuals.

[1978] 4. Identifying suspicious individuals and issuing alerts

[1979] server

[1980] The server compares the analyzed data in real time with the characteristic data stored in the database. Based on the comparison results, if a person not registered in the system or suspicious behavior or voice patterns are detected, the person is identified as a suspicious person. If a suspicious person is identified, the characteristics, time of detection, and location are summarized and a notification is generated. This notification is sent to the user's smartphone as a push notification.

[1981] Terminal

[1982] The device receives a notification from the server and alerts those around it by sounding an alarm or flashing a warning light.

[1983] 5. User Response

[1984] User

[1985] Users receive a notification on their smartphone or PC that a suspicious person has been identified and can check the details of the suspicious person. Through the application, users can check the suspicious person's video and characteristic information, and if necessary, contact the police or security company and take appropriate action. Specifically, they can call the police by pressing the emergency contact button within the app.

[1986] Prompt Sentence Examples

[1987] Here is an example of a prompt for a generative AI model:

[1988] "Explain a security system that can reliably distinguish between family members and suspicious individuals by registering their faces, body shapes, and movements in advance and analyzing the video and audio data collected by the camera and microphone. Describe in detail the components of this system, the initial setup procedure, the data collection and analysis methods, the process for identifying suspicious individuals, and the notifications received by the user. Also, explain with concrete examples."

[1989] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1990] Step 1:

[1991] Data registration by users

[1992] Input: Smartphone or PC operated by the user, facial photo, body shape data, movement video, and voice data.

[1993] Specific operation: The user launches the security system app, creates a new account, and logs in. They then use their smartphone camera to take photos of their family members or other people involved and tap the upload button. Next, they enter their body shape data, such as height and weight, into the text fields. They then take a video of their movements with the smartphone camera and press the action registration button to upload it. Finally, they use the microphone to record the voices of their family members or other people involved, and once the recording is complete, they press the save button to send it to the server.

[1994] Output: Facial photo data, body shape data, motion video data, and voice data are sent to the server.

[1995] Step 2:

[1996] Data analysis by server

[1997] Input: Face photo, body shape data, motion video, and audio data sent by the user.

[1998] Specific operation: The server receives the uploaded data and begins analysis. A facial recognition model is used to extract feature points from the facial photo, and a body shape recognition model is used to analyze height and weight data. The motion video is analyzed using a motion recognition model to identify unique motion patterns. A voice recognition model is used to analyze the voice data and extract voice features.

[1999] Data processing and data calculation: Analyze each data using a generative AI model and extract feature information.

[2000] Output: The extracted feature information is registered in a database.

[2001] Step 3:

[2002] Real-time data collection via terminal

[2003] Input: Camera and microphone installed on the monitoring terminal.

[2004] How it works: The monitoring device collects video and audio 24 hours a day and converts them into the appropriate format. For example, a camera installed at the entrance records continuously, and a microphone collects the surrounding audio. This data is then sent to the server in real time.

[2005] Output: The video and audio data collected in real time is sent to the server.

[2006] Step 4:

[2007] Real-time analysis and verification by the server

[2008] Input: Real-time video and audio data transmitted from the monitoring terminal.

[2009] Specific operation: The server analyzes the received data in real time. The collected video and audio data is analyzed using a generative AI model to extract feature information.

[2010] Data processing and calculation: The characteristic information obtained through real-time analysis is compared with the characteristics of family members and related parties registered in advance.

[2011] Output: Matching results are generated.

[2012] Step 5:

[2013] Identifying suspicious individuals by the server

[2014] Input: Real-time analysis and matching results.

[2015] Specific operation: Based on the matching results, the server identifies unregistered individuals and suspicious behavior and voices. It summarizes the characteristics of the suspicious individual and the data at the time of detection, and generates a notification.

[2016] Data processing and data calculation: Based on the suspicious person detection algorithm, the characteristics of suspicious persons are stored in a database and summary information is generated.

[2017] Output: A summary of the suspicious person notification is generated.

[2018] Step 6:

[2019] Server sends notifications

[2020] Input: Summary information of the suspect.

[2021] Specific operation: The server generates a notification based on the summary information and sends it to the user or security company as a push notification. The notification includes the specific time, location, and characteristics of the suspicious person.

[2022] Output: A suspicious person notification is sent to the user's smartphone and the security company.

[2023] Step 7:

[2024] Alarm operation by terminal

[2025] Input: Suspicious person notification from the server.

[2026] Specific operation: The device receives a notification from the server and alerts those around it by sounding an alarm or flashing a warning light.

[2027] Output: An alarm is sounded to the surrounding area.

[2028] Step 8:

[2029] User confirmation and response

[2030] Input: Suspicious person notification sent to the user's smartphone or PC.

[2031] Specific operation: The user receives a notification and checks the details of the suspicious person through the application, checks the video and characteristics of the suspicious person, and if necessary, presses the emergency contact button to contact the police or a security company.

[2032] Output: The user requests the police or a security company to take appropriate action.

[2033] (Application example 1)

[2034] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2035] Strengthening safety management and security is an essential issue in modern manufacturing sites. However, it is difficult to completely prevent the intrusion of suspicious individuals into a work environment, and advanced identification technology is required. Furthermore, conventional security systems have difficulty in accurately identifying individuals using motion and voice data, and are unable to cope with complex work environments. Therefore, there is a need for an effective security system specialized for manufacturing sites.

[2036] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2037] In this invention, the server includes means for registering the faces, body shapes, and movements of workers at the manufacturing site, means for analyzing video and audio data collected by the camera and microphone, means for comparing the analyzed data with the characteristics of registered workers, means for emitting an alarm sound or light when a suspicious person is identified, and means for summarizing the characteristics of the suspicious person and notifying the manager and security organization. This makes it possible to strengthen security at the manufacturing site in real time, immediately identify suspicious people, and take appropriate measures.

[2038] A "manufacturing site" is a place where products are produced or assembled, and generally refers to a factory or workshop.

[2039] A "worker" is an individual who performs a specific task or work, and in a manufacturing site, refers to an employee engaged in production or assembly work.

[2040] "Facial recognition" is the process of identifying an individual's face from images captured by a camera and comparing it with pre-registered facial data.

[2041] "Body shape recognition" is the process of analyzing an individual's body shape and posture from video data and comparing it with pre-registered data.

[2042] "Motion recognition" is a technology that analyzes video data acquired from a camera to identify and distinguish individual movements and behaviors.

[2043] "Voice recognition" is a technology that analyzes voice data acquired by a microphone and identifies an individual's voice.

[2044] A "suspicious person" refers to a person who is not registered at the manufacturing site or who behaves inappropriately.

[2045] An "alarm sound" is an audio signal emitted by the system to alert the user to the intrusion of a suspicious person or an abnormality.

[2046] The "light" refers to the optical signal emitted by the system to alert the user to the intrusion of a suspicious person or an abnormality.

[2047] "Signature summarization" is the process of generating a concise summary of a person's appearance and behavioral characteristics when they are identified as a suspicious person.

[2048] "Manager" refers to the person in charge of the operation and safety management of the manufacturing site.

[2049] The "security organization" is a specialized organization responsible for security at the manufacturing site and dealing with suspicious individuals.

[2050] The present invention relates to a system for enhancing security at a manufacturing site. Specific embodiments of the present invention will be described below.

[2051] System Configuration

[2052] The system mainly includes the following components:

[2053] User device: A smartphone or PC operated by an administrator.

[2054] Surveillance terminal: A surveillance device equipped with a camera and microphone.

[2055] Server: A central management server responsible for collecting, analyzing, and storing data.

[2056] Initial Setup

[2057] User

[2058] 1. The administrator logs in to the system and uploads the worker's face photo, body shape, movement video, and voice data.

[2059] 2. The administrator records each person's voice and sends it to the server.

[2060] server

[2061] 1. The server receives the uploaded data and analyzes it using face recognition, body shape recognition, movement recognition, and voice recognition models.

[2062] 2. Extract features and store them in a database.

[2063] Daily operations

[2064] a. Data collection

[2065] Terminal

[2066] 1. Monitoring terminals (cameras and microphones) installed within the factory monitor the plant 24 hours a day, collecting video and audio in real time and sending it to a server.

[2067] b. Data analysis and feature matching

[2068] server

[2069] 1. Analyze collected video and audio data in real time to extract features of workers and suspicious individuals.

[2070] 2. The extracted features are compared with data registered in a database to identify suspicious individuals.

[2071] c. Alarms / Notifications

[2072] server

[2073] 1. If a suspicious individual is identified, summarize their characteristics and notify management and security organizations.

[2074] 2. Determine the timing and method of notification (immediate notification, periodic reporting, etc.) according to the system settings.

[2075] Terminal

[2076] 1. Receives a notification from the server and activates an alarm sound or warning light to alert those in the vicinity.

[2077] User response

[2078] User

[2079] 1. The administrator checks the notification on a smartphone or PC and checks the suspicious person's detailed information and video through the system application.

[2080] 2. If necessary, contact security organizations and take appropriate action.

[2081] Program processing explanation

[2082] server

[2083] Hardware: A server with a powerful GPU.

[2084] Software: Facial recognition model (OpenCV, Dlib), body shape recognition model (BodyPix), movement recognition model (OpenPose), speech recognition model (Google Speech-to-Text API), notification system (Firebase, Push notifications), cloud storage (AWS S3), real-time streaming (WebRTC).

[2085] Processing: Analyzes the received data, stores the features in a database, identifies suspicious individuals, generates notifications, and sends them to administrators and security companies.

[2086] Specific examples

[2087] For example, when installing this security system in a factory, the following specific example can be considered.

[2088] 1. Initial Setup

[2089] The administrator logs into the application and registers the face photos, height, weight, and walking video of all workers, as well as recording and registering each worker's voice.

[2090] 2. Daily operations

[2091] Cameras and microphones are installed at entrances and exits and important areas within the factory, allowing for 24-hour monitoring.

[2092] The monitoring terminal sends the collected data to a server, which analyzes the data and identifies workers and suspicious individuals.

[2093] 3. Identifying suspicious individuals

[2094] If the server identifies a suspicious person, it summarizes their characteristics (such as suspicious behavior or voice) and notifies the administrator's smartphone.

[2095] The administrator will review the notification and contact security or police if necessary.

[2096] Prompt Sentence Examples

[2097] Below are some examples of prompt sentences.

[2098] Implement a security system that analyzes video and audio data collected by cameras and microphones in a factory and compares it with the faces, body shapes, movements, and voices of pre-registered workers to identify suspicious individuals.Specific methods include using facial recognition, body shape recognition, movement recognition, and voice recognition models, and combining alarm and notification systems to build an application that provides managers and security organizations with information about suspicious individuals in real time.

[2099] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2100] Step 1:

[2101] The user logs into an application for administrators and uploads the worker's face photo, body shape data (height, weight), walking video, and voice data. A smartphone or PC is used as input to collect various data (images, videos, and audio files). This input data is then sent to the server.

[2102] Step 2:

[2103] The server receives the uploaded data and analyzes it using face recognition, body recognition, movement recognition, and voice recognition models. Specifically, it analyzes face data using a face recognition model (OpenCV, Dlib), body data using a body recognition model (BodyPix), movement data using a movement recognition model (OpenPose), and voice data using a voice recognition model (Google Speech-to-Text API). Features are extracted and stored in a database.

[2104] Step 3:

[2105] Terminals (monitoring terminals) are installed in the factory and monitor it 24 hours a day using cameras and microphones. The input data collected is real-time video and audio, and this data is sent to the server.

[2106] Step 4:

[2107] The server receives the video and audio data transmitted in real time and analyzes the collected data using each model. Specifically, the video and audio data is analyzed using a face recognition model, a body shape recognition model, a motion recognition model, and a voice recognition model to extract features. These features are then compared and verified with data registered in an existing database. As a result, the worker or suspicious individual is identified.

[2108] Step 5:

[2109] When the server identifies a suspicious individual, it summarizes their characteristics, including suspicious behavior and voice, and generates a warning notification based on this.

[2110] Step 6:

[2111] Server-generated alert notifications are sent to administrators and security organizations. The timing and method of notifications are determined by the system configuration, and can be immediate or periodic.

[2112] Step 7:

[2113] The device receives the notification from the server and activates the alarm device, specifically emitting an alarm sound or flashing a warning light.

[2114] Step 8:

[2115] The user (administrator) checks the notification on their smartphone or PC. The notification contains detailed information and video of the suspicious individual. The notification content is input data, and the administrator checks the situation as output. If necessary, they contact the security organization and take appropriate action.

[2116] Through the above processing steps, this system strengthens the security of the manufacturing site in real time, making it possible to immediately identify suspicious individuals and take appropriate measures.

[2117] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2118] The present invention relates to a security system that accurately distinguishes between family members and suspicious individuals by registering their faces, body shapes, and movements in advance and analyzing video and audio data collected using a camera and microphone. Furthermore, by combining it with an emotion engine, the system aims to recognize the user's emotions and automatically provide appropriate countermeasures. Specific embodiments are described in detail below.

[2119] 1. System Configuration

[2120] The system consists of the following components:

[2121] User device: A smartphone or PC operated by the user.

[2122] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[2123] Server: A central server that collects, analyzes, stores, and recognizes data.

[2124] 2. Initial Setup

[2125] User

[2126] 1. The user installs the security system application, creates a new account and logs in.

[2127] 2. Upload facial photos, body shapes, video of movements, and voice data of family members and related parties to the application and send them to the server.

[2128] server

[2129] 1. Create a database based on the received data and register the characteristics of family members and related parties.

[2130] 3. Daily operations

[2131] a. Data collection

[2132] Terminal

[2133] 1. The monitoring terminal (camera and microphone) monitors the monitored area 24 hours a day and collects video and audio in real time.

[2134] 2. Send the collected data to the server.

[2135] b. Data analysis and feature matching

[2136] server

[2137] 1. Analyze the received video data and extract facial features using a facial recognition algorithm.

[2138] 2. Analyze the received voice data and extract voice features using a voice recognition algorithm.

[2139] 3. The extracted features are compared with family and related person data in the database to identify suspicious individuals.

[2140] 4. Emotion recognition and countermeasures

[2141] a. Emotion recognition

[2142] server

[2143] 1. During the process of analyzing voice data, the emotion engine is used to determine the user's emotions.

[2144] 2. Emotion recognition algorithms analyze the tone, pitch, rhythm, and speed of the voice to determine whether the user is in an emotional state (e.g., fear or panic).

[2145] b. Emotion-based measures

[2146] server

[2147] 1. Adjust the intensity and method of warning sounds and lights based on emotion recognition results. For example, if the user is feeling fear, a stronger warning will be issued.

[2148] 2. Based on the emotion recognition results, the system proposes and implements appropriate countermeasures. For example, if the user feels fear, the system automatically notifies a security company.

[2149] 3. If necessary, notify the user's friends and family and ask for their assistance.

[2150] 5. User Response

[2151] User

[2152] 1. Check the notification on your smartphone or PC and check the suspicious person's details and recorded video through the system application.

[2153] 2. If necessary, contact a security company or the police and take appropriate action.

[2154] Program processing explanation

[2155] The processing of the program of the present invention will be specifically described below.

[2156] User Initial Settings

[2157] User

[2158] 1. The user logs in to the security system and uploads photos of the faces, body shapes, and movement videos of family members and related parties to the application.

[2159] 2. The user uses a microphone to record their own voice or that of their family members and sends it to the server.

[2160] server

[2161] 1. The server receives the uploaded data and analyzes it using facial recognition, body shape recognition, movement recognition, and voice recognition algorithms.

[2162] 2. Extract features and store them in a database.

[2163] Data collection and analysis

[2164] Terminal

[2165] 1. The monitoring terminal collects video and audio in real time and sends them to the server.

[2166] server

[2167] 1. The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[2168] 2. Analyze the voice data and extract voice features using a voice recognition algorithm.

[2169] 3. The server compares the features and identifies suspicious individuals.

[2170] Emotion recognition and countermeasures

[2171] server

[2172] 1. Use an emotion engine to recognize user emotions while analyzing voice data.

[2173] 2. Based on the emotion recognition results, set and implement appropriate warnings and notification methods.

[2174] 3. Automatically notify security companies or contacts for assistance if necessary.

[2175] Specific examples

[2176] For example, when installing this security system in a home, the following specific examples are conceivable.

[2177] 1. Initial Setup

[2178] A user logs into a security system application and configures the cameras and microphones in their home.

[2179] Register photos of each family member's face, height, body shape, and video of their movements. Also, record and register each person's voice.

[2180] 2. Daily operations

[2181] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[2182] The monitoring terminal sends the collected data to a server, which analyzes it.

[2183] 3. Identifying suspicious individuals

[2184] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[2185] The user checks the notification and confirms the details of the suspicious person.

[2186] 4. Emotion Recognition and Countermeasures

[2187] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[2188] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[2189] In this way, the crime prevention system of the present invention enables highly accurate and automatic crime prevention measures. Furthermore, the emotion recognition function allows for quick and appropriate responses even when the user feels fear, providing a safer and more secure living environment.

[2190] The processing flow will be explained below.

[2191] User Initial Settings

[2192] User

[2193] Step 1: The user installs the security system application on their smartphone or PC, creates a new account and logs in.

[2194] Step 2: The user takes and uploads photos of family members and related people, along with their height, body shape, and movement videos, to the application.

[2195] Step 3: The user uses the microphone to record their own voice or that of a family member and sends it to the server via the application.

[2196] server

[2197] Step 4: The server receives the uploaded face photo, height, body shape, motion video, and audio data.

[2198] Step 5: The server analyzes the received data and extracts features from each data using algorithms for face recognition, body shape recognition, movement recognition, and voice recognition.

[2199] Step 6: The server stores the extracted features in a database and creates profiles of family members and related parties.

[2200] Data Collection and Monitoring

[2201] Terminal

[2202] Step 1: The monitoring terminal (camera and microphone) monitors the monitored area 24 hours a day and collects video and audio in real time.

[2203] Step 2: The device sends the collected video and audio data to the server.

[2204] Data analysis and feature matching

[2205] server

[2206] Step 3: The server analyzes the received video data and extracts facial features using a facial recognition algorithm.

[2207] Step 4: The server analyzes the received voice data and extracts voice features using a voice recognition algorithm.

[2208] Step 5: The server compares the extracted features with family and related person data in the database and performs identification.

[2209] Emotion recognition and countermeasures

[2210] server

[2211] Step 6: The server uses the emotion engine while analyzing the voice data to determine the user's emotion by analyzing the tone, pitch, rhythm, speed, etc. of the voice.

[2212] Step 7: The server adjusts the intensity and method of the warning sound and light based on the emotion recognition result. For example, if the user feels fear, it will emit a strong warning sound.

[2213] Step 8: The server sends notifications to the user's friends and family to ask for help, if necessary.

[2214] Step 9: Based on the user's emotion recognition results, a function to automatically notify a security company is executed.

[2215] Identifying and alerting suspicious individuals

[2216] server

[2217] Step 10: If the server identifies a suspicious person, it summarizes their characteristics (gender, clothing, body shape, movement patterns, etc.).

[2218] Step 11: The server sends a notification to the user and the security company.

[2219] Step 12: Notifications are sent via message, email, in-app notification, or other methods according to the user's settings.

[2220] Terminal

[2221] Step 13: The device receives the notification from the server and issues an alert or turns on a light.

[2222] User response

[2223] User

[2224] Step 14: The user checks the notification on their smartphone or PC and checks the suspicious person's details and recorded video through the system application.

[2225] Step 15: If necessary, the user contacts a security company or the police and takes appropriate action.

[2226] As a concrete example, when installing this security system at home, the process is as follows:

[2227] 1. Initial Setup

[2228] The user logs in to the application and registers face photos, body shapes, motion videos, and voices of all family members.

[2229] 2. Daily operations

[2230] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[2231] The collected data is sent to a server and analyzed in real time.

[2232] 3. Identifying suspicious individuals

[2233] If the server identifies a suspicious person, it notifies the user of their characteristics.

[2234] The user checks the notification and confirms the details of the suspicious person.

[2235] 4. Emotion Recognition and Countermeasures

[2236] The server analyzes the voice characteristics and if the user is feeling fearful, it will emit a strong warning sound.

[2237] If necessary, notify a security company or acquaintances and request assistance.

[2238] As a result, the security system of the present invention eliminates the hassle of setup and operation and provides highly accurate, automatic security measures. It can also be easily used by elderly people, helping to create a safe and secure living environment.

[2239] Example 2

[2240] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2241] The main function of existing security systems is to identify suspicious individuals by registering the faces, body shapes, and movements of family members and related parties and analyzing the data collected by cameras and microphones. However, these systems lack the ability to recognize the user's emotional state and take appropriate action based on that. This makes it difficult to respond quickly and appropriately when the user becomes frightened or panics.

[2242] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for registering personal recognition data of family members, means for analyzing video and audio data collected by the image capture device and the audio capture device, means for comparing the analyzed data with registered family characteristics, means for emitting an alarm sound or light when a suspicious person is identified, means for summarizing the characteristics of the suspicious person and notifying the user and a security agency, means for analyzing the user's emotional state using voice recognition, and means for adjusting the method of warning and notification based on the emotional state. This enables a quick and appropriate response even when the user is feeling fear or panic.

[2243] "Family member personal recognition data" refers to data including facial photographs, body shapes, motion videos, and voice features of family members.

[2244] "Image capture device" refers to a camera or other video capture device that collects video data of a monitored area.

[2245] "Audio capture device" refers to a microphone or other audio capture device that collects audio data from a monitored area.

[2246] "Means of analysis" refers to algorithms and software for processing collected video and audio data and extracting meaningful features from it.

[2247] The "matching means" refers to the technology and method for comparing the analyzed features with pre-registered personal recognition data of family members and determining the degree of match.

[2248] "Means for emitting warning sounds or lights" refers to a device or system that emits sounds or lights to warn a user when a suspicious person is identified.

[2249] The "means for summarizing characteristics and notifying the user and security agency" refers to the technology and method for summarizing the identification information of a suspicious person and transmitting that information to the user's terminal and the security agency.

[2250] "Means for analyzing a user's emotional state using voice recognition" refers to algorithms or software that analyze a user's voice data to determine their emotional state (e.g., fear, anger, calm, etc.).

[2251] "Means for adjusting warning and notification methods based on emotional state" refers to techniques and methods for changing the intensity and method of warning sounds and notifications depending on the analyzed emotional state of the user.

[2252] This invention relates to a crime prevention system that can accurately distinguish between family members and suspicious individuals by registering face, body shape, and movement data of family members in advance and analyzing video and audio data collected using an image capture device and an audio capture device. Furthermore, by combining it with an emotion engine, the system aims to recognize the user's emotional state and automatically provide appropriate countermeasures.

[2253] System Configuration

[2254] The system consists of the following components:

[2255] User device: A smartphone or PC operated by the user.

[2256] Surveillance terminal: A surveillance device with a built-in camera and microphone.

[2257] Server: A central server that collects, analyzes, stores, and recognizes data.

[2258] Initial Setup

[2259] User

[2260] 1. The user installs the security system application on their smartphone or PC, creates a new account and logs in.

[2261] 2. Within the application, take or select and upload photos of your family members' faces, body shapes and movement videos. Also, use the microphone to record the voices of all family members and send them to the server.

[2262] server

[2263] 1. The server analyzes the received facial photos, body shape, motion video, and audio data and extracts each feature. TensorFlow is used for facial recognition, Google Cloud Speech-to-Text API for voice recognition, and OpenCV for body shape and motion recognition.

[2264] 2. Save the extracted features in a database (e.g., MySQL).

[2265] Daily operations

[2266] Terminal

[2267] 1. Surveillance devices (e.g. security cameras and microphones) monitor the area 24 hours a day.

[2268] 2. The monitoring device sends the collected video and audio data to the server in real time using AWS Lambda.

[2269] server

[2270] 1. The server analyzes the received video data using a facial recognition algorithm (e.g., DeepFace) and extracts facial features.

[2271] 2. The voice data is analyzed using a voice recognition algorithm (e.g., DeepSpeech) to extract voice features.

[2272] 3. The features are matched with existing data in the database to identify suspicious individuals. This matching is performed using SQL queries.

[2273] Emotion recognition and countermeasures

[2274] server

[2275] 1. The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions based on the voice data collected in real time.

[2276] 2. Emotion recognition algorithms are used to analyze the tone, pitch, rhythm, and speed of the voice to determine the user's emotional state (e.g., fear or panic).

[2277] 3. Depending on the emotional state, take action, such as "set alarm sound to maximum level" or "increase lighting."

[2278] 4. If necessary, use the Twilio API to notify security companies or friends.

[2279] Specific examples

[2280] For example, when installing this security system in a home, the following specific examples are possible:

[2281] 1. Initial Setup

[2282] A user logs into a security system application (e.g., FamilySafe) and configures their home camera and microphone.

[2283] Register photos of each family member's face, body shape, and movement videos, and also record and register each person's voice.

[2284] 2. Daily operations

[2285] Cameras and microphones are installed near the entrance and windows of the home to monitor the home 24 hours a day.

[2286] The monitoring terminal sends the collected data to a server, which analyzes it.

[2287] 3. Identifying suspicious individuals

[2288] If the server identifies a suspicious person, it summarizes their characteristics and notifies the user's smartphone.

[2289] The user checks the notification and checks the suspicious person's details and recorded video.

[2290] 4. Emotion Recognition and Countermeasures

[2291] The server analyzes the user's voice and issues a strong warning if they are feeling fear or panic.

[2292] If necessary, it will automatically notify security companies or acquaintances to request appropriate assistance.

[2293] Prompt Sentence Examples

[2294] Here are some example prompts to input to a generative AI model:

[2295] "Please explain a system that identifies suspicious faces and voices and analyzes user emotions from video and audio data collected by home surveillance cameras."

[2296] The crime prevention system of this invention enables highly accurate and automatic crime prevention measures. Furthermore, the emotion recognition function allows for quick and appropriate responses even when the user feels fear, providing a safer and more secure living environment.

[2297] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2298] Step 1: User Data Entry

[2299] User

[2300] Users install the security system application on their smartphone or PC, create a new account, and log in. Within the application, they take or select and upload photos of their family members' faces, body shapes, and movement videos. They also use a microphone to record the voices of all family members and send them to the server.

[2301] Input: Face photo, body shape, motion video, audio recording

[2302] Output: Personal identification data sent to the server

[2303] Step 2: Data analysis and registration by the server

[2304] server

[2305] The server analyzes the received facial photos, body shape, motion video, and audio data, extracting each feature. This process uses TensorFlow and the Google Cloud Speech-to-Text API, and image data is analyzed using OpenCV. For example, it performs tasks such as "extracting facial feature points," "measuring height," and "identifying motion patterns." The extracted features are then stored in a database (e.g., MySQL).

[2306] Input: Personal identification data submitted by the user

[2307] Output: Analyzed feature data registered in the database

[2308] Step 3: Collect data from the device

[2309] Terminal

[2310] The monitoring device (e.g., security camera and microphone) monitors the area 24 hours a day. The monitoring device transmits the collected video and audio data to the server in real time. This is done using AWS Lambda. For example, the camera frame rate is set to 30 fps (30 frames per second), and the microphone is highly sensitive.

[2311] Input: Video and audio data of the monitored area

[2312] Output: Video and audio data sent to the server in real time

[2313] Step 4: Data analysis and collation by the server

[2314] server

[2315] The server analyzes the received video data using a facial recognition algorithm (e.g., DeepFace) to extract facial features. It also analyzes the audio data using a speech recognition algorithm (e.g., DeepSpeech) to extract voice features. These features are then compared with existing data in a database to identify suspicious individuals. For...

Claims

1. A way to register the faces, body shapes, and movements of family members, means for analyzing the video and audio data collected by the camera and microphone; a means of matching the analyzed data with registered family characteristics; A means for emitting a warning sound or light when a suspicious person is identified; A system including a means for summarizing the characteristics of a suspicious person and notifying the user and the security company.

2. 2. The system according to claim 1, further comprising a voice data analysis means for registering and verifying the user's voice and speech.

3. 2. The system according to claim 1, further comprising a user interface that allows a user to set the timing and method of notification when a suspicious person is identified.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A