system

The system addresses the inadequacies of current security measures by integrating facial recognition and anomaly detection with real-time alert systems and local network sharing to enhance community safety through rapid response to intrusions.

JP2026100556APending Publication Date: 2026-06-19SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-09
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Current security measures are inadequate in preventing intrusions by suspicious persons and often rely on post-incident responses, lacking real-time information sharing and effective community-wide safety enhancements.

Method used

A system that integrates facial recognition, anomaly detection in structures, and real-time alarm systems, coupled with local network information sharing to identify and alert users to suspicious individuals and structural anomalies, enabling rapid community responses.

Benefits of technology

Enhances community safety by promptly detecting and responding to intrusions, integrating multiple sensors and cameras for real-time information processing and sharing, and providing user-friendly alert systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026100556000001_ABST
    Figure 2026100556000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for acquiring video data and performing face recognition processing on said data, A method for analyzing a person's behavior and facial expressions and scoring their level of danger, A means for detecting abnormalities in structures such as windows and doors, A means of issuing an alarm and sending an alarm notification to the user, A means of sharing information about suspicious individuals, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, there has been a demand to protect residents from the increasing damage caused by intrusions into houses and suspicious persons. However, current security measures are not always effective in preventing the intrusion of suspicious persons and often only involve post - incident responses. Also, there is a current situation where real - time sharing of information about suspicious persons and prompt alarm transmission are not sufficiently carried out, and it has not contributed to improving the safety of the entire region.

Means for Solving the Problems

[0005] This invention includes means for acquiring video data and performing facial recognition processing based on it, and means for scoring the degree of danger by analyzing a person's actions and facial expressions. It also includes means for detecting abnormalities in structures such as windows and doors, and when an abnormality occurs, it issues an alarm and sends an alarm notification to the user. Furthermore, by linking and sharing information on suspicious persons with a local network, it enables a swift response throughout the entire area. This system aims to prevent intrusions in advance, enable a rapid response, and improve the safety of the community.

[0006] "Video data" refers to real-time image information acquired from surveillance cameras and other sources.

[0007] "Face recognition processing" refers to the process of detecting a person's face from acquired video data and comparing it with a database.

[0008] "Analysis of behavior and facial expressions" refers to the technology of analyzing the movements and facial expressions of a person in a video and evaluating that person's behavior.

[0009] "Risk scoring" refers to a method of expressing a person's potential risk level as a score based on the results of analyzed behavior and facial expressions.

[0010] "Anomaly detection in structures" refers to technology that detects unusual movements or conditions in security targets such as windows and doors.

[0011] An "alarm" refers to a warning system that uses sound, light, or notification to alert the system when an anomaly is detected.

[0012] "Warning notification to the user" refers to a communication method that sends information about detected anomalies to the user's terminal to promptly issue a warning.

[0013] "Suspicious person information" refers to data on individuals who are not registered, as well as information including their behavioral history that may have been deemed potentially dangerous.

[0014] A "local network" refers to a communication network for sharing information and collaborating with local crime prevention organizations and related institutions. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, when an emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] The security system of the present invention is implemented in a manner that enables the detection of suspicious individuals and a rapid response through the functions of the server, terminal, and user.

[0037] The server continuously receives video data from multiple surveillance cameras. This data is processed in real time and compared against faces registered in an existing database via a facial recognition algorithm. If there is no match, the person is recorded as a suspicious individual. This facial recognition process makes it possible to identify suspicious individuals.

[0038] The terminal continuously monitors using sensors installed in structures such as windows and doors. When an anomaly is detected, it immediately issues an alarm based on instructions from the server and notifies the user of the relevant location. This notification is sent to the user's smartphone or tablet device.

[0039] Upon receiving a notification, users can immediately operate the application and access live video on the server. This allows users to check the situation on site even when away from home and make arrangements to contact the police or security company directly if necessary.

[0040] Furthermore, the server has the function of sharing analyzed information about suspicious individuals with the local crime prevention network. This information sharing can improve safety within the community and strengthen cooperation among multiple organizations. For example, when the server detects the same suspicious individual in the area, it immediately transmits the information so that nearby residents and cooperating organizations can respond quickly.

[0041] Thus, this invention ensures the safety of not only individual residences but also entire communities by combining multi-layered security measures. It achieves effective crime prevention functionality through a design that enhances user convenience and system responsiveness.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The server continuously receives video data from the surveillance cameras and prepares to process that data in real time.

[0045] Step 2:

[0046] The server applies a facial recognition algorithm to detect faces in the video data. It then compares these faces with a registered database and identifies any unmatched faces as suspicious individuals.

[0047] Step 3:

[0048] The server analyzes a person's behavior and facial expressions and scores their level of risk. If specific abnormal behavior is detected, a high risk score is assigned.

[0049] Step 4:

[0050] The terminal constantly monitors sensors attached to the structure and detects abnormalities in windows and doors.

[0051] Step 5:

[0052] If an anomaly is detected, the terminal immediately issues an alarm, receives instructions from the server, and sends an alarm notification to the user.

[0053] Step 6:

[0054] Users receive alarm notifications on their smartphones or tablets, launch the application, and view real-time live video.

[0055] Step 7:

[0056] Establish a system that allows users to assess the situation and contact the police or security company as needed.

[0057] Step 8:

[0058] The server shares information about suspicious individuals with the local crime prevention network, coordinating efforts to strengthen security measures.

[0059] (Example 1)

[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0061] In recent years, with the increasing importance of security, there has been a growing need for systems that can efficiently detect suspicious individuals and identify anomalies early. However, conventional systems have struggled to integrate multiple sensors and cameras, process information in real time, and appropriately share information about suspicious individuals. Furthermore, users have limited means to quickly confirm the situation and take emergency action. To address these challenges, the development of new security systems is needed.

[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0063] In this invention, the server includes means for acquiring video data and performing face recognition processing on the data; means for comparing the acquired face features with an existing database and recording the person as a suspicious person if there is a mismatch; means for detecting anomalies using sensors installed on structures; means for enabling access to live video using a communication terminal; and means for transmitting suspicious person information to a shared platform. This enables real-time detection of suspicious persons and information sharing by integrating information from multiple data sources, and provides users with means to quickly confirm the situation and respond immediately as needed.

[0064] "Video data" refers to visual information acquired by surveillance cameras and other recording devices.

[0065] "Facial recognition processing" is an analytical method that detects facial features contained in video data and uses a specific algorithm to identify individuals.

[0066] A "database" is a collection of information organized according to specific rules, and is used to match retrieved data within a system.

[0067] A "suspicious person" refers to an individual who is not registered in the database or who has not been given prior permission.

[0068] A "detector" is a device used to detect changes in the state of structures such as windows and doors.

[0069] An "alarm" is a visual or auditory signal issued to draw attention when an abnormality is detected.

[0070] A "communication terminal" refers to an electronic device used by a user to receive or send information.

[0071] "Live video" refers to footage filmed in real time, providing information that shows the situation as it is currently unfolding.

[0072] A "shared infrastructure" is a platform for securely and efficiently communicating information between multiple organizations and devices.

[0073] This security system provides efficient intruder detection and rapid response to anomalies through the functions of the server, terminals, and users.

[0074] The server continuously acquires video data from multiple cameras. This video data is processed in real time and analyzed by a face recognition algorithm. Software libraries such as OpenCV and dlib are used here. The acquired face features are compared with an existing database, and if there is no match, the person is recorded as a suspicious individual. This record functions as a log for taking appropriate action depending on the situation.

[0075] The terminal uses sensors installed in windows and doors to detect any abnormalities. Examples of these sensors include magnetic sensors and infrared sensors, which can instantly detect unauthorized opening or movement of structures. When an abnormality is detected, the terminal issues an alarm. The alarm alerts the affected person with sound and light. Simultaneously, an alarm notification is sent to the user's communication terminal. Firebase Cloud Messaging can be used as the communication protocol here.

[0076] Based on the received notification, the user accesses live video footage provided by the server by operating a dedicated application on their communication terminal. This live video allows the user to check the situation at the scene remotely. Depending on the situation, the user can immediately contact the relevant authorities through the application's functions. For example, if an anomaly is notified while the user is away from home, they can check the scene from their mobile device and contact the security company directly.

[0077] Furthermore, the server transmits information recorded as suspicious individuals to a shared platform, allowing the information to be shared with relevant parties through the regional communication platform. This enables neighboring areas and related organizations to take similar actions in real time and cooperate with each other.

[0078] This system is designed not only to protect individual homes but also to improve the safety of the entire community. Furthermore, it is technically flexible and can be operated by combining various sensor and communication technologies. The following example prompt can be used in the generated AI model: "Please describe the process for detecting suspicious individuals and sharing information."

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The server acquires video data from the monitoring device. The acquired video data is broken down into frames, and each frame is used as input to a face recognition algorithm. Specifically, it detects faces of people within the frame and extracts their features. The output is face feature data.

[0082] Step 2:

[0083] The server compares the extracted facial features with those stored in the existing database. In this step, the server compares the facial features with those stored in the database, and if there is a mismatch, the newly detected face is recorded as a suspicious person. Cosine similarity is used for similarity determination here. The output is the result of the determination as to whether or not to record the face as a suspicious person.

[0084] Step 3:

[0085] The terminal monitors data from sensors installed on windows and doors. The input is real-time open / close data transmitted from the sensors. When abnormal opening or closing is detected, that information is sent to the server. The output is an alarm message indicating that an abnormality has been detected.

[0086] Step 4:

[0087] The terminal receives instructions from the server and issues an alarm. The input is alarm information from the server, and the output is a physical alarm signal (sound or light). It also has the function of sending this information as a notification to communication terminals.

[0088] Step 5:

[0089] Users receive alarm notifications on their communication terminals. Upon receiving a notification, users can launch a dedicated application and access live video from the server. This live video is real-time data streamed from the server and serves as input data for users to check the situation.

[0090] Step 6:

[0091] The server transmits recorded information about suspicious individuals to a local shared platform. The input is the information recorded as a suspicious person, and the output is the shared information transmitted to the local communication platform. This allows for information sharing with neighboring areas and related organizations, enabling a rapid response.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] In modern society, personal safety and community security are critical issues. However, systems that enable the immediate identification of suspicious individuals, rapid information sharing, and appropriate responses remain inadequate. In particular, when a suspicious person is identified, quickly sharing that information within the local community and improving overall community safety is a challenging task.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for acquiring video data and performing facial recognition processing, means for sending alarm notifications to users using an information and communication network, and means for sharing suspicious person information in cooperation with the local community. This enables the rapid issuance of alarms when a suspicious person is detected in the area, immediate sharing of suspicious person data, and improvement of local safety.

[0097] "Video data" refers to visual information acquired from recording devices such as surveillance cameras.

[0098] "Face recognition processing" refers to the process of identifying the face of a specific person from acquired video data using an algorithm.

[0099] "Methods for scoring risk" refer to methods that analyze a person's behavior and facial expressions and quantify their level of risk.

[0100] "Means for detecting abnormalities in structures" refers to methods of checking for abnormalities using sensors attached to physical structures such as windows and doors.

[0101] An "information and communication network" refers to a system that sends and receives data via the internet or other digital communications.

[0102] "Alert notifications" refer to information that warns the user in response to the detection of anomalies or suspicious individuals.

[0103] A "local community" refers to a group of people and related organizations that live in a particular area.

[0104] "Suspicious person information" refers to data about suspicious individuals identified through facial recognition or other means.

[0105] "Information terminals" refer to electronic devices that users can carry with them and that can connect to a network, such as smartphones and tablets.

[0106] "Live video" refers to visual information that is filmed in real time and distributed via a network.

[0107] To implement this invention, specific hardware and software are required to build a security system. Specific embodiments are described below.

[0108] server

[0109] The server continuously receives video data from multiple surveillance cameras. This requires network-connected cameras and a server with real-time data processing capabilities. For facial recognition processing, software libraries capable of executing facial recognition algorithms, such as OpenCV or TENSORFLOW®, are used. The server analyzes the video data using the facial recognition algorithm to identify suspicious individuals. This information about suspicious individuals is shared to improve the safety of the local community.

[0110] terminal

[0111] The terminal constantly monitors data from sensors installed on windows and doors. These sensors detect physical anomalies and issue an alarm when an anomaly is detected. Furthermore, based on instructions from a server, it sends warnings to the user's information terminal via the network. Communication technologies such as Wi-Fi and Bluetooth are used for information communication.

[0112] User

[0113] Users can access live video feeds on the server via smartphones and other information devices. This allows users to check the situation even when they are away from home and to make emergency contact with local communities and security agencies as needed. Users also play a role in quickly sharing information about suspicious individuals to ensure local safety.

[0114] Specific example

[0115] For example, if a suspicious person is detected in the area one night, the server immediately sends a push notification to the user's smartphone, allowing the user to view real-time video. This enables the user to quickly understand the situation and, if necessary, contact the police or local crime prevention organizations.

[0116] Example of a prompt

[0117] Please explain the mechanism of a system that identifies suspicious individuals from security camera footage and shares that information with local residents in real time. Please include specific technologies and procedures in detail.

[0118] This format is expected to significantly improve local safety and individual peace of mind.

[0119] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0120] Step 1:

[0121] The server receives video data from multiple surveillance cameras. This data is prepared for processing in real time. Once the video data arrives at the server, it is first input into a face recognition algorithm. The server analyzes each frame of the video and uses libraries such as OpenCV to identify clear facial features. As a result, a list of potential suspicious individuals is output.

[0122] Step 2:

[0123] The server compares the facial information identified by the facial recognition algorithm with an existing database. The facial data obtained as input is compared with known data in the database. If no matching entry is found in the database, the person is recorded as a suspicious person. This process completes the identification of the suspicious person, and the information is sent to the next step.

[0124] Step 3:

[0125] The terminal monitors signals from sensors installed on windows and doors. When an anomaly is detected, it processes the signal and records the location. This anomaly information is immediately sent to the server. Based on this information, the server generates instructions to issue an alarm at the location in question.

[0126] Step 4:

[0127] The server sends an alert notification to the user's information terminal based on anomaly and suspicious person information. Using the suspicious person and anomaly information obtained as input, it generates an integrated alert if both are related. This alert is delivered to the user via a push notification service such as Firebase Cloud Messaging. Upon receiving this notification, the user can immediately check the warning on their smartphone.

[0128] Step 5:

[0129] Users access the server using their information terminals to view live video. The server streams the video from the relevant surveillance camera using the RTSP protocol in response to the user's request. This allows users to understand the situation in real time and prepare to cooperate with the police or security organizations if necessary.

[0130] Step 6:

[0131] The server shares information about suspicious individuals with the local community's safety network. The entered data is sent to a database linked to local crime prevention organizations and residents. This information sharing raises crime prevention awareness throughout the community and enables residents to respond quickly.

[0132] The above outlines the system processing steps required to implement the application example. Each step works in conjunction with the others, aiming to enhance safety throughout the entire region.

[0133] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0134] This invention relates to a security system equipped with an emotion engine in addition to acquiring and processing video data. It recognizes the user's emotions in addition to detecting suspicious individuals, and allows for adjustment of the content of alarm notifications.

[0135] The server has the capability to process video data received from multiple surveillance cameras in real time. This processing uses facial recognition technology and behavioral / expression analysis technology to identify individuals and score their risk level. This allows for the immediate identification of individuals suspected of being intruders.

[0136] In addition, an emotion engine installed on the device constantly monitors the user's emotional state. This emotion engine infers emotions from the user's voice, behavioral patterns, and how they operate the device, and sends this data to a server. The server uses this emotional data to adjust alert notifications to match the user's psychological state. Specifically, if the server detects that the user is experiencing stress or anxiety, it can provide more reassuring content and adjust the frequency of notifications.

[0137] For example, if a window anomaly is detected while the user is showing signs of anxiety, the server will use gentle and considerate language in the notification and provide an option to contact a support desk that can respond quickly and professionally. The emotion engine also contributes to improving the accuracy of suspicious person information; by including emotion data when the server shares information in cooperation with the local network, it helps in developing safety measures based on the suspect's state and reactions.

[0138] Thus, the present invention provides an advanced type of security system that combines further enhancement of crime prevention functions with consideration for the psychological aspects of the user experience.

[0139] The following describes the processing flow.

[0140] Step 1:

[0141] The server receives video data from multiple surveillance cameras and prepares to process the data in real time.

[0142] Step 2:

[0143] The server applies a facial recognition algorithm to compare the faces of people in the video with existing data in the database. If a face is unfamiliar, it is marked and recorded as a suspicious person.

[0144] Step 3:

[0145] The server analyzes a person's behavior and facial expressions, and scores their level of danger based on the data obtained. If abnormal behavior is detected, a higher score is assigned.

[0146] Step 4:

[0147] The device monitors the user's emotional state using an emotion engine and infers emotions from the user's voice and operation patterns.

[0148] Step 5:

[0149] Emotional data from the emotion engine is sent to the server, which uses this data to adjust the content and method of alarm notifications.

[0150] Step 6:

[0151] If the server detects an anomaly using sensors installed on structures such as windows and doors, it immediately issues an alarm and sends an emotion-based, tailored notification to the user.

[0152] Step 7:

[0153] Users receive notifications on their smartphones or tablets and, depending on the situation, check live video to understand the current state of affairs.

[0154] Step 8:

[0155] If a user suspects an intruder, they can decide to contact a security company or the police and choose an appropriate support option that is sensitive to their feelings.

[0156] Step 9:

[0157] The server will share information on suspicious individuals and user sentiment data collected with the local network to enhance safety throughout the region.

[0158] (Example 2)

[0159] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0160] Modern security systems focus on detecting suspicious individuals but often neglect the psychological burden on users, with warning notifications sometimes causing anxiety and stress. Furthermore, there are limited means to improve the accuracy of suspicious person information and to efficiently share safety information with the community. To address these challenges, security features that consider the emotions of users are needed.

[0161] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0162] In this invention, the server includes means for acquiring video information and performing face recognition processing from the information, means for analyzing the target's behavior and facial expressions and assessing the degree of danger, and means for performing emotion analysis and adjusting the content of warning notifications based on the user's psychological state. This enables warning notifications that take into account the user's psychological state and improves crime prevention functions by enabling effective information sharing with the local community.

[0163] "Video information" refers to visual data captured by surveillance equipment. This data is used to understand the movements and circumstances of the monitored subject in real time.

[0164] "Facial recognition processing" refers to the technology that identifies human faces from video information, extracts their features, and uses them for identification. This technology is used in security systems to recognize individuals and match them with related data.

[0165] "Analyzing behavior and facial expressions" refers to the technique of interpreting underlying emotions and intentions from a subject's movement patterns and facial expressions. This technique is useful for predicting risky behavior and understanding psychological states.

[0166] "Assessing the level of risk" refers to the process of quantifying and evaluating potential risks based on analyzed behavior, facial expressions, and other sensor data.

[0167] "Sentiment analysis" refers to a technology that infers a user's emotional state from voice, text, and behavioral data. This technology makes it possible to adapt warning content according to the user's psychological state.

[0168] A "warning notification" refers to information issued to users in response to detected anomalies or risks. This information is typically communicated using visual and auditory means.

[0169] "Adjusting warning notification content based on psychological state" refers to the process of optimizing the wording and expression of warnings by taking into account the user's current emotions and stress level.

[0170] "Suspicious person information" refers to data on individuals or behaviors deemed to pose a potential threat detected within a monitored area. This information is used to strengthen crime prevention measures and to share information with the local community.

[0171] This invention utilizes advanced monitoring and emotion analysis technologies in security systems to provide warning notifications that take into account the user's psychological state.

[0172] The server acquires and processes video information in real time from multiple monitoring devices. The hardware used consists of network-connected cameras and recorders, and the data is efficiently managed through a stream processing system. The server processes this video information using general image processing libraries and APIs (for example, image analysis software for face recognition and software specialized for behavioral analysis) to perform face recognition and behavioral / facial expression analysis of subjects.

[0173] The device has an engine for sentiment analysis and is equipped with sensors for voice recognition and behavioral analysis. The device captures the user's voice and actions and analyzes their emotional state using a built-in sentiment analysis algorithm. This information is sent to the server in real time, and the sentiment data received by the server is used to adjust the content of warning notifications.

[0174] As a concrete example, consider a scenario where a user is at home and experiencing anxiety. When the monitoring system detects an anomaly in a window, it takes emotional data into account and quickly issues a warning using gentle, reassuring language. This process allows the user to understand the situation in a timely manner and respond without experiencing unnecessary stress.

[0175] As an example of a prompt message, we will use content such as, "Hello, how are you feeling right now? Please let us know if you have any comments or suggestions. We will use them to help us provide a safe and secure environment," to collect user feedback and improve the accuracy of the information in the security system.

[0176] In this way, in addition to regular warning notifications, the security system analyzes the user's emotional state, enabling more personalized notifications and responses, thereby providing a high level of safety and peace of mind.

[0177] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0178] Step 1:

[0179] The server starts a data stream using video information acquired in real time from monitoring equipment as input. Based on this video information, the server performs face recognition processing and extracts facial feature points. This outputs basic data for recognizing specific individuals. An image processing library is used for this process.

[0180] Step 2:

[0181] The server uses the facial feature data extracted in Step 1 as input to perform behavioral and facial expression analysis. By analyzing the movement patterns in the video and capturing changes in facial expressions, it infers the emotional state of the subject. Using a motion analysis algorithm, it outputs these analysis results as a risk score.

[0182] Step 3:

[0183] The device captures the user's voice using a microphone and uses this as input for emotion analysis. The emotion engine analyzes the voice data and identifies emotional states such as joy, anger, sadness, and happiness. The output is data that represents the user's emotional state using numbers and tags.

[0184] Step 4:

[0185] The device collects additional operation logs and behavioral patterns and sends them to the server. By analyzing what actions the user performed, psychological tendencies are monitored in more detail. This data is input, and the server outputs an overall emotional assessment.

[0186] Step 5:

[0187] The server utilizes the output data from steps 2 and 4 to generate a warning notification. The notification content and wording are customized to match the user's psychological state, adjusting to encourage caution while reducing psychological burden. The final notification is output to the terminal and communicated to the user.

[0188] Step 6:

[0189] The server shares generated suspicious person information and sentiment data with the local network. This operation strengthens security measures throughout the region and enables coordinated crime prevention activities. The output information is provided to relevant agencies.

[0190] (Application Example 2)

[0191] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0192] Security systems are required not only to detect suspicious individuals but also to provide appropriate alarm notifications that take into account the user's emotional state. However, current systems struggle to respond flexibly to the user's psychological state, failing to achieve both suspicious individual detection and user reassurance. Therefore, the challenge is to enhance security functions while improving the user experience by providing alarm notifications that take the user's emotions into consideration.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0194] In this invention, the server includes means for acquiring video information and performing facial recognition processing from the information; means for analyzing a person's movements and emotions and evaluating the degree of risk; means for detecting abnormalities in structures; means for adjusting the content of alarm notifications based on the user's emotional state; and means for linking and sharing suspicious person information with a community network. This makes it possible to adjust alarm notifications according to the user's emotional state, thereby strengthening security measures while maintaining the user's sense of security.

[0195] "Visual information" refers to visual data acquired from surveillance cameras and sensor devices, and this data is used for person identification and activity analysis.

[0196] "Face recognition processing" is a technology that detects the face of a specific person from acquired video information and analyzes related data.

[0197] "Analysis of actions and emotions" is a process that evaluates emotions and risks based on a person's behavior and facial expressions.

[0198] "Risk assessment" is the process of quantifying potential risks based on information obtained from analyzing a person's movements and facial expressions.

[0199] "Detection of structural abnormalities" is a technology that identifies security risks by detecting unusual conditions in physical facilities such as windows and doors.

[0200] "User emotional state" refers to the user's current psychological state, inferred from their voice, behavior, and how they operate the interface.

[0201] "Adjusting the content of alarm notifications" means customizing alarm and notification messages to appropriate content based on the user's emotional state.

[0202] A "community network" is an information-sharing platform that shares crime prevention information with relevant organizations and residents within the community to enhance safety throughout the region.

[0203] This invention utilizes video information and emotion data in a security system to adjust alarm notifications to users according to their emotions. Specifically, a server receives video information in real time from multiple surveillance cameras and sensor devices and performs face recognition processing. This identifies specific individuals and performs analysis of their actions and emotions. Amazon Rekognition is used for face recognition, and image processing libraries such as OpenCV and dlib are used for action and emotion analysis.

[0204] The server also detects anomalies in windows and doors using sensor devices attached to the structure. This anomaly information is combined with video information to assess the level of risk. The risk assessment uses a proprietary algorithm to quantify the risk and issues an alarm as needed.

[0205] The device collects the user's voice and behavioral data and uses Google® Cloud Speech-to-Text to infer their emotional state. Based on this data, the server adjusts the content of alert notifications and generates messages tailored to the user's psychological state. For example, if the user is feeling anxious or stressed, the notification will be delivered in a soft voice and gentle language.

[0206] Furthermore, the server integrates suspicious person information with the community network, improving local safety through an information sharing platform. This strengthens crime prevention capabilities throughout the entire community.

[0207] As a concrete example, consider a scenario where a user is relaxing at home at night, but an anomaly is detected near the window. In this case, the device recognizes the user's anxiety, and the server sends a gentle notification such as, "Please remain calm and pay immediate attention to what is happening outside the window."

[0208] An example of a prompt for a generative AI model is: "When I feel uneasy while walking down the street, please suggest what kind of notification would be appropriate if a suspicious person is detected. Example: notification in a gentle voice, option to contact support."

[0209] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0210] Step 1:

[0211] The server acquires video information in real time from surveillance cameras. The input is video data from the surveillance cameras, and the output is processed video frames. Amazon Rekognition is used to identify faces of people in the video and pass that data to the next step.

[0212] Step 2:

[0213] The server analyzes the behavior and emotions of face-recognized video data using OpenCV and dlib. The input is face-recognized video frames, and the output is data indicating the user's behavior and emotional state. Through this analysis, the server evaluates emotions from the person's behavior patterns and facial expressions, and quantifies the level of danger.

[0214] Step 3:

[0215] The device collects user voice and motion data and converts it to text using Google Cloud Speech-to-Text. Inputs are audio files and sensor data, while outputs are transcribed audio data and motion analysis results. This allows the system to infer the user's emotional state and send that information to a server.

[0216] Step 4:

[0217] The server acquires information from sensor devices to detect structural anomalies and determines whether an anomaly exists. The input is data from the sensor devices, and the output is a flag indicating the presence or absence of an anomaly. If an anomaly is detected, it is reflected in the risk assessment.

[0218] Step 5:

[0219] The server adjusts the content of the alarm notification based on the user's emotional state and risk assessment. The input is emotional state data and risk score, and the output is the adjusted alarm notification message. A generative AI model is used to create appropriate notification content and send it to the device.

[0220] Step 6:

[0221] The terminal receives pre-configured alarm notifications from the server and notifies the user gently via voice or text message. The input is the alarm notification message from the server, and the output is the voice or visual alarm delivered to the user. This provides security while maintaining the user's sense of security.

[0222] Step 7:

[0223] The server shares information about suspicious individuals in conjunction with the community network. Inputs include information about suspicious individuals and network connectivity, while output is shared information to enhance local safety. This improves crime prevention awareness throughout the community.

[0224] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0225] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0226] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0227] [Second Embodiment]

[0228] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0229] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0230] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0231] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0232] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0233] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0234] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0235] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0236] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0237] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0238] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0239] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0240] The security system of the present invention is implemented in a manner that enables the detection of suspicious individuals and a rapid response through the functions of the server, terminal, and user.

[0241] The server continuously receives video data from multiple surveillance cameras. This data is processed in real time and compared against faces registered in an existing database via a facial recognition algorithm. If there is no match, the person is recorded as a suspicious individual. This facial recognition process makes it possible to identify suspicious individuals.

[0242] The terminal continuously monitors using sensors installed in structures such as windows and doors. When an anomaly is detected, it immediately issues an alarm based on instructions from the server and notifies the user of the relevant location. This notification is sent to the user's smartphone or tablet device.

[0243] Upon receiving a notification, users can immediately operate the application and access live video on the server. This allows users to check the situation on site even when away from home and make arrangements to contact the police or security company directly if necessary.

[0244] Furthermore, the server has the function of sharing analyzed information about suspicious individuals with the local crime prevention network. This information sharing can improve safety within the community and strengthen cooperation among multiple organizations. For example, when the server detects the same suspicious individual in the area, it immediately transmits the information so that nearby residents and cooperating organizations can respond quickly.

[0245] Thus, this invention ensures the safety of not only individual residences but also entire communities by combining multi-layered security measures. It achieves effective crime prevention functionality through a design that enhances user convenience and system responsiveness.

[0246] The following describes the processing flow.

[0247] Step 1:

[0248] The server continuously receives video data from the surveillance cameras and prepares to process that data in real time.

[0249] Step 2:

[0250] The server applies a facial recognition algorithm to detect faces in the video data. It then compares these faces with a registered database and identifies any unmatched faces as suspicious individuals.

[0251] Step 3:

[0252] The server analyzes a person's behavior and facial expressions and scores their level of risk. If specific abnormal behavior is detected, a high risk score is assigned.

[0253] Step 4:

[0254] The terminal constantly monitors sensors attached to the structure and detects abnormalities in windows and doors.

[0255] Step 5:

[0256] If an anomaly is detected, the terminal immediately issues an alarm, receives instructions from the server, and sends an alarm notification to the user.

[0257] Step 6:

[0258] Users receive alarm notifications on their smartphones or tablets, launch the application, and view real-time live video.

[0259] Step 7:

[0260] Establish a system that allows users to assess the situation and contact the police or security company as needed.

[0261] Step 8:

[0262] The server shares information about suspicious individuals with the local crime prevention network, coordinating efforts to strengthen security measures.

[0263] (Example 1)

[0264] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0265] In recent years, with the increasing importance of security, there has been a growing need for systems that can efficiently detect suspicious individuals and identify anomalies early. However, conventional systems have struggled to integrate multiple sensors and cameras, process information in real time, and appropriately share information about suspicious individuals. Furthermore, users have limited means to quickly confirm the situation and take emergency action. To address these challenges, the development of new security systems is needed.

[0266] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0267] In this invention, the server includes means for acquiring video data and performing face recognition processing on the data; means for comparing the acquired face features with an existing database and recording the person as a suspicious person if there is a mismatch; means for detecting anomalies using sensors installed on structures; means for enabling access to live video using a communication terminal; and means for transmitting suspicious person information to a shared platform. This enables real-time detection of suspicious persons and information sharing by integrating information from multiple data sources, and provides users with means to quickly confirm the situation and respond immediately as needed.

[0268] "Video data" refers to visual information acquired by surveillance cameras and other recording devices.

[0269] "Facial recognition processing" is an analytical method that detects facial features contained in video data and uses a specific algorithm to identify individuals.

[0270] A "database" is a collection of information organized according to specific rules, and is used to match retrieved data within a system.

[0271] A "suspicious person" refers to an individual who is not registered in the database or who has not been given prior permission.

[0272] A "detector" is a device used to detect changes in the state of structures such as windows and doors.

[0273] An "alarm" is a visual or auditory signal issued to draw attention when an abnormality is detected.

[0274] A "communication terminal" refers to an electronic device used by a user to receive or send information.

[0275] "Live video" refers to footage filmed in real time, providing information that shows the situation as it is currently unfolding.

[0276] A "shared infrastructure" is a platform for securely and efficiently communicating information between multiple organizations and devices.

[0277] This security system provides efficient intruder detection and rapid response to anomalies through the functions of the server, terminals, and users.

[0278] The server continuously acquires video data from multiple cameras. This video data is processed in real time and analyzed by a face recognition algorithm. Software libraries such as OpenCV and dlib are used here. The acquired face features are compared with an existing database, and if there is no match, the person is recorded as a suspicious individual. This record functions as a log for taking appropriate action depending on the situation.

[0279] The terminal uses sensors installed in windows and doors to detect any abnormalities. Examples of these sensors include magnetic sensors and infrared sensors, which can instantly detect unauthorized opening or movement of structures. When an abnormality is detected, the terminal issues an alarm. The alarm alerts the affected person with sound and light. Simultaneously, an alarm notification is sent to the user's communication terminal. Firebase Cloud Messaging can be used as the communication protocol here.

[0280] Based on the received notifications, the user accesses the live video provided by the server by operating a dedicated application on the communication terminal. Through this live video, the user can check the on-site situation from a remote location. Depending on the situation, the user can immediately contact the relevant authorities through the functions of the application. For example, when an abnormality is notified while away from home, the user can check the site from the mobile information terminal and directly contact the security company.

[0281] Furthermore, the server can send the information recorded as a suspicious person to the sharing platform and share the information with relevant parties through the regional communication platform. This enables neighboring areas and related organizations to take similar actions in real time and cooperate.

[0282] This system is designed not only to protect individual houses but also to improve the safety of the entire region. Also, technically, it has flexibility and can be operated by combining various sensor technologies and communication technologies. The following example of a prompt sentence can be used for the generative AI model. "Please explain the process for detecting suspicious persons and sharing information."

[0283] The flow of the specific process in Example 1 will be described using FIG. 11.

[0284] Step 1:

[0285] The server acquires video data from the monitoring device. The acquired video data is decomposed into frames, and each frame is used as the input for the face recognition algorithm. Specifically, the faces of the people in the frame are detected, and their feature amounts are extracted. The output is the face feature amount data.

[0286] Step 2:

[0287] The server compares the extracted facial features with those stored in the existing database. In this step, the server compares the facial features with those stored in the database, and if there is a mismatch, the newly detected face is recorded as a suspicious person. Cosine similarity is used for similarity determination here. The output is the result of the determination as to whether or not to record the face as a suspicious person.

[0288] Step 3:

[0289] The terminal monitors data from sensors installed on windows and doors. The input is real-time open / close data transmitted from the sensors. When abnormal opening or closing is detected, that information is sent to the server. The output is an alarm message indicating that an abnormality has been detected.

[0290] Step 4:

[0291] The terminal receives instructions from the server and issues an alarm. The input is alarm information from the server, and the output is a physical alarm signal (sound or light). It also has the function of sending this information as a notification to communication terminals.

[0292] Step 5:

[0293] Users receive alarm notifications on their communication terminals. Upon receiving a notification, users can launch a dedicated application and access live video from the server. This live video is real-time data streamed from the server and serves as input data for users to check the situation.

[0294] Step 6:

[0295] The server transmits recorded information about suspicious individuals to a local shared platform. The input is the information recorded as a suspicious person, and the output is the shared information transmitted to the local communication platform. This allows for information sharing with neighboring areas and related organizations, enabling a rapid response.

[0296] (Application Example 1)

[0297] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0298] In modern society, personal safety and community security are critical issues. However, systems that enable the immediate identification of suspicious individuals, rapid information sharing, and appropriate responses remain inadequate. In particular, when a suspicious person is identified, quickly sharing that information within the local community and improving overall community safety is a challenging task.

[0299] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0300] In this invention, the server includes means for acquiring video data and performing facial recognition processing, means for sending alarm notifications to users using an information and communication network, and means for sharing suspicious person information in cooperation with the local community. This enables the rapid issuance of alarms when a suspicious person is detected in the area, immediate sharing of suspicious person data, and improvement of local safety.

[0301] "Video data" refers to visual information acquired from recording devices such as surveillance cameras.

[0302] "Face recognition processing" refers to the process of identifying the face of a specific person from acquired video data using an algorithm.

[0303] "Methods for scoring risk" refer to methods that analyze a person's behavior and facial expressions and quantify their level of risk.

[0304] "Means for detecting abnormalities in structures" refers to methods of checking for abnormalities using sensors attached to physical structures such as windows and doors.

[0305] "Information and communication network" refers to a system that transmits and receives data via the Internet or other digital communications.

[0306] "Alarm notification" refers to information that issues a warning to a user in response to the detection of an abnormality or a suspicious person.

[0307] "Regional community" refers to a group that includes people living in a specific region and related organizations.

[0308] "Suspicious person information" refers to data related to a suspicious person identified by face recognition or other means.

[0309] "Information terminal" refers to an electronic device that a user can carry around, such as a smartphone or a tablet, and can be connected to a network.

[0310] "Live video" refers to visual information that is captured in real time and distributed through a network.

[0311] To implement this invention, specific hardware and software for constructing a security system are required. The specific embodiments are shown below.

[0312] Server

[0313] The server continuously receives video data from multiple surveillance cameras. For this, cameras connected to a network and a server with real-time data processing capabilities are required. For face recognition processing, a software library that can execute face recognition algorithms such as OpenCV or TensorFlow is used. The server analyzes the video data using the face recognition algorithm and identifies suspicious persons. This suspicious person information is shared to improve the security of the regional community.

[0314] Terminal

[0315] The terminal constantly monitors data from sensors installed on windows and doors. These sensors detect physical anomalies and issue an alarm when an anomaly is detected. Furthermore, based on instructions from a server, it sends warnings to the user's information terminal via the network. Communication technologies such as Wi-Fi and Bluetooth are used for information communication.

[0316] User

[0317] Users can access live video feeds on the server via smartphones and other information devices. This allows users to check the situation even when they are away from home and to make emergency contact with local communities and security agencies as needed. Users also play a role in quickly sharing information about suspicious individuals to ensure local safety.

[0318] Specific example

[0319] For example, if a suspicious person is detected in the area one night, the server immediately sends a push notification to the user's smartphone, allowing the user to view real-time video. This enables the user to quickly understand the situation and, if necessary, contact the police or local crime prevention organizations.

[0320] Example of a prompt

[0321] Please explain the mechanism of a system that identifies suspicious individuals from security camera footage and shares that information with local residents in real time. Please include specific technologies and procedures in detail.

[0322] This format is expected to significantly improve local safety and individual peace of mind.

[0323] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0324] Step 1:

[0325] The server receives video data from multiple surveillance cameras. This data is prepared for processing in real time. Once the video data arrives at the server, it is first input into a face recognition algorithm. The server analyzes each frame of the video and uses libraries such as OpenCV to identify clear facial features. As a result, a list of potential suspicious individuals is output.

[0326] Step 2:

[0327] The server compares the facial information identified by the facial recognition algorithm with an existing database. The facial data obtained as input is compared with known data in the database. If no matching entry is found in the database, the person is recorded as a suspicious person. This process completes the identification of the suspicious person, and the information is sent to the next step.

[0328] Step 3:

[0329] The terminal monitors signals from sensors installed on windows and doors. When an anomaly is detected, it processes the signal and records the location. This anomaly information is immediately sent to the server. Based on this information, the server generates instructions to issue an alarm at the location in question.

[0330] Step 4:

[0331] The server sends an alert notification to the user's information terminal based on anomaly and suspicious person information. Using the suspicious person and anomaly information obtained as input, it generates an integrated alert if both are related. This alert is delivered to the user via a push notification service such as Firebase Cloud Messaging. Upon receiving this notification, the user can immediately check the warning on their smartphone.

[0332] Step 5:

[0333] Users access the server using their information terminals to view live video. The server streams the video from the relevant surveillance camera using the RTSP protocol in response to the user's request. This allows users to understand the situation in real time and prepare to cooperate with the police or security organizations if necessary.

[0334] Step 6:

[0335] The server shares information about suspicious individuals with the local community's safety network. The entered data is sent to a database linked to local crime prevention organizations and residents. This information sharing raises crime prevention awareness throughout the community and enables residents to respond quickly.

[0336] The above outlines the system processing steps required to implement the application example. Each step works in conjunction with the others, aiming to enhance safety throughout the entire region.

[0337] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0338] This invention relates to a security system equipped with an emotion engine in addition to acquiring and processing video data. It recognizes the user's emotions in addition to detecting suspicious individuals, and allows for adjustment of the content of alarm notifications.

[0339] The server has the capability to process video data received from multiple surveillance cameras in real time. This processing uses facial recognition technology and behavioral / expression analysis technology to identify individuals and score their risk level. This allows for the immediate identification of individuals suspected of being intruders.

[0340] In addition, an emotion engine installed on the device constantly monitors the user's emotional state. This emotion engine infers emotions from the user's voice, behavioral patterns, and how they operate the device, and sends this data to a server. The server uses this emotional data to adjust alert notifications to match the user's psychological state. Specifically, if the server detects that the user is experiencing stress or anxiety, it can provide more reassuring content and adjust the frequency of notifications.

[0341] For example, if a window anomaly is detected while the user is showing signs of anxiety, the server will use gentle and considerate language in the notification and provide an option to contact a support desk that can respond quickly and professionally. The emotion engine also contributes to improving the accuracy of suspicious person information; by including emotion data when the server shares information in cooperation with the local network, it helps in developing safety measures based on the suspect's state and reactions.

[0342] Thus, the present invention provides an advanced type of security system that combines further enhancement of crime prevention functions with consideration for the psychological aspects of the user experience.

[0343] The following describes the processing flow.

[0344] Step 1:

[0345] The server receives video data from multiple surveillance cameras and prepares to process the data in real time.

[0346] Step 2:

[0347] The server applies a facial recognition algorithm to compare the faces of people in the video with existing data in the database. If a face is unfamiliar, it is marked and recorded as a suspicious person.

[0348] Step 3:

[0349] The server analyzes a person's behavior and facial expressions, and scores their level of danger based on the data obtained. If abnormal behavior is detected, a higher score is assigned.

[0350] Step 4:

[0351] The device monitors the user's emotional state using an emotion engine and infers emotions from the user's voice and operation patterns.

[0352] Step 5:

[0353] Emotional data from the emotion engine is sent to the server, which uses this data to adjust the content and method of alarm notifications.

[0354] Step 6:

[0355] If the server detects an anomaly using sensors installed on structures such as windows and doors, it immediately issues an alarm and sends an emotion-based, tailored notification to the user.

[0356] Step 7:

[0357] Users receive notifications on their smartphones or tablets and, depending on the situation, check live video to understand the current state of affairs.

[0358] Step 8:

[0359] If a user suspects an intruder, they can decide to contact a security company or the police and choose an appropriate support option that is sensitive to their feelings.

[0360] Step 9:

[0361] The server will share information on suspicious individuals and user sentiment data collected with the local network to enhance safety throughout the region.

[0362] (Example 2)

[0363] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0364] Modern security systems focus on detecting suspicious individuals but often neglect the psychological burden on users, with warning notifications sometimes causing anxiety and stress. Furthermore, there are limited means to improve the accuracy of suspicious person information and to efficiently share safety information with the community. To address these challenges, security features that consider the emotions of users are needed.

[0365] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0366] In this invention, the server includes means for acquiring video information and performing face recognition processing from the information, means for analyzing the target's behavior and facial expressions and assessing the degree of danger, and means for performing emotion analysis and adjusting the content of warning notifications based on the user's psychological state. This enables warning notifications that take into account the user's psychological state and improves crime prevention functions by enabling effective information sharing with the local community.

[0367] "Video information" refers to visual data captured by surveillance equipment. This data is used to understand the movements and circumstances of the monitored subject in real time.

[0368] "Facial recognition processing" refers to the technology that identifies human faces from video information, extracts their features, and uses them for identification. This technology is used in security systems to recognize individuals and match them with related data.

[0369] "Analyzing behavior and facial expressions" refers to the technique of interpreting underlying emotions and intentions from a subject's movement patterns and facial expressions. This technique is useful for predicting risky behavior and understanding psychological states.

[0370] "Assessing the level of risk" refers to the process of quantifying and evaluating potential risks based on analyzed behavior, facial expressions, and other sensor data.

[0371] "Sentiment analysis" refers to a technology that infers a user's emotional state from voice, text, and behavioral data. This technology makes it possible to adapt warning content according to the user's psychological state.

[0372] A "warning notification" refers to information issued to users in response to detected anomalies or risks. This information is typically communicated using visual and auditory means.

[0373] "Adjusting warning notification content based on psychological state" refers to the process of optimizing the wording and expression of warnings by taking into account the user's current emotions and stress level.

[0374] "Suspicious person information" refers to data on individuals or behaviors deemed to pose a potential threat detected within a monitored area. This information is used to strengthen crime prevention measures and to share information with the local community.

[0375] This invention utilizes advanced monitoring and emotion analysis technologies in security systems to provide warning notifications that take into account the user's psychological state.

[0376] The server acquires and processes video information in real time from multiple monitoring devices. The hardware used consists of network-connected cameras and recorders, and the data is efficiently managed through a stream processing system. The server processes this video information using general image processing libraries and APIs (for example, image analysis software for face recognition and software specialized for behavioral analysis) to perform face recognition and behavioral / facial expression analysis of subjects.

[0377] The device has an engine for sentiment analysis and is equipped with sensors for voice recognition and behavioral analysis. The device captures the user's voice and actions and analyzes their emotional state using a built-in sentiment analysis algorithm. This information is sent to the server in real time, and the sentiment data received by the server is used to adjust the content of warning notifications.

[0378] As a concrete example, consider a scenario where a user is at home and experiencing anxiety. When the monitoring system detects an anomaly in a window, it takes emotional data into account and quickly issues a warning using gentle, reassuring language. This process allows the user to understand the situation in a timely manner and respond without experiencing unnecessary stress.

[0379] As an example of a prompt message, we will use content such as, "Hello, how are you feeling right now? Please let us know if you have any comments or suggestions. We will use them to help us provide a safe and secure environment," to collect user feedback and improve the accuracy of the information in the security system.

[0380] In this way, in addition to regular warning notifications, the security system analyzes the user's emotional state, enabling more personalized notifications and responses, thereby providing a high level of safety and peace of mind.

[0381] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0382] Step 1:

[0383] The server starts a data stream using video information acquired in real time from monitoring equipment as input. Based on this video information, the server performs face recognition processing and extracts facial feature points. This outputs basic data for recognizing specific individuals. An image processing library is used for this process.

[0384] Step 2:

[0385] The server uses the facial feature data extracted in Step 1 as input to perform behavioral and facial expression analysis. By analyzing the movement patterns in the video and capturing changes in facial expressions, it infers the emotional state of the subject. Using a motion analysis algorithm, it outputs these analysis results as a risk score.

[0386] Step 3:

[0387] The device captures the user's voice using a microphone and uses this as input for emotion analysis. The emotion engine analyzes the voice data and identifies emotional states such as joy, anger, sadness, and happiness. The output is data that represents the user's emotional state using numbers and tags.

[0388] Step 4:

[0389] The device collects additional operation logs and behavioral patterns and sends them to the server. By analyzing what actions the user performed, psychological tendencies are monitored in more detail. This data is input, and the server outputs an overall emotional assessment.

[0390] Step 5:

[0391] The server utilizes the output data from steps 2 and 4 to generate a warning notification. The notification content and wording are customized to match the user's psychological state, adjusting to encourage caution while reducing psychological burden. The final notification is output to the terminal and communicated to the user.

[0392] Step 6:

[0393] The server shares generated suspicious person information and sentiment data with the local network. This operation strengthens security measures throughout the region and enables coordinated crime prevention activities. The output information is provided to relevant agencies.

[0394] (Application Example 2)

[0395] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0396] Security systems are required not only to detect suspicious individuals but also to provide appropriate alarm notifications that take into account the user's emotional state. However, current systems struggle to respond flexibly to the user's psychological state, failing to achieve both suspicious individual detection and user reassurance. Therefore, the challenge is to enhance security functions while improving the user experience by providing alarm notifications that take the user's emotions into consideration.

[0397] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0398] In this invention, the server includes means for acquiring video information and performing facial recognition processing from the information; means for analyzing a person's movements and emotions and evaluating the degree of risk; means for detecting abnormalities in structures; means for adjusting the content of alarm notifications based on the user's emotional state; and means for linking and sharing suspicious person information with a community network. This makes it possible to adjust alarm notifications according to the user's emotional state, thereby strengthening security measures while maintaining the user's sense of security.

[0399] "Visual information" refers to visual data acquired from surveillance cameras and sensor devices, and this data is used for person identification and activity analysis.

[0400] "Face recognition processing" is a technology that detects the face of a specific person from acquired video information and analyzes related data.

[0401] "Analysis of actions and emotions" is a process that evaluates emotions and risks based on a person's behavior and facial expressions.

[0402] "Risk assessment" is the process of quantifying potential risks based on information obtained from analyzing a person's movements and facial expressions.

[0403] "Detection of structural abnormalities" is a technology that identifies security risks by detecting unusual conditions in physical facilities such as windows and doors.

[0404] "User emotional state" refers to the user's current psychological state, inferred from their voice, behavior, and how they operate the interface.

[0405] "Adjusting the content of alarm notifications" means customizing alarm and notification messages to appropriate content based on the user's emotional state.

[0406] A "community network" is an information-sharing platform that shares crime prevention information with relevant organizations and residents within the community to enhance safety throughout the region.

[0407] This invention utilizes video information and emotion data in a security system to adjust alarm notifications to users according to their emotions. Specifically, a server receives video information in real time from multiple surveillance cameras and sensor devices and performs face recognition processing. This identifies specific individuals and performs analysis of their actions and emotions. Amazon Rekognition is used for face recognition, and image processing libraries such as OpenCV and dlib are used for action and emotion analysis.

[0408] The server also detects anomalies in windows and doors using sensor devices attached to the structure. This anomaly information is combined with video information to assess the level of risk. The risk assessment uses a proprietary algorithm to quantify the risk and issues an alarm as needed.

[0409] The device collects the user's voice and behavioral data and uses Google Cloud Speech-to-Text to infer their emotional state. Based on this data, the server adjusts the content of alert notifications and generates messages tailored to the user's psychological state. For example, if the user is feeling anxious or stressed, the notification will be delivered in a soft voice and gentle language.

[0410] Furthermore, the server integrates suspicious person information with the community network, improving local safety through an information sharing platform. This strengthens crime prevention capabilities throughout the entire community.

[0411] As a concrete example, consider a scenario where a user is relaxing at home at night, but an anomaly is detected near the window. In this case, the device recognizes the user's anxiety, and the server sends a gentle notification such as, "Please remain calm and pay immediate attention to what is happening outside the window."

[0412] An example of a prompt for a generative AI model is: "When I feel uneasy while walking down the street, please suggest what kind of notification would be appropriate if a suspicious person is detected. Example: notification in a gentle voice, option to contact support."

[0413] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0414] Step 1:

[0415] The server acquires video information in real time from surveillance cameras. The input is video data from the surveillance cameras, and the output is processed video frames. Amazon Rekognition is used to identify faces of people in the video and pass that data to the next step.

[0416] Step 2:

[0417] The server analyzes the behavior and emotions of face-recognized video data using OpenCV and dlib. The input is face-recognized video frames, and the output is data indicating the user's behavior and emotional state. Through this analysis, the server evaluates emotions from the person's behavior patterns and facial expressions, and quantifies the level of danger.

[0418] Step 3:

[0419] The device collects user voice and motion data and converts it to text using Google Cloud Speech-to-Text. Inputs are audio files and sensor data, while outputs are transcribed audio data and motion analysis results. This allows the system to infer the user's emotional state and send that information to a server.

[0420] Step 4:

[0421] The server acquires information from sensor devices to detect structural anomalies and determines whether an anomaly exists. The input is data from the sensor devices, and the output is a flag indicating the presence or absence of an anomaly. If an anomaly is detected, it is reflected in the risk assessment.

[0422] Step 5:

[0423] The server adjusts the content of the alarm notification based on the user's emotional state and risk assessment. The input is emotional state data and risk score, and the output is the adjusted alarm notification message. A generative AI model is used to create appropriate notification content and send it to the device.

[0424] Step 6:

[0425] The terminal receives pre-configured alarm notifications from the server and notifies the user gently via voice or text message. The input is the alarm notification message from the server, and the output is the voice or visual alarm delivered to the user. This provides security while maintaining the user's sense of security.

[0426] Step 7:

[0427] The server shares information about suspicious individuals in conjunction with the community network. Inputs include information about suspicious individuals and network connectivity, while output is shared information to enhance local safety. This improves crime prevention awareness throughout the community.

[0428] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0429] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0430] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0431] [Third Embodiment]

[0432] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0433] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0434] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0435] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0436] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0437] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0438] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0439] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0440] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0441] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0442] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0443] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0444] The security system of the present invention is implemented in a manner that enables the detection of suspicious individuals and a rapid response through the functions of the server, terminal, and user.

[0445] The server continuously receives video data from multiple surveillance cameras. This data is processed in real time and compared against faces registered in an existing database via a facial recognition algorithm. If there is no match, the person is recorded as a suspicious individual. This facial recognition process makes it possible to identify suspicious individuals.

[0446] The terminal continuously monitors using sensors installed in structures such as windows and doors. When an anomaly is detected, it immediately issues an alarm based on instructions from the server and notifies the user of the relevant location. This notification is sent to the user's smartphone or tablet device.

[0447] Upon receiving a notification, users can immediately operate the application and access live video on the server. This allows users to check the situation on site even when away from home and make arrangements to contact the police or security company directly if necessary.

[0448] Furthermore, the server has the function of sharing analyzed information about suspicious individuals with the local crime prevention network. This information sharing can improve safety within the community and strengthen cooperation among multiple organizations. For example, when the server detects the same suspicious individual in the area, it immediately transmits the information so that nearby residents and cooperating organizations can respond quickly.

[0449] Thus, this invention ensures the safety of not only individual residences but also entire communities by combining multi-layered security measures. It achieves effective crime prevention functionality through a design that enhances user convenience and system responsiveness.

[0450] The following describes the processing flow.

[0451] Step 1:

[0452] The server continuously receives video data from the surveillance cameras and prepares to process that data in real time.

[0453] Step 2:

[0454] The server applies a facial recognition algorithm to detect faces in the video data. It then compares these faces with a registered database and identifies any unmatched faces as suspicious individuals.

[0455] Step 3:

[0456] The server analyzes a person's behavior and facial expressions and scores their level of risk. If specific abnormal behavior is detected, a high risk score is assigned.

[0457] Step 4:

[0458] The terminal constantly monitors sensors attached to the structure and detects abnormalities in windows and doors.

[0459] Step 5:

[0460] If an anomaly is detected, the terminal immediately issues an alarm, receives instructions from the server, and sends an alarm notification to the user.

[0461] Step 6:

[0462] Users receive alarm notifications on their smartphones or tablets, launch the application, and view real-time live video.

[0463] Step 7:

[0464] Establish a system that allows users to assess the situation and contact the police or security company as needed.

[0465] Step 8:

[0466] The server shares information about suspicious individuals with the local crime prevention network, coordinating efforts to strengthen security measures.

[0467] (Example 1)

[0468] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0469] In recent years, with the increasing importance of security, there has been a growing need for systems that can efficiently detect suspicious individuals and identify anomalies early. However, conventional systems have struggled to integrate multiple sensors and cameras, process information in real time, and appropriately share information about suspicious individuals. Furthermore, users have limited means to quickly confirm the situation and take emergency action. To address these challenges, the development of new security systems is needed.

[0470] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0471] In this invention, the server includes means for acquiring video data and performing face recognition processing on the data; means for comparing the acquired face features with an existing database and recording the person as a suspicious person if there is a mismatch; means for detecting anomalies using sensors installed on structures; means for enabling access to live video using a communication terminal; and means for transmitting suspicious person information to a shared platform. This enables real-time detection of suspicious persons and information sharing by integrating information from multiple data sources, and provides users with means to quickly confirm the situation and respond immediately as needed.

[0472] "Video data" refers to visual information acquired by surveillance cameras and other recording devices.

[0473] "Facial recognition processing" is an analytical method that detects facial features contained in video data and uses a specific algorithm to identify individuals.

[0474] A "database" is a collection of information organized according to specific rules, and is used to match retrieved data within a system.

[0475] A "suspicious person" refers to an individual who is not registered in the database or who has not been given prior permission.

[0476] A "detector" is a device used to detect changes in the state of structures such as windows and doors.

[0477] An "alarm" is a visual or auditory signal issued to draw attention when an abnormality is detected.

[0478] A "communication terminal" refers to an electronic device used by a user to receive or send information.

[0479] "Live video" refers to footage filmed in real time, providing information that shows the situation as it is currently unfolding.

[0480] A "shared infrastructure" is a platform for securely and efficiently communicating information between multiple organizations and devices.

[0481] This security system provides efficient intruder detection and rapid response to anomalies through the functions of the server, terminals, and users.

[0482] The server continuously acquires video data from multiple cameras. This video data is processed in real time and analyzed by a face recognition algorithm. Software libraries such as OpenCV and dlib are used here. The acquired face features are compared with an existing database, and if there is no match, the person is recorded as a suspicious individual. This record functions as a log for taking appropriate action depending on the situation.

[0483] The terminal uses sensors installed in windows and doors to detect any abnormalities. Examples of these sensors include magnetic sensors and infrared sensors, which can instantly detect unauthorized opening or movement of structures. When an abnormality is detected, the terminal issues an alarm. The alarm alerts the affected person with sound and light. Simultaneously, an alarm notification is sent to the user's communication terminal. Firebase Cloud Messaging can be used as the communication protocol here.

[0484] Based on the received notification, the user accesses live video footage provided by the server by operating a dedicated application on their communication terminal. This live video allows the user to check the situation at the scene remotely. Depending on the situation, the user can immediately contact the relevant authorities through the application's functions. For example, if an anomaly is notified while the user is away from home, they can check the scene from their mobile device and contact the security company directly.

[0485] Furthermore, the server transmits information recorded as suspicious individuals to a shared platform, allowing the information to be shared with relevant parties through the regional communication platform. This enables neighboring areas and related organizations to take similar actions in real time and cooperate with each other.

[0486] This system is designed not only to protect individual homes but also to improve the safety of the entire community. Furthermore, it is technically flexible and can be operated by combining various sensor and communication technologies. The following example prompt can be used in the generated AI model: "Please describe the process for detecting suspicious individuals and sharing information."

[0487] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0488] Step 1:

[0489] The server acquires video data from the monitoring device. The acquired video data is broken down into frames, and each frame is used as input to a face recognition algorithm. Specifically, it detects faces of people within the frame and extracts their features. The output is face feature data.

[0490] Step 2:

[0491] The server compares the extracted facial features with those stored in the existing database. In this step, the server compares the facial features with those stored in the database, and if there is a mismatch, the newly detected face is recorded as a suspicious person. Cosine similarity is used for similarity determination here. The output is the result of the determination as to whether or not to record the face as a suspicious person.

[0492] Step 3:

[0493] The terminal monitors data from sensors installed on windows and doors. The input is real-time open / close data transmitted from the sensors. When abnormal opening or closing is detected, that information is sent to the server. The output is an alarm message indicating that an abnormality has been detected.

[0494] Step 4:

[0495] The terminal receives instructions from the server and issues an alarm. The input is alarm information from the server, and the output is a physical alarm signal (sound or light). It also has the function of sending this information as a notification to communication terminals.

[0496] Step 5:

[0497] Users receive alarm notifications on their communication terminals. Upon receiving a notification, users can launch a dedicated application and access live video from the server. This live video is real-time data streamed from the server and serves as input data for users to check the situation.

[0498] Step 6:

[0499] The server transmits recorded information about suspicious individuals to a local shared platform. The input is the information recorded as a suspicious person, and the output is the shared information transmitted to the local communication platform. This allows for information sharing with neighboring areas and related organizations, enabling a rapid response.

[0500] (Application Example 1)

[0501] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0502] In modern society, personal safety and community security are critical issues. However, systems that enable the immediate identification of suspicious individuals, rapid information sharing, and appropriate responses remain inadequate. In particular, when a suspicious person is identified, quickly sharing that information within the local community and improving overall community safety is a challenging task.

[0503] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0504] In this invention, the server includes means for acquiring video data and performing facial recognition processing, means for sending alarm notifications to users using an information and communication network, and means for sharing suspicious person information in cooperation with the local community. This enables the rapid issuance of alarms when a suspicious person is detected in the area, immediate sharing of suspicious person data, and improvement of local safety.

[0505] "Video data" refers to visual information acquired from recording devices such as surveillance cameras.

[0506] "Face recognition processing" refers to the process of identifying the face of a specific person from acquired video data using an algorithm.

[0507] "Methods for scoring risk" refer to methods that analyze a person's behavior and facial expressions and quantify their level of risk.

[0508] "Means for detecting abnormalities in structures" refers to methods of checking for abnormalities using sensors attached to physical structures such as windows and doors.

[0509] An "information and communication network" refers to a system that sends and receives data via the internet or other digital communications.

[0510] "Alert notifications" refer to information that warns the user in response to the detection of anomalies or suspicious individuals.

[0511] A "local community" refers to a group of people and related organizations that live in a particular area.

[0512] "Suspicious person information" refers to data about suspicious individuals identified through facial recognition or other means.

[0513] "Information terminals" refer to electronic devices that users can carry with them and that can connect to a network, such as smartphones and tablets.

[0514] "Live video" refers to visual information that is filmed in real time and distributed via a network.

[0515] To implement this invention, specific hardware and software are required to build a security system. Specific embodiments are described below.

[0516] server

[0517] The server continuously receives video data from multiple surveillance cameras. This requires networked cameras and a server with real-time data processing capabilities. For facial recognition, software libraries capable of executing facial recognition algorithms, such as OpenCV or TensorFlow, are used. The server analyzes the video data using the facial recognition algorithm to identify suspicious individuals. This information about suspicious individuals is shared to improve the safety of the local community.

[0518] terminal

[0519] The terminal constantly monitors data from sensors installed on windows and doors. These sensors detect physical anomalies and issue an alarm when an anomaly is detected. Furthermore, based on instructions from a server, it sends warnings to the user's information terminal via the network. Communication technologies such as Wi-Fi and Bluetooth are used for information communication.

[0520] User

[0521] Users can access live video feeds on the server via smartphones and other information devices. This allows users to check the situation even when they are away from home and to make emergency contact with local communities and security agencies as needed. Users also play a role in quickly sharing information about suspicious individuals to ensure local safety.

[0522] Specific example

[0523] For example, if a suspicious person is detected in the area one night, the server immediately sends a push notification to the user's smartphone, allowing the user to view real-time video. This enables the user to quickly understand the situation and, if necessary, contact the police or local crime prevention organizations.

[0524] Example of a prompt

[0525] Please explain the mechanism of a system that identifies suspicious individuals from security camera footage and shares that information with local residents in real time. Please include specific technologies and procedures in detail.

[0526] This format is expected to significantly improve local safety and individual peace of mind.

[0527] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0528] Step 1:

[0529] The server receives video data from multiple surveillance cameras. This data is prepared for processing in real time. Once the video data arrives at the server, it is first input into a face recognition algorithm. The server analyzes each frame of the video and uses libraries such as OpenCV to identify clear facial features. As a result, a list of potential suspicious individuals is output.

[0530] Step 2:

[0531] The server compares the facial information identified by the facial recognition algorithm with an existing database. The facial data obtained as input is compared with known data in the database. If no matching entry is found in the database, the person is recorded as a suspicious person. This process completes the identification of the suspicious person, and the information is sent to the next step.

[0532] Step 3:

[0533] The terminal monitors signals from sensors installed on windows and doors. When an anomaly is detected, it processes the signal and records the location. This anomaly information is immediately sent to the server. Based on this information, the server generates instructions to issue an alarm at the location in question.

[0534] Step 4:

[0535] The server sends an alert notification to the user's information terminal based on anomaly and suspicious person information. Using the suspicious person and anomaly information obtained as input, it generates an integrated alert if both are related. This alert is delivered to the user via a push notification service such as Firebase Cloud Messaging. Upon receiving this notification, the user can immediately check the warning on their smartphone.

[0536] Step 5:

[0537] Users access the server using their information terminals to view live video. The server streams the video from the relevant surveillance camera using the RTSP protocol in response to the user's request. This allows users to understand the situation in real time and prepare to cooperate with the police or security organizations if necessary.

[0538] Step 6:

[0539] The server shares information about suspicious individuals with the local community's safety network. The entered data is sent to a database linked to local crime prevention organizations and residents. This information sharing raises crime prevention awareness throughout the community and enables residents to respond quickly.

[0540] The above outlines the system processing steps required to implement the application example. Each step works in conjunction with the others, aiming to enhance safety throughout the entire region.

[0541] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0542] This invention relates to a security system equipped with an emotion engine in addition to acquiring and processing video data. It recognizes the user's emotions in addition to detecting suspicious individuals, and allows for adjustment of the content of alarm notifications.

[0543] The server has the capability to process video data received from multiple surveillance cameras in real time. This processing uses facial recognition technology and behavioral / expression analysis technology to identify individuals and score their risk level. This allows for the immediate identification of individuals suspected of being intruders.

[0544] In addition, an emotion engine installed on the device constantly monitors the user's emotional state. This emotion engine infers emotions from the user's voice, behavioral patterns, and how they operate the device, and sends this data to a server. The server uses this emotional data to adjust alert notifications to match the user's psychological state. Specifically, if the server detects that the user is experiencing stress or anxiety, it can provide more reassuring content and adjust the frequency of notifications.

[0545] For example, if a window anomaly is detected while the user is showing signs of anxiety, the server will use gentle and considerate language in the notification and provide an option to contact a support desk that can respond quickly and professionally. The emotion engine also contributes to improving the accuracy of suspicious person information; by including emotion data when the server shares information in cooperation with the local network, it helps in developing safety measures based on the suspect's state and reactions.

[0546] Thus, the present invention provides an advanced type of security system that combines further enhancement of crime prevention functions with consideration for the psychological aspects of the user experience.

[0547] The following describes the processing flow.

[0548] Step 1:

[0549] The server receives video data from multiple surveillance cameras and prepares to process the data in real time.

[0550] Step 2:

[0551] The server applies a facial recognition algorithm to compare the faces of people in the video with existing data in the database. If a face is unfamiliar, it is marked and recorded as a suspicious person.

[0552] Step 3:

[0553] The server analyzes a person's behavior and facial expressions, and scores their level of danger based on the data obtained. If abnormal behavior is detected, a higher score is assigned.

[0554] Step 4:

[0555] The device monitors the user's emotional state using an emotion engine and infers emotions from the user's voice and operation patterns.

[0556] Step 5:

[0557] Emotional data from the emotion engine is sent to the server, which uses this data to adjust the content and method of alarm notifications.

[0558] Step 6:

[0559] If the server detects an anomaly using sensors installed on structures such as windows and doors, it immediately issues an alarm and sends an emotion-based, tailored notification to the user.

[0560] Step 7:

[0561] Users receive notifications on their smartphones or tablets and, depending on the situation, check live video to understand the current state of affairs.

[0562] Step 8:

[0563] If a user suspects an intruder, they can decide to contact a security company or the police and choose an appropriate support option that is sensitive to their feelings.

[0564] Step 9:

[0565] The server will share information on suspicious individuals and user sentiment data collected with the local network to enhance safety throughout the region.

[0566] (Example 2)

[0567] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0568] Modern security systems focus on detecting suspicious individuals but often neglect the psychological burden on users, with warning notifications sometimes causing anxiety and stress. Furthermore, there are limited means to improve the accuracy of suspicious person information and to efficiently share safety information with the community. To address these challenges, security features that consider the emotions of users are needed.

[0569] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0570] In this invention, the server includes means for acquiring video information and performing face recognition processing from the information, means for analyzing the target's behavior and facial expressions and assessing the degree of danger, and means for performing emotion analysis and adjusting the content of warning notifications based on the user's psychological state. This enables warning notifications that take into account the user's psychological state and improves crime prevention functions by enabling effective information sharing with the local community.

[0571] "Video information" refers to visual data captured by surveillance equipment. This data is used to understand the movements and circumstances of the monitored subject in real time.

[0572] "Facial recognition processing" refers to the technology that identifies human faces from video information, extracts their features, and uses them for identification. This technology is used in security systems to recognize individuals and match them with related data.

[0573] "Analyzing behavior and facial expressions" refers to the technique of interpreting underlying emotions and intentions from a subject's movement patterns and facial expressions. This technique is useful for predicting risky behavior and understanding psychological states.

[0574] "Assessing the level of risk" refers to the process of quantifying and evaluating potential risks based on analyzed behavior, facial expressions, and other sensor data.

[0575] "Sentiment analysis" refers to a technology that infers a user's emotional state from voice, text, and behavioral data. This technology makes it possible to adapt warning content according to the user's psychological state.

[0576] A "warning notification" refers to information issued to users in response to detected anomalies or risks. This information is typically communicated using visual and auditory means.

[0577] "Adjusting warning notification content based on psychological state" refers to the process of optimizing the wording and expression of warnings by taking into account the user's current emotions and stress level.

[0578] "Suspicious person information" refers to data on individuals or behaviors deemed to pose a potential threat detected within a monitored area. This information is used to strengthen crime prevention measures and to share information with the local community.

[0579] This invention utilizes advanced monitoring and emotion analysis technologies in security systems to provide warning notifications that take into account the user's psychological state.

[0580] The server acquires and processes video information in real time from multiple monitoring devices. The hardware used consists of network-connected cameras and recorders, and the data is efficiently managed through a stream processing system. The server processes this video information using general image processing libraries and APIs (for example, image analysis software for face recognition and software specialized for behavioral analysis) to perform face recognition and behavioral / facial expression analysis of subjects.

[0581] The device has an engine for sentiment analysis and is equipped with sensors for voice recognition and behavioral analysis. The device captures the user's voice and actions and analyzes their emotional state using a built-in sentiment analysis algorithm. This information is sent to the server in real time, and the sentiment data received by the server is used to adjust the content of warning notifications.

[0582] As a concrete example, consider a scenario where a user is at home and experiencing anxiety. When the monitoring system detects an anomaly in a window, it takes emotional data into account and quickly issues a warning using gentle, reassuring language. This process allows the user to understand the situation in a timely manner and respond without experiencing unnecessary stress.

[0583] As an example of a prompt message, we will use content such as, "Hello, how are you feeling right now? Please let us know if you have any comments or suggestions. We will use them to help us provide a safe and secure environment," to collect user feedback and improve the accuracy of the information in the security system.

[0584] In this way, in addition to regular warning notifications, the security system analyzes the user's emotional state, enabling more personalized notifications and responses, thereby providing a high level of safety and peace of mind.

[0585] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0586] Step 1:

[0587] The server starts a data stream using video information acquired in real time from monitoring equipment as input. Based on this video information, the server performs face recognition processing and extracts facial feature points. This outputs basic data for recognizing specific individuals. An image processing library is used for this process.

[0588] Step 2:

[0589] The server uses the facial feature data extracted in Step 1 as input to perform behavioral and facial expression analysis. By analyzing the movement patterns in the video and capturing changes in facial expressions, it infers the emotional state of the subject. Using a motion analysis algorithm, it outputs these analysis results as a risk score.

[0590] Step 3:

[0591] The device captures the user's voice using a microphone and uses this as input for emotion analysis. The emotion engine analyzes the voice data and identifies emotional states such as joy, anger, sadness, and happiness. The output is data that represents the user's emotional state using numbers and tags.

[0592] Step 4:

[0593] The device collects additional operation logs and behavioral patterns and sends them to the server. By analyzing what actions the user performed, psychological tendencies are monitored in more detail. This data is input, and the server outputs an overall emotional assessment.

[0594] Step 5:

[0595] The server utilizes the output data from steps 2 and 4 to generate a warning notification. The notification content and wording are customized to match the user's psychological state, adjusting to encourage caution while reducing psychological burden. The final notification is output to the terminal and communicated to the user.

[0596] Step 6:

[0597] The server shares generated suspicious person information and sentiment data with the local network. This operation strengthens security measures throughout the region and enables coordinated crime prevention activities. The output information is provided to relevant agencies.

[0598] (Application Example 2)

[0599] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0600] Security systems are required not only to detect suspicious individuals but also to provide appropriate alarm notifications that take into account the user's emotional state. However, current systems struggle to respond flexibly to the user's psychological state, failing to achieve both suspicious individual detection and user reassurance. Therefore, the challenge is to enhance security functions while improving the user experience by providing alarm notifications that take the user's emotions into consideration.

[0601] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0602] In this invention, the server includes means for acquiring video information and performing facial recognition processing from the information; means for analyzing a person's movements and emotions and evaluating the degree of risk; means for detecting abnormalities in structures; means for adjusting the content of alarm notifications based on the user's emotional state; and means for linking and sharing suspicious person information with a community network. This makes it possible to adjust alarm notifications according to the user's emotional state, thereby strengthening security measures while maintaining the user's sense of security.

[0603] "Visual information" refers to visual data acquired from surveillance cameras and sensor devices, and this data is used for person identification and activity analysis.

[0604] "Face recognition processing" is a technology that detects the face of a specific person from acquired video information and analyzes related data.

[0605] "Analysis of actions and emotions" is a process that evaluates emotions and risks based on a person's behavior and facial expressions.

[0606] "Risk assessment" is the process of quantifying potential risks based on information obtained from analyzing a person's movements and facial expressions.

[0607] "Detection of structural abnormalities" is a technology that identifies security risks by detecting unusual conditions in physical facilities such as windows and doors.

[0608] "User emotional state" refers to the user's current psychological state, inferred from their voice, behavior, and how they operate the interface.

[0609] "Adjusting the content of alarm notifications" means customizing alarm and notification messages to appropriate content based on the user's emotional state.

[0610] A "community network" is an information-sharing platform that shares crime prevention information with relevant organizations and residents within the community to enhance safety throughout the region.

[0611] This invention utilizes video information and emotion data in a security system to adjust alarm notifications to users according to their emotions. Specifically, a server receives video information in real time from multiple surveillance cameras and sensor devices and performs face recognition processing. This identifies specific individuals and performs analysis of their actions and emotions. Amazon Rekognition is used for face recognition, and image processing libraries such as OpenCV and dlib are used for action and emotion analysis.

[0612] The server also detects anomalies in windows and doors using sensor devices attached to the structure. This anomaly information is combined with video information to assess the level of risk. The risk assessment uses a proprietary algorithm to quantify the risk and issues an alarm as needed.

[0613] The device collects the user's voice and behavioral data and uses Google Cloud Speech-to-Text to infer their emotional state. Based on this data, the server adjusts the content of alert notifications and generates messages tailored to the user's psychological state. For example, if the user is feeling anxious or stressed, the notification will be delivered in a soft voice and gentle language.

[0614] Furthermore, the server integrates suspicious person information with the community network, improving local safety through an information sharing platform. This strengthens crime prevention capabilities throughout the entire community.

[0615] As a concrete example, consider a scenario where a user is relaxing at home at night, but an anomaly is detected near the window. In this case, the device recognizes the user's anxiety, and the server sends a gentle notification such as, "Please remain calm and pay immediate attention to what is happening outside the window."

[0616] An example of a prompt for a generative AI model is: "When I feel uneasy while walking down the street, please suggest what kind of notification would be appropriate if a suspicious person is detected. Example: notification in a gentle voice, option to contact support."

[0617] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0618] Step 1:

[0619] The server acquires video information in real time from surveillance cameras. The input is video data from the surveillance cameras, and the output is processed video frames. Amazon Rekognition is used to identify faces of people in the video and pass that data to the next step.

[0620] Step 2:

[0621] The server analyzes the behavior and emotions of face-recognized video data using OpenCV and dlib. The input is face-recognized video frames, and the output is data indicating the user's behavior and emotional state. Through this analysis, the server evaluates emotions from the person's behavior patterns and facial expressions, and quantifies the level of danger.

[0622] Step 3:

[0623] The device collects user voice and motion data and converts it to text using Google Cloud Speech-to-Text. Inputs are audio files and sensor data, while outputs are transcribed audio data and motion analysis results. This allows the system to infer the user's emotional state and send that information to a server.

[0624] Step 4:

[0625] The server acquires information from sensor devices to detect structural anomalies and determines whether an anomaly exists. The input is data from the sensor devices, and the output is a flag indicating the presence or absence of an anomaly. If an anomaly is detected, it is reflected in the risk assessment.

[0626] Step 5:

[0627] The server adjusts the content of the alarm notification based on the user's emotional state and risk assessment. The input is emotional state data and risk score, and the output is the adjusted alarm notification message. A generative AI model is used to create appropriate notification content and send it to the device.

[0628] Step 6:

[0629] The terminal receives pre-configured alarm notifications from the server and notifies the user gently via voice or text message. The input is the alarm notification message from the server, and the output is the voice or visual alarm delivered to the user. This provides security while maintaining the user's sense of security.

[0630] Step 7:

[0631] The server shares information about suspicious individuals in conjunction with the community network. Inputs include information about suspicious individuals and network connectivity, while output is shared information to enhance local safety. This improves crime prevention awareness throughout the community.

[0632] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0633] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0634] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0635] [Fourth Embodiment]

[0636] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0637] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0638] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0639] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0640] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0641] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0642] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0643] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0644] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0645] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0646] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0647] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0648] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0649] The security system of the present invention is implemented in a manner that enables the detection of suspicious individuals and a rapid response through the functions of the server, terminal, and user.

[0650] The server continuously receives video data from multiple surveillance cameras. This data is processed in real time and compared against faces registered in an existing database via a facial recognition algorithm. If there is no match, the person is recorded as a suspicious individual. This facial recognition process makes it possible to identify suspicious individuals.

[0651] The terminal continuously monitors using sensors installed in structures such as windows and doors. When an anomaly is detected, it immediately issues an alarm based on instructions from the server and notifies the user of the relevant location. This notification is sent to the user's smartphone or tablet device.

[0652] Upon receiving a notification, users can immediately operate the application and access live video on the server. This allows users to check the situation on site even when away from home and make arrangements to contact the police or security company directly if necessary.

[0653] Furthermore, the server has the function of sharing analyzed information about suspicious individuals with the local crime prevention network. This information sharing can improve safety within the community and strengthen cooperation among multiple organizations. For example, when the server detects the same suspicious individual in the area, it immediately transmits the information so that nearby residents and cooperating organizations can respond quickly.

[0654] Thus, this invention ensures the safety of not only individual residences but also entire communities by combining multi-layered security measures. It achieves effective crime prevention functionality through a design that enhances user convenience and system responsiveness.

[0655] The following describes the processing flow.

[0656] Step 1:

[0657] The server continuously receives video data from the surveillance cameras and prepares to process that data in real time.

[0658] Step 2:

[0659] The server applies a facial recognition algorithm to detect faces in the video data. It then compares these faces with a registered database and identifies any unmatched faces as suspicious individuals.

[0660] Step 3:

[0661] The server analyzes a person's behavior and facial expressions and scores their level of risk. If specific abnormal behavior is detected, a high risk score is assigned.

[0662] Step 4:

[0663] The terminal constantly monitors sensors attached to the structure and detects abnormalities in windows and doors.

[0664] Step 5:

[0665] If an anomaly is detected, the terminal immediately issues an alarm, receives instructions from the server, and sends an alarm notification to the user.

[0666] Step 6:

[0667] Users receive alarm notifications on their smartphones or tablets, launch the application, and view real-time live video.

[0668] Step 7:

[0669] Establish a system that allows users to assess the situation and contact the police or security company as needed.

[0670] Step 8:

[0671] The server shares information about suspicious individuals with the local crime prevention network, coordinating efforts to strengthen security measures.

[0672] (Example 1)

[0673] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0674] In recent years, with the increasing importance of security, there has been a growing need for systems that can efficiently detect suspicious individuals and identify anomalies early. However, conventional systems have struggled to integrate multiple sensors and cameras, process information in real time, and appropriately share information about suspicious individuals. Furthermore, users have limited means to quickly confirm the situation and take emergency action. To address these challenges, the development of new security systems is needed.

[0675] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0676] In this invention, the server includes means for acquiring video data and performing face recognition processing on the data; means for comparing the acquired face features with an existing database and recording the person as a suspicious person if there is a mismatch; means for detecting anomalies using sensors installed on structures; means for enabling access to live video using a communication terminal; and means for transmitting suspicious person information to a shared platform. This enables real-time detection of suspicious persons and information sharing by integrating information from multiple data sources, and provides users with means to quickly confirm the situation and respond immediately as needed.

[0677] "Video data" refers to visual information acquired by surveillance cameras and other recording devices.

[0678] "Facial recognition processing" is an analytical method that detects facial features contained in video data and uses a specific algorithm to identify individuals.

[0679] A "database" is a collection of information organized according to specific rules, and is used to match retrieved data within a system.

[0680] A "suspicious person" refers to an individual who is not registered in the database or who has not been given prior permission.

[0681] A "detector" is a device used to detect changes in the state of structures such as windows and doors.

[0682] An "alarm" is a visual or auditory signal issued to draw attention when an abnormality is detected.

[0683] A "communication terminal" refers to an electronic device used by a user to receive or send information.

[0684] "Live video" refers to footage filmed in real time, providing information that shows the situation as it is currently unfolding.

[0685] A "shared infrastructure" is a platform for securely and efficiently communicating information between multiple organizations and devices.

[0686] This security system provides efficient intruder detection and rapid response to anomalies through the functions of the server, terminals, and users.

[0687] The server continuously acquires video data from multiple cameras. This video data is processed in real time and analyzed by a face recognition algorithm. Software libraries such as OpenCV and dlib are used here. The acquired face features are compared with an existing database, and if there is no match, the person is recorded as a suspicious individual. This record functions as a log for taking appropriate action depending on the situation.

[0688] The terminal uses sensors installed in windows and doors to detect any abnormalities. Examples of these sensors include magnetic sensors and infrared sensors, which can instantly detect unauthorized opening or movement of structures. When an abnormality is detected, the terminal issues an alarm. The alarm alerts the affected person with sound and light. Simultaneously, an alarm notification is sent to the user's communication terminal. Firebase Cloud Messaging can be used as the communication protocol here.

[0689] Based on the received notification, the user accesses live video footage provided by the server by operating a dedicated application on their communication terminal. This live video allows the user to check the situation at the scene remotely. Depending on the situation, the user can immediately contact the relevant authorities through the application's functions. For example, if an anomaly is notified while the user is away from home, they can check the scene from their mobile device and contact the security company directly.

[0690] Furthermore, the server transmits information recorded as suspicious individuals to a shared platform, allowing the information to be shared with relevant parties through the regional communication platform. This enables neighboring areas and related organizations to take similar actions in real time and cooperate with each other.

[0691] This system is designed not only to protect individual homes but also to improve the safety of the entire community. Furthermore, it is technically flexible and can be operated by combining various sensor and communication technologies. The following example prompt can be used in the generated AI model: "Please describe the process for detecting suspicious individuals and sharing information."

[0692] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0693] Step 1:

[0694] The server acquires video data from the monitoring device. The acquired video data is broken down into frames, and each frame is used as input to a face recognition algorithm. Specifically, it detects faces of people within the frame and extracts their features. The output is face feature data.

[0695] Step 2:

[0696] The server compares the extracted facial features with those stored in the existing database. In this step, the server compares the facial features with those stored in the database, and if there is a mismatch, the newly detected face is recorded as a suspicious person. Cosine similarity is used for similarity determination here. The output is the result of the determination as to whether or not to record the face as a suspicious person.

[0697] Step 3:

[0698] The terminal monitors data from sensors installed on windows and doors. The input is real-time open / close data transmitted from the sensors. When abnormal opening or closing is detected, that information is sent to the server. The output is an alarm message indicating that an abnormality has been detected.

[0699] Step 4:

[0700] The terminal receives instructions from the server and issues an alarm. The input is alarm information from the server, and the output is a physical alarm signal (sound or light). It also has the function of sending this information as a notification to communication terminals.

[0701] Step 5:

[0702] Users receive alarm notifications on their communication terminals. Upon receiving a notification, users can launch a dedicated application and access live video from the server. This live video is real-time data streamed from the server and serves as input data for users to check the situation.

[0703] Step 6:

[0704] The server transmits recorded information about suspicious individuals to a local shared platform. The input is the information recorded as a suspicious person, and the output is the shared information transmitted to the local communication platform. This allows for information sharing with neighboring areas and related organizations, enabling a rapid response.

[0705] (Application Example 1)

[0706] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0707] In modern society, personal safety and community security are critical issues. However, systems that enable the immediate identification of suspicious individuals, rapid information sharing, and appropriate responses remain inadequate. In particular, when a suspicious person is identified, quickly sharing that information within the local community and improving overall community safety is a challenging task.

[0708] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0709] In this invention, the server includes means for acquiring video data and performing facial recognition processing, means for sending alarm notifications to users using an information and communication network, and means for sharing suspicious person information in cooperation with the local community. This enables the rapid issuance of alarms when a suspicious person is detected in the area, immediate sharing of suspicious person data, and improvement of local safety.

[0710] "Video data" refers to visual information acquired from recording devices such as surveillance cameras.

[0711] "Face recognition processing" refers to the process of identifying the face of a specific person from acquired video data using an algorithm.

[0712] "Methods for scoring risk" refer to methods that analyze a person's behavior and facial expressions and quantify their level of risk.

[0713] "Means for detecting abnormalities in structures" refers to methods of checking for abnormalities using sensors attached to physical structures such as windows and doors.

[0714] An "information and communication network" refers to a system that sends and receives data via the internet or other digital communications.

[0715] "Alert notifications" refer to information that warns the user in response to the detection of anomalies or suspicious individuals.

[0716] A "local community" refers to a group of people and related organizations that live in a particular area.

[0717] "Suspicious person information" refers to data about suspicious individuals identified through facial recognition or other means.

[0718] "Information terminals" refer to electronic devices that users can carry with them and that can connect to a network, such as smartphones and tablets.

[0719] "Live video" refers to visual information that is filmed in real time and distributed via a network.

[0720] To implement this invention, specific hardware and software are required to build a security system. Specific embodiments are described below.

[0721] server

[0722] The server continuously receives video data from multiple surveillance cameras. This requires networked cameras and a server with real-time data processing capabilities. For facial recognition, software libraries capable of executing facial recognition algorithms, such as OpenCV or TensorFlow, are used. The server analyzes the video data using the facial recognition algorithm to identify suspicious individuals. This information about suspicious individuals is shared to improve the safety of the local community.

[0723] terminal

[0724] The terminal constantly monitors data from sensors installed on windows and doors. These sensors detect physical anomalies and issue an alarm when an anomaly is detected. Furthermore, based on instructions from a server, it sends warnings to the user's information terminal via the network. Communication technologies such as Wi-Fi and Bluetooth are used for information communication.

[0725] User

[0726] Users can access live video feeds on the server via smartphones and other information devices. This allows users to check the situation even when they are away from home and to make emergency contact with local communities and security agencies as needed. Users also play a role in quickly sharing information about suspicious individuals to ensure local safety.

[0727] Specific example

[0728] For example, if a suspicious person is detected in the area one night, the server immediately sends a push notification to the user's smartphone, allowing the user to view real-time video. This enables the user to quickly understand the situation and, if necessary, contact the police or local crime prevention organizations.

[0729] Example of a prompt

[0730] Please explain the mechanism of a system that identifies suspicious individuals from security camera footage and shares that information with local residents in real time. Please include specific technologies and procedures in detail.

[0731] This format is expected to significantly improve local safety and individual peace of mind.

[0732] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0733] Step 1:

[0734] The server receives video data from multiple surveillance cameras. This data is prepared for processing in real time. Once the video data arrives at the server, it is first input into a face recognition algorithm. The server analyzes each frame of the video and uses libraries such as OpenCV to identify clear facial features. As a result, a list of potential suspicious individuals is output.

[0735] Step 2:

[0736] The server compares the facial information identified by the facial recognition algorithm with an existing database. The facial data obtained as input is compared with known data in the database. If no matching entry is found in the database, the person is recorded as a suspicious person. This process completes the identification of the suspicious person, and the information is sent to the next step.

[0737] Step 3:

[0738] The terminal monitors signals from sensors installed on windows and doors. When an anomaly is detected, it processes the signal and records the location. This anomaly information is immediately sent to the server. Based on this information, the server generates instructions to issue an alarm at the location in question.

[0739] Step 4:

[0740] The server sends an alert notification to the user's information terminal based on anomaly and suspicious person information. Using the suspicious person and anomaly information obtained as input, it generates an integrated alert if both are related. This alert is delivered to the user via a push notification service such as Firebase Cloud Messaging. Upon receiving this notification, the user can immediately check the warning on their smartphone.

[0741] Step 5:

[0742] Users access the server using their information terminals to view live video. The server streams the video from the relevant surveillance camera using the RTSP protocol in response to the user's request. This allows users to understand the situation in real time and prepare to cooperate with the police or security organizations if necessary.

[0743] Step 6:

[0744] The server shares information about suspicious individuals with the local community's safety network. The entered data is sent to a database linked to local crime prevention organizations and residents. This information sharing raises crime prevention awareness throughout the community and enables residents to respond quickly.

[0745] The above outlines the system processing steps required to implement the application example. Each step works in conjunction with the others, aiming to enhance safety throughout the entire region.

[0746] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0747] This invention relates to a security system equipped with an emotion engine in addition to acquiring and processing video data. It recognizes the user's emotions in addition to detecting suspicious individuals, and allows for adjustment of the content of alarm notifications.

[0748] The server has the capability to process video data received from multiple surveillance cameras in real time. This processing uses facial recognition technology and behavioral / expression analysis technology to identify individuals and score their risk level. This allows for the immediate identification of individuals suspected of being intruders.

[0749] In addition, an emotion engine installed on the device constantly monitors the user's emotional state. This emotion engine infers emotions from the user's voice, behavioral patterns, and how they operate the device, and sends this data to a server. The server uses this emotional data to adjust alert notifications to match the user's psychological state. Specifically, if the server detects that the user is experiencing stress or anxiety, it can provide more reassuring content and adjust the frequency of notifications.

[0750] For example, if a window anomaly is detected while the user is showing signs of anxiety, the server will use gentle and considerate language in the notification and provide an option to contact a support desk that can respond quickly and professionally. The emotion engine also contributes to improving the accuracy of suspicious person information; by including emotion data when the server shares information in cooperation with the local network, it helps in developing safety measures based on the suspect's state and reactions.

[0751] Thus, the present invention provides an advanced type of security system that combines further enhancement of crime prevention functions with consideration for the psychological aspects of the user experience.

[0752] The following describes the processing flow.

[0753] Step 1:

[0754] The server receives video data from multiple surveillance cameras and prepares to process the data in real time.

[0755] Step 2:

[0756] The server applies a facial recognition algorithm to compare the faces of people in the video with existing data in the database. If a face is unfamiliar, it is marked and recorded as a suspicious person.

[0757] Step 3:

[0758] The server analyzes a person's behavior and facial expressions, and scores their level of danger based on the data obtained. If abnormal behavior is detected, a higher score is assigned.

[0759] Step 4:

[0760] The device monitors the user's emotional state using an emotion engine and infers emotions from the user's voice and operation patterns.

[0761] Step 5:

[0762] Emotional data from the emotion engine is sent to the server, which uses this data to adjust the content and method of alarm notifications.

[0763] Step 6:

[0764] If the server detects an anomaly using sensors installed on structures such as windows and doors, it immediately issues an alarm and sends an emotion-based, tailored notification to the user.

[0765] Step 7:

[0766] Users receive notifications on their smartphones or tablets and, depending on the situation, check live video to understand the current state of affairs.

[0767] Step 8:

[0768] If a user suspects an intruder, they can decide to contact a security company or the police and choose an appropriate support option that is sensitive to their feelings.

[0769] Step 9:

[0770] The server will share information on suspicious individuals and user sentiment data collected with the local network to enhance safety throughout the region.

[0771] (Example 2)

[0772] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0773] Modern security systems focus on detecting suspicious individuals but often neglect the psychological burden on users, with warning notifications sometimes causing anxiety and stress. Furthermore, there are limited means to improve the accuracy of suspicious person information and to efficiently share safety information with the community. To address these challenges, security features that consider the emotions of users are needed.

[0774] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0775] In this invention, the server includes means for acquiring video information and performing face recognition processing from the information, means for analyzing the target's behavior and facial expressions and assessing the degree of danger, and means for performing emotion analysis and adjusting the content of warning notifications based on the user's psychological state. This enables warning notifications that take into account the user's psychological state and improves crime prevention functions by enabling effective information sharing with the local community.

[0776] "Video information" refers to visual data captured by surveillance equipment. This data is used to understand the movements and circumstances of the monitored subject in real time.

[0777] "Facial recognition processing" refers to the technology that identifies human faces from video information, extracts their features, and uses them for identification. This technology is used in security systems to recognize individuals and match them with related data.

[0778] "Analyzing behavior and facial expressions" refers to the technique of interpreting underlying emotions and intentions from a subject's movement patterns and facial expressions. This technique is useful for predicting risky behavior and understanding psychological states.

[0779] "Assessing the level of risk" refers to the process of quantifying and evaluating potential risks based on analyzed behavior, facial expressions, and other sensor data.

[0780] "Sentiment analysis" refers to a technology that infers a user's emotional state from voice, text, and behavioral data. This technology makes it possible to adapt warning content according to the user's psychological state.

[0781] A "warning notification" refers to information issued to users in response to detected anomalies or risks. This information is typically communicated using visual and auditory means.

[0782] "Adjusting warning notification content based on psychological state" refers to the process of optimizing the wording and expression of warnings by taking into account the user's current emotions and stress level.

[0783] "Suspicious person information" refers to data on individuals or behaviors deemed to pose a potential threat detected within a monitored area. This information is used to strengthen crime prevention measures and to share information with the local community.

[0784] This invention utilizes advanced monitoring and emotion analysis technologies in security systems to provide warning notifications that take into account the user's psychological state.

[0785] The server acquires and processes video information in real time from multiple monitoring devices. The hardware used consists of network-connected cameras and recorders, and the data is efficiently managed through a stream processing system. The server processes this video information using general image processing libraries and APIs (for example, image analysis software for face recognition and software specialized for behavioral analysis) to perform face recognition and behavioral / facial expression analysis of subjects.

[0786] The device has an engine for sentiment analysis and is equipped with sensors for voice recognition and behavioral analysis. The device captures the user's voice and actions and analyzes their emotional state using a built-in sentiment analysis algorithm. This information is sent to the server in real time, and the sentiment data received by the server is used to adjust the content of warning notifications.

[0787] As a concrete example, consider a scenario where a user is at home and experiencing anxiety. When the monitoring system detects an anomaly in a window, it takes emotional data into account and quickly issues a warning using gentle, reassuring language. This process allows the user to understand the situation in a timely manner and respond without experiencing unnecessary stress.

[0788] As an example of a prompt message, we will use content such as, "Hello, how are you feeling right now? Please let us know if you have any comments or suggestions. We will use them to help us provide a safe and secure environment," to collect user feedback and improve the accuracy of the information in the security system.

[0789] In this way, in addition to regular warning notifications, the security system analyzes the user's emotional state, enabling more personalized notifications and responses, thereby providing a high level of safety and peace of mind.

[0790] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0791] Step 1:

[0792] The server starts a data stream using video information acquired in real time from monitoring equipment as input. Based on this video information, the server performs face recognition processing and extracts facial feature points. This outputs basic data for recognizing specific individuals. An image processing library is used for this process.

[0793] Step 2:

[0794] The server uses the facial feature data extracted in Step 1 as input to perform behavioral and facial expression analysis. By analyzing the movement patterns in the video and capturing changes in facial expressions, it infers the emotional state of the subject. Using a motion analysis algorithm, it outputs these analysis results as a risk score.

[0795] Step 3:

[0796] The device captures the user's voice using a microphone and uses this as input for emotion analysis. The emotion engine analyzes the voice data and identifies emotional states such as joy, anger, sadness, and happiness. The output is data that represents the user's emotional state using numbers and tags.

[0797] Step 4:

[0798] The device collects additional operation logs and behavioral patterns and sends them to the server. By analyzing what actions the user performed, psychological tendencies are monitored in more detail. This data is input, and the server outputs an overall emotional assessment.

[0799] Step 5:

[0800] The server utilizes the output data from steps 2 and 4 to generate a warning notification. The notification content and wording are customized to match the user's psychological state, adjusting to encourage caution while reducing psychological burden. The final notification is output to the terminal and communicated to the user.

[0801] Step 6:

[0802] The server shares generated suspicious person information and sentiment data with the local network. This operation strengthens security measures throughout the region and enables coordinated crime prevention activities. The output information is provided to relevant agencies.

[0803] (Application Example 2)

[0804] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0805] Security systems are required not only to detect suspicious individuals but also to provide appropriate alarm notifications that take into account the user's emotional state. However, current systems struggle to respond flexibly to the user's psychological state, failing to achieve both suspicious individual detection and user reassurance. Therefore, the challenge is to enhance security functions while improving the user experience by providing alarm notifications that take the user's emotions into consideration.

[0806] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0807] In this invention, the server includes means for acquiring video information and performing facial recognition processing from the information; means for analyzing a person's movements and emotions and evaluating the degree of risk; means for detecting abnormalities in structures; means for adjusting the content of alarm notifications based on the user's emotional state; and means for linking and sharing suspicious person information with a community network. This makes it possible to adjust alarm notifications according to the user's emotional state, thereby strengthening security measures while maintaining the user's sense of security.

[0808] "Visual information" refers to visual data acquired from surveillance cameras and sensor devices, and this data is used for person identification and activity analysis.

[0809] "Face recognition processing" is a technology that detects the face of a specific person from acquired video information and analyzes related data.

[0810] "Analysis of actions and emotions" is a process that evaluates emotions and risks based on a person's behavior and facial expressions.

[0811] "Risk assessment" is the process of quantifying potential risks based on information obtained from analyzing a person's movements and facial expressions.

[0812] "Detection of structural abnormalities" is a technology that identifies security risks by detecting unusual conditions in physical facilities such as windows and doors.

[0813] "User emotional state" refers to the user's current psychological state, inferred from their voice, behavior, and how they operate the interface.

[0814] "Adjusting the content of alarm notifications" means customizing alarm and notification messages to appropriate content based on the user's emotional state.

[0815] A "community network" is an information-sharing platform that shares crime prevention information with relevant organizations and residents within the community to enhance safety throughout the region.

[0816] This invention utilizes video information and emotion data in a security system to adjust alarm notifications to users according to their emotions. Specifically, a server receives video information in real time from multiple surveillance cameras and sensor devices and performs face recognition processing. This identifies specific individuals and performs analysis of their actions and emotions. Amazon Rekognition is used for face recognition, and image processing libraries such as OpenCV and dlib are used for action and emotion analysis.

[0817] The server also detects anomalies in windows and doors using sensor devices attached to the structure. This anomaly information is combined with video information to assess the level of risk. The risk assessment uses a proprietary algorithm to quantify the risk and issues an alarm as needed.

[0818] The device collects the user's voice and behavioral data and uses Google Cloud Speech-to-Text to infer their emotional state. Based on this data, the server adjusts the content of alert notifications and generates messages tailored to the user's psychological state. For example, if the user is feeling anxious or stressed, the notification will be delivered in a soft voice and gentle language.

[0819] Furthermore, the server integrates suspicious person information with the community network, improving local safety through an information sharing platform. This strengthens crime prevention capabilities throughout the entire community.

[0820] As a concrete example, consider a scenario where a user is relaxing at home at night, but an anomaly is detected near the window. In this case, the device recognizes the user's anxiety, and the server sends a gentle notification such as, "Please remain calm and pay immediate attention to what is happening outside the window."

[0821] An example of a prompt for a generative AI model is: "When I feel uneasy while walking down the street, please suggest what kind of notification would be appropriate if a suspicious person is detected. Example: notification in a gentle voice, option to contact support."

[0822] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0823] Step 1:

[0824] The server acquires video information in real time from surveillance cameras. The input is video data from the surveillance cameras, and the output is processed video frames. Amazon Rekognition is used to identify faces of people in the video and pass that data to the next step.

[0825] Step 2:

[0826] The server analyzes the behavior and emotions of face-recognized video data using OpenCV and dlib. The input is face-recognized video frames, and the output is data indicating the user's behavior and emotional state. Through this analysis, the server evaluates emotions from the person's behavior patterns and facial expressions, and quantifies the level of danger.

[0827] Step 3:

[0828] The device collects user voice and motion data and converts it to text using Google Cloud Speech-to-Text. Inputs are audio files and sensor data, while outputs are transcribed audio data and motion analysis results. This allows the system to infer the user's emotional state and send that information to a server.

[0829] Step 4:

[0830] The server acquires information from sensor devices to detect structural anomalies and determines whether an anomaly exists. The input is data from the sensor devices, and the output is a flag indicating the presence or absence of an anomaly. If an anomaly is detected, it is reflected in the risk assessment.

[0831] Step 5:

[0832] The server adjusts the content of the alarm notification based on the user's emotional state and risk assessment. The input is emotional state data and risk score, and the output is the adjusted alarm notification message. A generative AI model is used to create appropriate notification content and send it to the device.

[0833] Step 6:

[0834] The terminal receives pre-configured alarm notifications from the server and notifies the user gently via voice or text message. The input is the alarm notification message from the server, and the output is the voice or visual alarm delivered to the user. This provides security while maintaining the user's sense of security.

[0835] Step 7:

[0836] The server shares information about suspicious individuals in conjunction with the community network. Inputs include information about suspicious individuals and network connectivity, while output is shared information to enhance local safety. This improves crime prevention awareness throughout the community.

[0837] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0838] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0839] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0840] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0841] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0842] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0843] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0844] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0845] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0846] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0847] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0848] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0849] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0850] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0851] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0852] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0853] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0854] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0855] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0856] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0857] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0858] The following is further disclosed regarding the embodiments described above.

[0859] (Claim 1)

[0860] A means for acquiring video data and performing face recognition processing on said data,

[0861] A method for analyzing a person's behavior and facial expressions and scoring their level of danger,

[0862] A means for detecting abnormalities in structures such as windows and doors,

[0863] A means of issuing an alarm and sending an alarm notification to the user,

[0864] A means of sharing information about suspicious individuals,

[0865] A system that includes this.

[0866] (Claim 2)

[0867] The system according to claim 1, wherein the user can view live video using a mobile device.

[0868] (Claim 3)

[0869] The system according to claim 1, which links and shares information about suspicious persons with a local network.

[0870] "Example 1"

[0871] (Claim 1)

[0872] A means for acquiring video data and performing face recognition processing on said data,

[0873] A method for comparing acquired facial features with an existing database and recording a person as suspicious if there is a mismatch,

[0874] A means of detecting abnormalities using sensors installed on a structure,

[0875] A means of issuing an alarm and sending an alarm notification to a communication terminal,

[0876] A means of enabling access to live video using a communication terminal,

[0877] A means of transmitting information about suspicious persons to a shared platform,

[0878] A system that includes this.

[0879] (Claim 2)

[0880] The system according to claim 1, wherein a customer uses a mobile device to view live video.

[0881] (Claim 3)

[0882] The system according to claim 1, which links and shares information about suspicious persons with a community exchange platform.

[0883] "Application Example 1"

[0884] (Claim 1)

[0885] A means for acquiring video data and performing face recognition processing on said data,

[0886] A method for analyzing a person's behavior and facial expressions and scoring their level of danger,

[0887] A means for detecting abnormalities in structures such as windows and doors,

[0888] A means of issuing an alarm using an information and communication network and sending an alarm notification to the user,

[0889] A means of sharing information about suspicious individuals in cooperation with the local community,

[0890] A means of delivering video in real time to the user's information terminal when a suspicious person is detected,

[0891] A support system that allows users to use information terminals to share information about suspicious individuals with organizations and make direct contact,

[0892] A system that includes this.

[0893] (Claim 2)

[0894] The system according to claim 1, which allows users to view live video using an information terminal and improves their sense of social security.

[0895] (Claim 3)

[0896] The system according to claim 1, which links information on suspicious persons with a local safety network to enable efficient information sharing.

[0897] "Example 2 of combining an emotion engine"

[0898] (Claim 1)

[0899] A means for acquiring video information and performing face recognition processing from said information,

[0900] A means of analyzing the subject's behavior and facial expressions to assess the degree of danger,

[0901] A means for detecting abnormalities in entrances, exits, openings, etc.

[0902] A means of issuing a warning and sending a warning notification to the user,

[0903] A means of performing emotional analysis and adjusting the content of warning notifications based on the user's psychological state,

[0904] A means of integrating and sharing information on suspicious individuals and emotional data,

[0905] A system that includes this.

[0906] (Claim 2)

[0907] The system according to claim 1, wherein a user checks real-time video using a mobile device.

[0908] (Claim 3)

[0909] The system according to claim 1, which links and shares information about suspicious persons with a regional network.

[0910] "Application example 2 when combining with an emotional engine"

[0911] (Claim 1)

[0912] A means for acquiring video information and performing face recognition processing from said information,

[0913] A means of analyzing a person's actions and emotions and evaluating the degree of risk,

[0914] Means for detecting abnormalities in structures,

[0915] A means of issuing an alarm and sending an alarm notification to the user,

[0916] A means of adjusting the content of alarm notifications based on the user's emotional state,

[0917] A means of sharing information about suspicious individuals in cooperation with a community network,

[0918] A system that includes this.

[0919] (Claim 2)

[0920] The system according to claim 1, wherein a user views real-time video using a mobile device.

[0921] (Claim 3)

[0922] The system according to claim 1, which links and shares information about suspicious persons with a local network. [Explanation of symbols]

[0923] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring video data and performing face recognition processing on said data, A method for analyzing a person's behavior and facial expressions and scoring their level of danger, A means for detecting abnormalities in structures such as windows and doors, A means of issuing an alarm and sending an alarm notification to the user, A means of sharing information about suspicious individuals, A system that includes this.

2. The system according to claim 1, wherein the user can view live video using a mobile device.

3. The system according to claim 1, which links and shares information about suspicious persons with a local network.