System

The integration of security cameras with generative AI allows for real-time detection and response to suspicious behavior, preventing crimes through audio warnings and alerts to security staff and residents.

JP2026014956APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116430
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional security camera systems are unable to detect suspicious behavior in real time and respond immediately, making it difficult to take appropriate measures before a crime occurs.

Method used

A system that integrates security cameras with generative AI to analyze video data for suspicious behavior, output audio warnings, and transmit information to security staff and residents, using computer vision and machine learning algorithms for 24-hour surveillance.

Benefits of technology

Enables early detection and prevention of crimes by identifying suspicious individuals and alerting relevant parties, enhancing safety in various locations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014956000001_ABST
    Figure 2026014956000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a means for receiving video data from a security camera, a means for analyzing the video data and detecting suspicious behavior, a means for outputting a voice alarm when a suspicious person is detected, and a means for transmitting suspicious person information to a security staff or a resident.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional security camera systems simply record video footage and are unable to detect suspicious behavior in real time and respond immediately. This makes it difficult to take appropriate measures before a crime occurs. The present invention aims to solve this problem by combining security cameras with generative AI and provide a system that prevents crime before it occurs. [Means for solving the problem]

[0005] The present invention is a system that includes the following means: means for receiving video data from security cameras, means for analyzing the video data and detecting suspicious behavior, means for outputting an audio warning when a suspicious individual is detected, and means for transmitting information about the suspicious individual to security staff and residents. Furthermore, the system uses computer vision technology and machine learning algorithms to detect specific behavioral patterns from the video data, and transmits an alert message containing information about the suspicious individual, including video capture, location information, and a danger level, thereby realizing 24-hour surveillance and preventing various crimes.

[0006] A "security camera" is a device installed to monitor and record suspicious activity.

[0007] "Video data" refers to information about images and videos captured by security cameras.

[0008] "Suspicious behavior" refers to behavior that differs from normal patterns of behavior and is considered to be potentially criminal.

[0009] "Generative AI" refers to artificial intelligence systems that use computer vision techniques and machine learning algorithms to analyze data and generate insights.

[0010] "Audio warning" refers to an audio warning message that is immediately sent to a suspicious person.

[0011] "Security staff" refers to the specialized personnel who monitor security camera systems.

[0012] "Residents" refers to people who live in apartment buildings or detached houses.

[0013] "Computer vision technology" refers to the technology that allows computers to extract and understand information from images and videos.

[0014] A "machine learning algorithm" refers to a computational rule that learns from data and makes predictions and judgments.

[0015] "Video capture" refers to the means of storing and displaying image data from a specific moment in time.

[0016] "Location Information" means geographic coordinates or other location data used to identify a particular location.

[0017] "Danger level" refers to an indicator that shows the degree of potential danger posed by suspicious behavior.

[0018] "Alert message" refers to a notification message sent to notify of the presence of a suspicious person.

[0019] "24-hour system" means that the system is always operational. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be explained.

[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0028] [First embodiment]

[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0041] The present invention provides a system that combines security cameras and generative AI to detect suspicious individuals and prevent crimes before they occur.

[0042] System Configuration

[0043] 1. Security camera and server integration

[0044] The server receives video data from security cameras in real time and uses computer vision techniques and machine learning algorithms to process this video data.

[0045] 2. Analysis of video data

[0046] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns. For example, it detects people who stay in a particular location for a long time or people who leave their luggage behind. This analysis uses technologies such as YOLO (You Only Look Once) and OpenCV.

[0047] 3. Identifying suspicious individuals and storing their information

[0048] When a suspicious person is detected, the server stores the information in a database, including the person's location, time, and behavioral patterns. This information provides information for later use in making appropriate decisions.

[0049] 4. Audio warning

[0050] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which plays an audio warning such as "Security is being strengthened." This audio warning is effective in alerting not only the suspicious person but also other people in the area.

[0051] 5. Sending alerts

[0052] When a suspicious person is detected, the server sends an alert to security staff and residents. The alert includes a video capture of the suspicious person, their location, and a danger level, allowing relevant parties to take immediate action.

[0053] Specific examples

[0054] For example, suppose footage from a security camera installed at a station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server determines that Person A's behavior is abnormal and immediately saves the information. The server then sends a command to the speaker next to the camera, which plays a warning audio message saying "Security is being strengthened." At the same time, an alert is sent to the security staff's device, providing location information and danger level along with the video capture. This allows security staff to respond to the scene quickly.

[0055] This system operates 24 hours a day and can be used in a variety of locations, including stations, airports, stores, apartment buildings, and detached homes. It is expected to make a significant contribution to crime prevention.

[0056] Natural language description of the program

[0057] The system's program analyzes video data sent from security cameras in real time to detect suspicious individuals. Specifically, the server receives the video data and analyzes it using a machine learning algorithm. If suspicious activity is detected, the server issues an audio warning through a speaker and simultaneously sends an alert to security staff and residents. The alert includes a video capture of the suspicious individual, their location, and a danger level.

[0058] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[0059] The processing flow will be explained below.

[0060] Step 1:

[0061] The server receives video data from the security camera in real time, specifically by continuously capturing video frames using a streaming protocol (e.g., RTSP).

[0062] Step 2:

[0063] The server analyzes the received video data frame by frame. Specifically, the video frames are input into a machine learning algorithm (such as YOLO or OpenCV) to identify the position and movement of people.

[0064] Step 3:

[0065] The server detects specific patterns of behavior or suspicious activity by using anomaly detection algorithms to identify behavior that differs from normal patterns (such as prolonged stays, rapid movements, or unnatural changes in location).

[0066] Step 4:

[0067] The server determines who is suspected of being a suspicious person by evaluating the detected behavioral patterns and identifying those who meet pre-defined criteria (e.g., staying for a long time or leaving luggage behind).

[0068] Step 5:

[0069] The server stores the suspect's information in a database, including details of location, time, and behavior, and uses it as data for future responses.

[0070] Step 6:

[0071] The server sends an audio warning to the speaker installed in the location where the suspicious person was detected. Specifically, it connects to the speaker controller and sends an instruction to play the warning audio, "Security being strengthened, security being strengthened."

[0072] Step 7:

[0073] The server then sends alerts to security staff and residents, including video capture of the suspicious individual, their location, and a risk level, allowing the relevant parties to initiate a rapid response.

[0074] Step 8:

[0075] Security staff and residents receive an alert and rush to the scene to identify and respond to suspicious individuals. Specifically, they conduct on-site investigations and take necessary measures based on the information contained in the alert.

[0076] Example 1

[0077] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0078] Conventional security camera systems have difficulty detecting suspicious individuals in real time, making it difficult to take immediate action to prevent crimes. Furthermore, they lack a means to accurately and quickly communicate suspicious individual information to relevant parties, which can result in delays in appropriate response. The present invention aims to solve these problems and provide a system for efficiently and effectively detecting and responding to suspicious individuals.

[0079] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0080] In this invention, the server includes means for receiving video data from security cameras, means for preprocessing the video data, means for analyzing the preprocessed video data and identifying suspicious behavior patterns, means for storing information on a suspicious individual in a database when the suspicious individual is detected, means for outputting an audio warning when the suspicious individual is detected, and means for transmitting information on the suspicious individual to security personnel and users. This enables early detection of suspicious individuals and prompt response, making it possible to prevent crimes from occurring.

[0081] A "security camera" is a device installed to capture video within a surveillance area and deter crimes and suspicious behavior.

[0082] "Video Data" means a digital representation of visual information captured by a security camera.

[0083] "Preprocessing" refers to initial processing such as adjusting resolution and removing noise, which is performed to make video data easier to analyze.

[0084] A "behavioral pattern" is an identified series of actions or characteristics of behavior taken by an object in video data.

[0085] A "database" is a system for efficiently storing, managing, and retrieving information.

[0086] A "voice warning" is a voice message that alerts suspicious people and people in the area.

[0087] An "alert message" is a notification containing information about a suspicious person, and is sent promptly to security personnel and users.

[0088] "Image processing technology" refers to a set of computer techniques for analyzing, transforming, and understanding digital images.

[0089] A "learning algorithm" is an algorithm that learns patterns and knowledge from data and makes predictions and classifications.

[0090] "Suspicious person information" is detailed information such as the suspicious person's behavior, location, time, and danger level.

[0091] "Location data" is information that indicates the geographical location of an object or person.

[0092] The "risk level" is an index that indicates the degree of risk based on the behavior of a suspicious person and the situation.

[0093] MODE FOR CARRYING OUT THE INVENTION

[0094] The purpose of this invention is to detect suspicious individuals and prevent crimes before they occur by using a system that combines security cameras and generative AI. This system consists of security cameras, a server, a speaker, and a terminal.

[0095] The server receives video data from security cameras in real time. It is always connected via the network and receives the video data captured by the security cameras sequentially. A multipurpose digital signal processing device is used to analyze the video data. At this stage, the video data is unprocessed raw data.

[0096] The server then preprocesses the received video data. This preprocessing involves using image processing techniques such as the OpenCV library to adjust the resolution and remove noise. Specifically, the video is converted to grayscale and a noise reduction filter is applied. This makes subsequent analysis easier.

[0097] After preprocessing is complete, the server performs object detection and behavioral pattern analysis using YOLO (You Only Look Once) or similar object detection algorithms. The analysis is performed frame by frame to identify people who linger for long periods of time or who behave abnormally.

[0098] If a suspicious person is detected, the server stores the information in a database, including the person's location, time, and behavioral patterns. For example, it connects to an SQL database and stores the information using an INSERT statement.

[0099] The server also issues an audio warning to detected suspicious individuals. It sends an HTTP request to the speaker and uses a speech synthesis library to output an audio warning such as "Security is being strengthened." This audio warning alerts not only the suspicious individual but also those around them.

[0100] The server then sends an alert message to the security personnel or the user's device, which includes a video capture of the suspicious person, their location, and the danger level. For example, the alert message can be sent via email or SMS API (e.g., Twilio).

[0101] Specific examples

[0102] For example, video data from a security camera installed at a station is sent to a server, and after analysis, person A is detected behaving suspiciously near a bench. The server determines that person A's behavior is abnormal and stores this information in a database. The server then issues an instruction to a speaker to play a warning audio message such as "Security is being strengthened." At the same time, an alert is sent to the security officer's device, providing a video capture, location information, and danger level. This allows security officers to respond to the scene quickly.

[0103] Example prompts for generative AI models

[0104] 1. "How can I analyze security camera footage at a station to detect suspicious individuals?"

[0105] 2. "How can I build an automation system that detects suspicious activity from security cameras and sends audio warnings and alerts?"

[0106] As described above, this invention utilizes security cameras and generative AI to enable early detection of suspicious individuals and rapid response. Using this system can greatly contribute to crime prevention in a variety of locations.

[0107] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0108] Program processing steps

[0109] Step 1:

[0110] The server receives video data from security cameras in real time. It takes the video stream (e.g., RTSP) from the security cameras as input and stores the video data in the server's memory. The output at this stage is the raw video data stored in memory.

[0111] Step 2:

[0112] The server preprocesses the received video data. The input is the video data saved in step 1, and the server uses the OpenCV library to adjust the resolution and remove noise. Specifically, it converts the video to grayscale and applies a noise reduction filter. The output is the preprocessed video data.

[0113] Step 3:

[0114] The server analyzes the preprocessed video data to identify objects and behavioral patterns. The input is the preprocessed video data created in step 2, which is analyzed using an object detection algorithm such as YOLO. Specifically, the video data is analyzed frame by frame to identify the location and movement of people and luggage. The output is data on detected objects and behavioral patterns.

[0115] Step 4:

[0116] When a suspicious person is detected, the server stores the information in a database. The input is the suspicious person information detected in step 3, and the server connects to an SQL database and uses the INSERT statement to store information such as the suspicious person's location, time, and behavioral patterns. The output is a new record stored in the database.

[0117] Step 5:

[0118] The server issues an audio warning if a suspicious person is detected. The input is the suspicious person information detected in step 3, and an HTTP request is sent to the speaker, which uses a Python speech synthesis library to play the audio message "Security is being strengthened." The output is an audio warning emitted from the speaker.

[0119] Step 6:

[0120] The server sends suspicious person information to the device of the security officer or user. The input is the suspicious person information detected in step 3, and an alert message is sent using email or an SMS API (for example, Twilio). The alert includes a video capture image, location information, and danger level. The output is an alert message that is displayed on the device of the security officer or user.

[0121] The above are the specific processing steps of this system. We have explained in detail how the input data is processed and what kind of output is obtained at each step. We have also clearly stated what specific operations are performed at each step. This makes the overall flow and function of this system clear.

[0122] (Application example 1)

[0123] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0124] Ensuring safety in autonomous vehicles is an important issue. In particular, there is a need to detect suspicious or dangerous behavior in the vehicle early and take appropriate measures promptly. However, conventional methods make it difficult to detect suspicious behavior in the vehicle in real time and respond immediately. For this reason, a new system is needed to improve safety in autonomous vehicles.

[0125] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0126] In this invention, the server includes means for receiving video data from security cameras, means for analyzing the video data and detecting suspicious behavior, means for outputting an audio warning when a suspicious individual is detected, means for transmitting suspicious individual information to security personnel and residents, means for receiving video data from an in-vehicle camera of the autonomous vehicle and detecting suspicious behavior of passengers, means for issuing an audio warning from an in-vehicle speaker, and means for transmitting an alert to an operation control center including video capture of the suspicious individual, location information, and danger level. This enables suspicious behavior in the autonomous vehicle to be quickly detected and an immediate response to be taken.

[0127] A "security camera" is a recording device that captures video data within a specific area and is used for safety and crime prevention purposes.

[0128] "Video data" refers to images and video information obtained from security cameras, in-car cameras, etc.

[0129] "Analysis" refers to the methods and techniques used to process video data frame by frame and understand and distinguish its content.

[0130] "Suspicious activity" refers to movements or behavior that deviate from normal patterns of behavior and may indicate dangerous or criminal activity.

[0131] A "voice warning" is a voice message that is emitted from a speaker to alert and warn people when suspicious behavior is detected.

[0132] "Security personnel" refers to professional people with safety and security responsibilities.

[0133] "Residents" refers to people who live in a particular building or area.

[0134] An "in-vehicle camera" is a camera device installed inside an autonomous vehicle to monitor passengers and the situation inside the vehicle.

[0135] "Passenger" refers to people using the self-driving vehicle.

[0136] "Operation control center" means a central control facility or organization for monitoring and managing the operation status of autonomous vehicles.

[0137] "Video capture" refers to obtaining a still image of a video at a specific moment.

[0138] "Location information" is data that indicates the geographic location of a particular object or person.

[0139] The "risk level" is a scale for evaluating the degree of danger of detected suspicious behavior or individuals.

[0140] An "alert" is a notification system for issuing warnings and alerts.

[0141] To realize this invention, the following system configuration and program are required. First, video data from security cameras and in-car cameras is received in real time by a server. The server analyzes this video data and uses machine learning algorithms to detect suspicious behavior or individuals. If suspicious behavior or individuals are detected, the server immediately and automatically takes appropriate action.

[0142] The system uses the following hardware and software:

[0143] In-car cameras and security cameras (any camera device)

[0144] Server (high performance computer)

[0145] OpenCV (image processing library)

[0146] YOLO (You Only Look Once) (machine learning model)

[0147] Warning audio playback system (speaker)

[0148] Alert sending system (network connection)

[0149] The server analyzes the video data sent from the camera in real time, analyzing each frame using the YOLO model, which is a technology that detects people and other objects in images quickly and accurately. Using this model, specific behavioral patterns and abnormal behavior can be quickly identified.

[0150] If suspicious behavior is detected, the server plays an audio warning through the speaker and simultaneously sends an alert to the operation control center. The alert includes a video capture of the suspicious person, their location, and a danger level, allowing the operation control center staff to immediately check the situation and take appropriate measures.

[0151] As a concrete example, consider its use in an autonomous vehicle. For example, if an in-car camera detects suspicious activity over a long period of time near the driver's seat, the server analyzes the information and issues an audio warning such as "Security is being strengthened," while simultaneously sending an alert to the operation control center containing details of the suspicious activity. This alert includes a video capture of the scene, the vehicle's location, and a danger level, allowing the administrator to take prompt action.

[0152] In addition, specific examples of prompts for system design using generative AI models are as follows:

[0153] "We need to analyze the in-car footage captured by the camera and detect suspicious individuals. Please help us develop a system that analyzes passenger behavior using the YOLO model and maintains safety."

[0154] As described above, the present invention provides an effective means for enhancing passenger safety in autonomous vehicles, enabling early detection of suspicious behavior and rapid response, thereby ensuring safety inside the vehicle.

[0155] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0156] Step 1:

[0157] The server receives video data from security cameras and in-car cameras in real time.

[0158] Input: Video stream from a camera device.

[0159] Output: Real-time video data.

[0160] How it works: The camera device captures video and sends the data over the network to a server, which receives it and prepares it for analysis.

[0161] Step 2:

[0162] The server analyzes the received video data frame by frame to detect suspicious behavior or people.

[0163] Input: Real-time video data.

[0164] Output: Detection results of suspicious behavior or people.

[0165] How it works: The server uses the YOLO model to analyze video frames, identify people and objects in each frame, and record any specific behavioral patterns or abnormal behaviors.

[0166] Step 3:

[0167] If suspicious behavior is detected, the server obtains video capture and location information to form suspicious person information.

[0168] Input: Suspicious behavior or person detection results.

[0169] Output: Video capture, location information, suspicious person information.

[0170] Specific operation: The server captures frames of any detected suspicious behavior or person and obtains location information such as GPS data. This information is then integrated to generate information about the suspicious person.

[0171] Step 4:

[0172] When a suspicious person is detected, the server plays a warning sound from the car's speakers.

[0173] Input: Suspicious person information.

[0174] Output: Warning audio.

[0175] Specific operation: The server sends instructions to the speaker to play a pre-prepared warning sound such as "Security is being strengthened," thereby alerting passengers and suspicious people inside the vehicle.

[0176] Step 5:

[0177] The server sends an alert containing information about the suspicious person to the operation control center.

[0178] Input: Suspicious person information.

[0179] Output: The alert message.

[0180] Specific operation: The server sends an alert message to the operation control center, including a video capture of the suspicious person, their location, and the danger level. Operation control personnel can use this information to take prompt action.

[0181] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0182] The present invention provides a system that combines security cameras, generative AI, and an emotion engine to recognize the user's emotions, detect suspicious individuals, and prevent crimes before they occur.

[0183] System Configuration

[0184] 1. Security camera and server integration

[0185] The server receives video data from security cameras in real time and uses computer vision techniques, machine learning algorithms, and an emotion engine to process the video data.

[0186] 2. Analysis of video data

[0187] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns or suspicious movements. For example, it can detect people who stay in a particular location for a long time or people who leave their luggage behind. This analysis uses technologies such as YOLO (You Only Look Once) and OpenCV.

[0188] 3. Emotional Recognition

[0189] The server analyzes the user's facial expressions from the video data and estimates their emotions. This emotion analysis uses an emotion engine (e.g., Facial Emotion Recognition). This allows it to identify whether the user is feeling anxiety, fear, anger, or other emotions.

[0190] 4. Identifying suspicious individuals and storing their information

[0191] When a suspicious individual is detected, the server stores the information in a database, including the individual's location, time, behavioral patterns, and perceived emotions, providing information for later appropriate response.

[0192] 5. Audio warning transmission and adjustment

[0193] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which plays an audio warning such as "Security is being strengthened." This warning not only alerts the suspicious person, but also other people in the area. Furthermore, the content and tone of the audio warning are adjusted appropriately based on the user's emotions recognized by the emotion engine.

[0194] 6. Sending Alerts

[0195] The server sends alerts to security staff and residents when a suspicious individual is detected, including video captures of the individual, their location, danger level, and perceived emotion, allowing relevant parties to take immediate action.

[0196] Specific examples

[0197] For example, suppose that footage from a security camera installed at a train station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server then detects the emotion of fear from Person A's facial expression. Based on this, the server determines that not only is Person A's behavior abnormal, but that fear is also behind that behavior.

[0198] This information is stored in a database, and the server sends a command to a speaker next to the camera, playing a warning message saying "security is being strengthened" and adjusting the tone of the message to a gentler tone. At the same time, an alert is sent to security staff, informing them that the video capture includes location information, danger level, and fear emotion. This allows security staff to respond to the scene quickly and with the underlying fear factor in mind.

[0199] In this way, by combining emotion engines, it is possible to create a system that can prevent crimes while also responding appropriately to the situation.

[0200] Natural language description of the program

[0201] The system's program analyzes video data sent from security cameras in real time to not only detect suspicious individuals but also analyze user emotions. Specifically, the server receives the video data and analyzes it using machine learning algorithms and an emotion engine. If suspicious behavior or a specific emotion is detected, the server plays an audio warning through a speaker and simultaneously sends an alert to security staff and residents. The alert includes the suspicious individual's video capture, location information, danger level, and recognized emotion.

[0202] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[0203] The processing flow will be explained below.

[0204] Step 1:

[0205] The server receives video data from the security camera in real time, specifically by continuously capturing video frames using a streaming protocol (e.g., RTSP).

[0206] Step 2:

[0207] The server analyzes the received video data frame by frame. Specifically, the video frames are input into a machine learning algorithm (such as YOLO or OpenCV) to identify the position and movement of people.

[0208] Step 3:

[0209] The server detects specific patterns of behavior or suspicious activity by using anomaly detection algorithms to identify behavior that differs from normal patterns (such as prolonged stays, rapid movements, or unnatural changes in location).

[0210] Step 4:

[0211] The server analyzes the user's facial expressions from the video data and estimates their emotions. Specifically, it uses an emotion engine to identify emotions such as anxiety, fear, and anger from the user's facial expressions.

[0212] Step 5:

[0213] The server determines whether a person is suspicious. Specifically, it evaluates both behavioral patterns and emotions to identify suspicious individuals who meet pre-defined criteria (e.g., long stay + emotion of fear).

[0214] Step 6:

[0215] The server stores the suspect's information in a database, including location, time, behavioral details, and perceived emotions, for use as data for future responses.

[0216] Step 7:

[0217] The server then sends an audio warning to a speaker installed in the location where a suspicious person was detected. Specifically, it connects to a speaker controller and sends instructions to play a warning voice saying, "Security being strengthened, security being strengthened." The tone may be adjusted based on emotion.

[0218] Step 8:

[0219] The server sends alerts to security staff and residents, including video capture of the suspicious individual, their location, danger level, and perceived emotion, allowing relevant parties to initiate a rapid response.

[0220] Step 9:

[0221] Security staff and residents receive an alert and rush to the scene to identify and respond to suspicious individuals. Specifically, they conduct on-site investigations and take necessary measures based on the information contained in the alert.

[0222] Example 2

[0223] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0224] While conventional security systems can detect suspicious behavior by analyzing video data, they face the challenge of being unable to analyze user emotions and respond appropriately. Furthermore, there are insufficient means for notifying security staff and residents of suspicious person information in real time, making it difficult to respond quickly. Furthermore, there is a lack of efficient and accurate means for storing and managing suspicious person information.

[0225] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving video data from a security camera, means for analyzing the video data and detecting suspicious behavior, means for analyzing a user's emotion from the video data, means for outputting an audio warning when a suspicious person is detected, means for transmitting suspicious person information to security staff or residents, means for saving the suspicious person information in a database, means for dividing the video data into frames and performing preprocessing, means for detecting suspicious behavior using a machine learning algorithm, means for estimating a user's emotion using an emotion engine, and means for adjusting the tone of the audio warning as appropriate. This enables the security system to comprehensively analyze the behavior and emotion of a suspicious person and take prompt and appropriate action.

[0226] A "security camera" is a device that monitors a specific area and records and transmits the video data.

[0227] "Video data" is digital data consisting of a series of image frames captured by a security camera.

[0228] A "server" is a computer system for storing, analyzing, and processing data.

[0229] "Analysis" is the process of extracting and evaluating specific information from acquired video data.

[0230] "Suspicious activity" refers to specific unusual or suspicious movements or behaviors in security camera footage.

[0231] "User" refers to a general user, including a person who is the subject of surveillance by a security camera and an administrator who uses the system.

[0232] "Voice warning" is a voice message played through a speaker to alert those around you.

[0233] "Security Staff" means the person or group responsible for monitoring and managing the security system.

[0234] "Resident" means a person who lives in or uses the area in which the security system is installed.

[0235] A "database" is a system for efficiently storing, managing, and retrieving information.

[0236] "Frame segmentation" is the process of breaking down video data into individual still images.

[0237] "Preprocessing" is the process of improving the image quality and removing noise from video data before analysis.

[0238] A "machine learning algorithm" is a learning program that analyzes data and recognizes patterns.

[0239] The "emotion engine" is software that analyzes a user's facial expressions from video data and estimates their emotions.

[0240] "Tone adjustment" is the process of changing the tone and content of an audio alert depending on the situation.

[0241] MODE FOR CARRYING OUT THE INVENTION

[0242] The present invention provides a system that combines security cameras, generative AI models, and an emotion engine to recognize user emotions, detect suspicious individuals, and prevent crimes before they occur.

[0243] System Configuration

[0244] This invention combines various technologies centered around the server, terminal, and user.

[0245] 1. Security camera and server integration

[0246] The server receives video data from security cameras in real time and uses computer vision techniques, machine learning algorithms, and an emotion engine to process the video data.

[0247] 2. Analysis of video data

[0248] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns and suspicious movements, using technologies such as YOLO (You Only Look Once) and OpenCV.

[0249] 3. Emotional Recognition

[0250] The server analyzes the user's facial expressions from the video data and estimates their emotions. This emotion analysis uses an emotion engine (e.g., Facial Emotion Recognition).

[0251] 4. Identifying suspicious individuals and storing their information

[0252] If a suspicious individual is detected, the server stores the information in a database, including the individual's location, time, behavioral patterns, and perceived emotions.

[0253] 5. Audio warning transmission and adjustment

[0254] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which issues an audio warning such as "Security is being strengthened." The content and tone of the audio warning are adjusted appropriately based on the user's emotions recognized by the emotion engine.

[0255] 6. Sending Alerts

[0256] The server sends alerts to security staff and residents when a suspicious individual is detected, including video captures of the individual, their location, danger level, and perceived emotion.

[0257] Specific examples

[0258] For example, suppose that footage from a security camera installed in a station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server then detects the emotion of fear from Person A's facial expression. Based on this, the server determines that not only is Person A's behavior abnormal, but that fear is also behind that behavior.

[0259] This information is stored in a database, and the server sends a command to the speaker next to the camera, which plays a warning message saying "security is being strengthened" and adjusts the tone of the message to a gentler tone. At the same time, an alert is sent to security staff, informing them that the video capture includes location information, danger level, and fear emotion. This allows security staff to respond to the scene quickly and take into account the underlying fear factor.

[0260] Prompt Sentence Examples

[0261] Based on the system log data, an analysis of security camera footage detected Person A behaving suspiciously near the bench. Furthermore, the emotion of fear was detected. Using this information, what information should be provided to the security staff, and what action should be taken next?

[0262] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[0263] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0264] Step 1:

[0265] The server receives video data from security cameras in real time. Specifically, the server receives video streams through a specific IP address and port. The input is the video data from the security cameras, and the output is the video frames stored in the server.

[0266] Step 2:

[0267] The server divides the received video data into frames, extracts the frames using the OpenCV library, and simultaneously performs preprocessing such as noise reduction and image quality improvement. The input is the raw video stream data, and the output is the processed frame images.

[0268] Step 3:

[0269] The server loads a machine learning algorithm (e.g., YOLO) and performs object recognition within each frame. Specifically, it uses the YOLO model to identify people and objects and analyze specific movement and behavior patterns. The input is the preprocessed frame image, and the output is the identified objects and their locations.

[0270] Step 4:

[0271] The server uses an emotion engine (e.g., Facial Emotion Recognition) to analyze the facial expressions of the detected person and estimate their emotions. This allows us to identify the emotions the user is feeling (e.g., anxiety, fear, anger). The input is a person's facial image, and the output is estimated emotional information.

[0272] Step 5:

[0273] The server stores the information of the detected suspicious person (location, time, behavioral patterns, and recognized emotions) in a database. Specifically, it inserts data into a MySQL database. The input is the suspicious person information, and the output is the success / failure status of saving the data to the database.

[0274] Step 6:

[0275] The server sends instructions to the speaker based on the suspicious person information, and issues an audio warning saying "Security is being strengthened." The tone and content of the audio are adjusted based on the results of emotion analysis. The input is suspicious person information and emotion information, and the output is the execution of the audio warning.

[0276] Step 7:

[0277] The server sends alerts to security staff and residents, including video captures of suspicious individuals, their location, danger level, and perceived emotions. The alerts are sent via email or SMS. The input is information about the suspicious individual and their emotions, and the output is the sending of an alert message.

[0278] (Application example 2)

[0279] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0280] In modern society, detecting crimes and suspicious individuals is a major issue. Although security cameras are commonly installed, their ability to detect suspicious individuals and analyze abnormal behavior in real time is insufficient, and there are few systems that can recognize emotions. As a result, it is difficult for users and security staff to take prompt countermeasures, and effective crime prevention measures are needed.

[0281] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0282] In this invention, the server includes means for receiving video data from security cameras, means for analyzing the video data and detecting suspicious behavior, means for analyzing facial expressions from the video data and recognizing emotions, means for sending a push notification to a smartphone when a suspicious person is detected, means for outputting an audio warning when a suspicious person is detected, and means for sending suspicious person information to security staff and residents. This enables real-time detection of suspicious people and recognition of emotions, enabling rapid and effective response.

[0283] A "security camera" is a device that records and transmits video in real time to monitor for crimes and suspicious activity.

[0284] "Video data" is visual information of a scene or object at a particular time obtained from a security camera.

[0285] "Analysis" is the process of using computer vision techniques and machine learning algorithms to identify specific behavioral patterns and anomalous behavior from video data.

[0286] "Suspicious behavior" is behavior that is abnormal to general patterns of behavior and indicates potential criminal activity or security risks.

[0287] "Facial expression" refers to the expression of emotions that can be read from the movement and arrangement of muscles in a person's face.

[0288] "Emotion recognition" is a technology that analyzes people's facial expressions from video data and estimates emotions such as joy, anger, and fear.

[0289] "Push notification" is a technology that allows specific information to be instantly displayed on the screen of a user's device, such as a smartphone.

[0290] A "voice warning" is a method of issuing an audio warning through a speaker to alert suspicious individuals and those around them.

[0291] "Suspicious person information" is data about a person who has engaged in suspicious behavior, and includes video capture, location information, emotion recognition results, danger level, and the like.

[0292] "Security staff" are people who monitor and patrol buildings and areas to ensure their safety.

[0293] This invention is a system that combines a security camera, generative AI, and an emotion recognition engine, and aims to recognize the user's emotions, detect suspicious individuals, and prevent crimes before they occur.

[0294] System Configuration

[0295] This system consists of the following elements:

[0296] 1. Security camera and server integration: The server receives video data from security cameras in real time. This video data is processed on the server. The hardware used can be a standard IP camera or a high-resolution security camera.

[0297] 2. Video data analysis: The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns and suspicious movements. This analysis uses computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., YOLO).

[0298] 3. Emotion Recognition: The server analyzes the user's facial expressions from the video data and estimates their emotions. For emotion analysis, an emotion recognition engine (e.g., Facial Emotion Recognition) is used. This allows the server to identify whether the user is feeling emotions such as anxiety, fear, or anger.

[0299] 4. Suspicious person detection and alert sending: If the server detects a suspicious person, it stores the information in a database and sends a push notification to the smartphone. It can also output an audio warning through the speaker. This notification includes a video capture of the suspicious person, their location, emotion recognition results, and danger level.

[0300] What the program does

[0301] The server processes video data received from security cameras in real time. First, it uses OpenCV to analyze the video data frame by frame to detect suspicious behavior. Next, it uses YOLO to identify specific behavioral patterns. At the same time, it uses an emotion recognition engine to analyze facial expressions from the video data and classify the user's emotions.

[0302] If a suspicious individual is detected, the information is stored in a database and a push notification is sent to the smartphone application. An audio warning is also issued from the speaker based on instructions from the server. The suspicious individual information includes video capture, location information, emotion recognition results, and danger level, allowing security staff and users to take prompt and appropriate action.

[0303] Specific examples

[0304] For example, if a home security camera detects a person staying in a specific location for a long period of time and recognizes the emotion of fear from the person's facial expression, the server will issue a voice warning in a gentle tone from the speaker saying "Security is being strengthened." At the same time, a push notification will be sent to the smartphone, containing information such as "Location: (104, 230, 150, 276), Emotion: Fear."

[0305] Prompt Sentence Examples

[0306] 1. "Security cameras analyze footage in real time to detect people who stay in the area for long periods of time."

[0307] 2. "Build a system that estimates emotions from facial expressions and issues audio warnings in the event of anxiety or fear."

[0308] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0309] Step 1:

[0310] Receive video data from security cameras

[0311] The server receives video data from security cameras in real time. This video data is the input for the entire system. Image data is sent to the server frame by frame, allowing it to proceed to the next analysis step.

[0312] Step 2:

[0313] Video data analysis

[0314] The server analyzes the received video data frame by frame. This analysis uses computer vision technology (OpenCV) and machine learning algorithms (YOLO). It detects people and objects in the frames and identifies suspicious movements and behavior patterns. The input is the video data for each frame, and the output is information on the location and behavior patterns of detected suspicious individuals.

[0315] Step 3:

[0316] Facial Expression Analysis and Emotion Recognition

[0317] The server analyzes the facial expressions of the detected person and uses an emotion recognition engine to recognize their emotions. Specifically, it cuts out the face region and inputs it into the Facial Emotion Recognition model. The data input used here is the image data of the detected face, and the output is the determined emotion (e.g., fear, anger, joy, etc.).

[0318] Step 4:

[0319] Saving suspicious person information

[0320] The server stores information about the detected suspicious individuals (location, behavioral patterns, emotions, etc.) in a database. This information is later used for analysis and countermeasures. The data input is the analysis result, and the record in the database is the output.

[0321] Step 5:

[0322] Sending push notifications

[0323] If a suspicious person is detected, the server sends a push notification to the smartphone. This notification includes the suspicious person's video capture, location information, emotion recognition results, danger level, etc. The data input is the stored suspicious person information, and the output is the notification sent to the smartphone.

[0324] Step 6:

[0325] Sending a voice warning

[0326] The server sends instructions to the speaker to play an audio warning such as "Security is being strengthened." The content and tone of the audio warning are adjusted based on the recognized emotion. The data input is the analysis results and emotion recognition results, and the output is the audio warning emitted from the speaker.

[0327] Step 7:

[0328] Sending alerts

[0329] The server sends alerts containing suspicious person information to security staff and residents. The alerts include video capture, location information, emotion recognition results, and danger level. The data input is the stored suspicious person information, and the output is the sending of an alert message.

[0330] In this way, the system can detect suspicious individuals in real time and take necessary measures immediately.

[0331] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0332] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0333] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0334] [Second embodiment]

[0335] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0336] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0337] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0338] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0339] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0340] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0341] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0342] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0343] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0344] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0345] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0346] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0347] The present invention provides a system that combines security cameras and generative AI to detect suspicious individuals and prevent crimes before they occur.

[0348] System Configuration

[0349] 1. Security camera and server integration

[0350] The server receives video data from security cameras in real time and uses computer vision techniques and machine learning algorithms to process this video data.

[0351] 2. Analysis of video data

[0352] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns. For example, it detects people who stay in a particular location for a long time or people who leave their luggage behind. This analysis uses technologies such as YOLO (You Only Look Once) and OpenCV.

[0353] 3. Identifying suspicious individuals and storing their information

[0354] When a suspicious person is detected, the server stores the information in a database, including the person's location, time, and behavioral patterns. This information provides information for later use in making appropriate decisions.

[0355] 4. Audio warning

[0356] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which plays an audio warning such as "Security is being strengthened." This audio warning is effective in alerting not only the suspicious person but also other people in the area.

[0357] 5. Sending alerts

[0358] When a suspicious person is detected, the server sends an alert to security staff and residents. The alert includes a video capture of the suspicious person, their location, and a danger level, allowing relevant parties to take immediate action.

[0359] Specific examples

[0360] For example, suppose footage from a security camera installed at a station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server determines that Person A's behavior is abnormal and immediately saves the information. The server then sends a command to the speaker next to the camera, which plays a warning audio message saying "Security is being strengthened." At the same time, an alert is sent to the security staff's device, providing location information and danger level along with the video capture. This allows security staff to respond to the scene quickly.

[0361] This system operates 24 hours a day and can be used in a variety of locations, including stations, airports, stores, apartment buildings, and detached homes. It is expected to make a significant contribution to crime prevention.

[0362] Natural language description of the program

[0363] The system's program analyzes video data sent from security cameras in real time to detect suspicious individuals. Specifically, the server receives the video data and analyzes it using a machine learning algorithm. If suspicious activity is detected, the server issues an audio warning through a speaker and simultaneously sends an alert to security staff and residents. The alert includes a video capture of the suspicious individual, their location, and a danger level.

[0364] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[0365] The processing flow will be explained below.

[0366] Step 1:

[0367] The server receives video data from the security camera in real time, specifically by continuously capturing video frames using a streaming protocol (e.g., RTSP).

[0368] Step 2:

[0369] The server analyzes the received video data frame by frame. Specifically, the video frames are input into a machine learning algorithm (such as YOLO or OpenCV) to identify the position and movement of people.

[0370] Step 3:

[0371] The server detects specific patterns of behavior or suspicious activity by using anomaly detection algorithms to identify behavior that differs from normal patterns (such as prolonged stays, rapid movements, or unnatural changes in location).

[0372] Step 4:

[0373] The server determines who is suspected of being a suspicious person by evaluating the detected behavioral patterns and identifying those who meet pre-defined criteria (e.g., staying for a long time or leaving luggage behind).

[0374] Step 5:

[0375] The server stores the suspect's information in a database, including details of location, time, and behavior, and uses it as data for future responses.

[0376] Step 6:

[0377] The server sends an audio warning to the speaker installed in the location where the suspicious person was detected. Specifically, it connects to the speaker controller and sends an instruction to play the warning audio, "Security being strengthened, security being strengthened."

[0378] Step 7:

[0379] The server then sends alerts to security staff and residents, including video capture of the suspicious individual, their location, and a risk level, allowing the relevant parties to initiate a rapid response.

[0380] Step 8:

[0381] Security staff and residents receive an alert and rush to the scene to identify and respond to suspicious individuals. Specifically, they conduct on-site investigations and take necessary measures based on the information contained in the alert.

[0382] Example 1

[0383] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0384] Conventional security camera systems have difficulty detecting suspicious individuals in real time, making it difficult to take immediate action to prevent crimes. Furthermore, they lack a means to accurately and quickly communicate suspicious individual information to relevant parties, which can result in delays in appropriate response. The present invention aims to solve these problems and provide a system for efficiently and effectively detecting and responding to suspicious individuals.

[0385] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0386] In this invention, the server includes means for receiving video data from security cameras, means for preprocessing the video data, means for analyzing the preprocessed video data and identifying suspicious behavior patterns, means for storing information on a suspicious individual in a database when the suspicious individual is detected, means for outputting an audio warning when the suspicious individual is detected, and means for transmitting information on the suspicious individual to security personnel and users. This enables early detection of suspicious individuals and prompt response, making it possible to prevent crimes from occurring.

[0387] A "security camera" is a device installed to capture video within a surveillance area and deter crimes and suspicious behavior.

[0388] "Video Data" means a digital representation of visual information captured by a security camera.

[0389] "Preprocessing" refers to initial processing such as adjusting resolution and removing noise, which is performed to make video data easier to analyze.

[0390] A "behavioral pattern" is an identified series of actions or characteristics of behavior taken by an object in video data.

[0391] A "database" is a system for efficiently storing, managing, and retrieving information.

[0392] A "voice warning" is a voice message that alerts suspicious people and people in the area.

[0393] An "alert message" is a notification containing information about a suspicious person, and is sent promptly to security personnel and users.

[0394] "Image processing technology" refers to a set of computer techniques for analyzing, transforming, and understanding digital images.

[0395] A "learning algorithm" is an algorithm that learns patterns and knowledge from data and makes predictions and classifications.

[0396] "Suspicious person information" is detailed information such as the suspicious person's behavior, location, time, and danger level.

[0397] "Location data" is information that indicates the geographical location of an object or person.

[0398] The "risk level" is an index that indicates the degree of risk based on the behavior of a suspicious person and the situation.

[0399] MODE FOR CARRYING OUT THE INVENTION

[0400] The purpose of this invention is to detect suspicious individuals and prevent crimes before they occur by using a system that combines security cameras and generative AI. This system consists of security cameras, a server, a speaker, and a terminal.

[0401] The server receives video data from security cameras in real time. It is always connected via the network and receives the video data captured by the security cameras sequentially. A multipurpose digital signal processing device is used to analyze the video data. At this stage, the video data is unprocessed raw data.

[0402] The server then preprocesses the received video data. This preprocessing involves using image processing techniques such as the OpenCV library to adjust the resolution and remove noise. Specifically, the video is converted to grayscale and a noise reduction filter is applied. This makes subsequent analysis easier.

[0403] After preprocessing is complete, the server performs object detection and behavioral pattern analysis using YOLO (You Only Look Once) or similar object detection algorithms. The analysis is performed frame by frame to identify people who linger for long periods of time or who behave abnormally.

[0404] If a suspicious person is detected, the server stores the information in a database, including the person's location, time, and behavioral patterns. For example, it connects to an SQL database and stores the information using an INSERT statement.

[0405] The server also issues an audio warning to detected suspicious individuals. It sends an HTTP request to the speaker and uses a speech synthesis library to output an audio warning such as "Security is being strengthened." This audio warning alerts not only the suspicious individual but also those around them.

[0406] The server then sends an alert message to the security personnel or the user's device, which includes a video capture of the suspicious person, their location, and the danger level. For example, the alert message can be sent via email or SMS API (e.g., Twilio).

[0407] Specific examples

[0408] For example, video data from a security camera installed at a station is sent to a server, and after analysis, person A is detected behaving suspiciously near a bench. The server determines that person A's behavior is abnormal and stores this information in a database. The server then issues an instruction to a speaker to play a warning audio message such as "Security is being strengthened." At the same time, an alert is sent to the security officer's device, providing a video capture, location information, and danger level. This allows security officers to respond to the scene quickly.

[0409] Example prompts for generative AI models

[0410] 1. "How can I analyze security camera footage at a station to detect suspicious individuals?"

[0411] 2. "How can I build an automation system that detects suspicious activity from security cameras and sends audio warnings and alerts?"

[0412] As described above, this invention utilizes security cameras and generative AI to enable early detection of suspicious individuals and rapid response. Using this system can greatly contribute to crime prevention in a variety of locations.

[0413] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0414] Program processing steps

[0415] Step 1:

[0416] The server receives video data from security cameras in real time. It takes the video stream (e.g., RTSP) from the security cameras as input and stores the video data in the server's memory. The output at this stage is the raw video data stored in memory.

[0417] Step 2:

[0418] The server preprocesses the received video data. The input is the video data saved in step 1, and the server uses the OpenCV library to adjust the resolution and remove noise. Specifically, it converts the video to grayscale and applies a noise reduction filter. The output is the preprocessed video data.

[0419] Step 3:

[0420] The server analyzes the preprocessed video data to identify objects and behavioral patterns. The input is the preprocessed video data created in step 2, which is analyzed using an object detection algorithm such as YOLO. Specifically, the video data is analyzed frame by frame to identify the location and movement of people and luggage. The output is data on detected objects and behavioral patterns.

[0421] Step 4:

[0422] When a suspicious person is detected, the server stores the information in a database. The input is the suspicious person information detected in step 3, and the server connects to an SQL database and uses the INSERT statement to store information such as the suspicious person's location, time, and behavioral patterns. The output is a new record stored in the database.

[0423] Step 5:

[0424] The server issues an audio warning if a suspicious person is detected. The input is the suspicious person information detected in step 3, and an HTTP request is sent to the speaker, which uses a Python speech synthesis library to play the audio message "Security is being strengthened." The output is an audio warning emitted from the speaker.

[0425] Step 6:

[0426] The server sends suspicious person information to the device of the security officer or user. The input is the suspicious person information detected in step 3, and an alert message is sent using email or an SMS API (for example, Twilio). The alert includes a video capture image, location information, and danger level. The output is an alert message that is displayed on the device of the security officer or user.

[0427] The above are the specific processing steps of this system. We have explained in detail how the input data is processed and what kind of output is obtained at each step. We have also clearly stated what specific operations are performed at each step. This makes the overall flow and function of this system clear.

[0428] (Application example 1)

[0429] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0430] Ensuring safety in autonomous vehicles is an important issue. In particular, there is a need to detect suspicious or dangerous behavior in the vehicle early and take appropriate measures promptly. However, conventional methods make it difficult to detect suspicious behavior in the vehicle in real time and respond immediately. For this reason, a new system is needed to improve safety in autonomous vehicles.

[0431] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0432] In this invention, the server includes means for receiving video data from security cameras, means for analyzing the video data and detecting suspicious behavior, means for outputting an audio warning when a suspicious individual is detected, means for transmitting suspicious individual information to security personnel and residents, means for receiving video data from an in-vehicle camera of the autonomous vehicle and detecting suspicious behavior of passengers, means for issuing an audio warning from an in-vehicle speaker, and means for transmitting an alert to an operation control center including video capture of the suspicious individual, location information, and danger level. This enables suspicious behavior in the autonomous vehicle to be quickly detected and an immediate response to be taken.

[0433] A "security camera" is a recording device that captures video data within a specific area and is used for safety and crime prevention purposes.

[0434] "Video data" refers to images and video information obtained from security cameras, in-car cameras, etc.

[0435] "Analysis" refers to the methods and techniques used to process video data frame by frame and understand and distinguish its content.

[0436] "Suspicious activity" refers to movements or behavior that deviate from normal patterns of behavior and may indicate dangerous or criminal activity.

[0437] A "voice warning" is a voice message that is emitted from a speaker to alert and warn people when suspicious behavior is detected.

[0438] "Security personnel" refers to professional people with safety and security responsibilities.

[0439] "Residents" refers to people who live in a particular building or area.

[0440] An "in-vehicle camera" is a camera device installed inside an autonomous vehicle to monitor passengers and the situation inside the vehicle.

[0441] "Passenger" refers to people using the self-driving vehicle.

[0442] "Operation control center" means a central control facility or organization for monitoring and managing the operation status of autonomous vehicles.

[0443] "Video capture" refers to obtaining a still image of a video at a specific moment.

[0444] "Location information" is data that indicates the geographic location of a particular object or person.

[0445] The "risk level" is a scale for evaluating the degree of danger of detected suspicious behavior or individuals.

[0446] An "alert" is a notification system for issuing warnings and alerts.

[0447] To realize this invention, the following system configuration and program are required. First, video data from security cameras and in-car cameras is received in real time by a server. The server analyzes this video data and uses machine learning algorithms to detect suspicious behavior or individuals. If suspicious behavior or individuals are detected, the server immediately and automatically takes appropriate action.

[0448] The system uses the following hardware and software:

[0449] In-car cameras and security cameras (any camera device)

[0450] Server (high performance computer)

[0451] OpenCV (image processing library)

[0452] YOLO (You Only Look Once) (machine learning model)

[0453] Warning audio playback system (speaker)

[0454] Alert sending system (network connection)

[0455] The server analyzes the video data sent from the camera in real time, analyzing each frame using the YOLO model, which is a technology that detects people and other objects in images quickly and accurately. Using this model, specific behavioral patterns and abnormal behavior can be quickly identified.

[0456] If suspicious behavior is detected, the server plays an audio warning through the speaker and simultaneously sends an alert to the operation control center. The alert includes a video capture of the suspicious person, their location, and a danger level, allowing the operation control center staff to immediately check the situation and take appropriate measures.

[0457] As a concrete example, consider its use in an autonomous vehicle. For example, if an in-car camera detects suspicious activity over a long period of time near the driver's seat, the server analyzes the information and issues an audio warning such as "Security is being strengthened," while simultaneously sending an alert to the operation control center containing details of the suspicious activity. This alert includes a video capture of the scene, the vehicle's location, and a danger level, allowing the administrator to take prompt action.

[0458] In addition, specific examples of prompts for system design using generative AI models are as follows:

[0459] "We need to analyze the in-car footage captured by the camera and detect suspicious individuals. Please help us develop a system that analyzes passenger behavior using the YOLO model and maintains safety."

[0460] As described above, the present invention provides an effective means for enhancing passenger safety in autonomous vehicles, enabling early detection of suspicious behavior and rapid response, thereby ensuring safety inside the vehicle.

[0461] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0462] Step 1:

[0463] The server receives video data from security cameras and in-car cameras in real time.

[0464] Input: Video stream from a camera device.

[0465] Output: Real-time video data.

[0466] How it works: The camera device captures video and sends the data over the network to a server, which receives it and prepares it for analysis.

[0467] Step 2:

[0468] The server analyzes the received video data frame by frame to detect suspicious behavior or people.

[0469] Input: Real-time video data.

[0470] Output: Detection results of suspicious behavior or people.

[0471] How it works: The server uses the YOLO model to analyze video frames, identify people and objects in each frame, and record any specific behavioral patterns or abnormal behaviors.

[0472] Step 3:

[0473] If suspicious behavior is detected, the server obtains video capture and location information to form suspicious person information.

[0474] Input: Suspicious behavior or person detection results.

[0475] Output: Video capture, location information, suspicious person information.

[0476] Specific operation: The server captures frames of any detected suspicious behavior or person and obtains location information such as GPS data. This information is then integrated to generate information about the suspicious person.

[0477] Step 4:

[0478] When a suspicious person is detected, the server plays a warning sound from the car's speakers.

[0479] Input: Suspicious person information.

[0480] Output: Warning audio.

[0481] Specific operation: The server sends instructions to the speaker to play a pre-prepared warning sound such as "Security is being strengthened," thereby alerting passengers and suspicious people inside the vehicle.

[0482] Step 5:

[0483] The server sends an alert containing information about the suspicious person to the operation control center.

[0484] Input: Suspicious person information.

[0485] Output: The alert message.

[0486] Specific operation: The server sends an alert message to the operation control center, including a video capture of the suspicious person, their location, and the danger level. Operation control personnel can use this information to take prompt action.

[0487] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0488] The present invention provides a system that combines security cameras, generative AI, and an emotion engine to recognize the user's emotions, detect suspicious individuals, and prevent crimes before they occur.

[0489] System Configuration

[0490] 1. Security camera and server integration

[0491] The server receives video data from security cameras in real time and uses computer vision techniques, machine learning algorithms, and an emotion engine to process the video data.

[0492] 2. Analysis of video data

[0493] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns or suspicious movements. For example, it can detect people who stay in a particular location for a long time or people who leave their luggage behind. This analysis uses technologies such as YOLO (You Only Look Once) and OpenCV.

[0494] 3. Emotional Recognition

[0495] The server analyzes the user's facial expressions from the video data and estimates their emotions. This emotion analysis uses an emotion engine (e.g., Facial Emotion Recognition). This allows it to identify whether the user is feeling anxiety, fear, anger, or other emotions.

[0496] 4. Identifying suspicious individuals and storing their information

[0497] When a suspicious individual is detected, the server stores the information in a database, including the individual's location, time, behavioral patterns, and perceived emotions, providing information for later appropriate response.

[0498] 5. Audio warning transmission and adjustment

[0499] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which plays an audio warning such as "Security is being strengthened." This warning not only alerts the suspicious person, but also other people in the area. Furthermore, the content and tone of the audio warning are adjusted appropriately based on the user's emotions recognized by the emotion engine.

[0500] 6. Sending Alerts

[0501] The server sends alerts to security staff and residents when a suspicious individual is detected, including video captures of the individual, their location, danger level, and perceived emotion, allowing relevant parties to take immediate action.

[0502] Specific examples

[0503] For example, suppose that footage from a security camera installed at a train station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server then detects the emotion of fear from Person A's facial expression. Based on this, the server determines that not only is Person A's behavior abnormal, but that fear is also behind that behavior.

[0504] This information is stored in a database, and the server sends a command to a speaker next to the camera, playing a warning message saying "security is being strengthened" and adjusting the tone of the message to a gentler tone. At the same time, an alert is sent to security staff, informing them that the video capture includes location information, danger level, and fear emotion. This allows security staff to respond to the scene quickly and with the underlying fear factor in mind.

[0505] In this way, by combining emotion engines, it is possible to create a system that can prevent crimes while also responding appropriately to the situation.

[0506] Natural language description of the program

[0507] The system's program analyzes video data sent from security cameras in real time to not only detect suspicious individuals but also analyze user emotions. Specifically, the server receives the video data and analyzes it using machine learning algorithms and an emotion engine. If suspicious behavior or a specific emotion is detected, the server plays an audio warning through a speaker and simultaneously sends an alert to security staff and residents. The alert includes the suspicious individual's video capture, location information, danger level, and recognized emotion.

[0508] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[0509] The processing flow will be explained below.

[0510] Step 1:

[0511] The server receives video data from the security camera in real time, specifically by continuously capturing video frames using a streaming protocol (e.g., RTSP).

[0512] Step 2:

[0513] The server analyzes the received video data frame by frame. Specifically, the video frames are input into a machine learning algorithm (such as YOLO or OpenCV) to identify the position and movement of people.

[0514] Step 3:

[0515] The server detects specific patterns of behavior or suspicious activity by using anomaly detection algorithms to identify behavior that differs from normal patterns (such as prolonged stays, rapid movements, or unnatural changes in location).

[0516] Step 4:

[0517] The server analyzes the user's facial expressions from the video data and estimates their emotions. Specifically, it uses an emotion engine to identify emotions such as anxiety, fear, and anger from the user's facial expressions.

[0518] Step 5:

[0519] The server determines whether a person is suspicious. Specifically, it evaluates both behavioral patterns and emotions to identify suspicious individuals who meet pre-defined criteria (e.g., long stay + emotion of fear).

[0520] Step 6:

[0521] The server stores the suspect's information in a database, including location, time, behavioral details, and perceived emotions, for use as data for future responses.

[0522] Step 7:

[0523] The server then sends an audio warning to a speaker installed in the location where a suspicious person was detected. Specifically, it connects to a speaker controller and sends instructions to play a warning voice saying, "Security being strengthened, security being strengthened." The tone may be adjusted based on emotion.

[0524] Step 8:

[0525] The server sends alerts to security staff and residents, including video capture of the suspicious individual, their location, danger level, and perceived emotion, allowing relevant parties to initiate a rapid response.

[0526] Step 9:

[0527] Security staff and residents receive an alert and rush to the scene to identify and respond to suspicious individuals. Specifically, they conduct on-site investigations and take necessary measures based on the information contained in the alert.

[0528] Example 2

[0529] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0530] While conventional security systems can detect suspicious behavior by analyzing video data, they face the challenge of being unable to analyze user emotions and respond appropriately. Furthermore, there are insufficient means for notifying security staff and residents of suspicious person information in real time, making it difficult to respond quickly. Furthermore, there is a lack of efficient and accurate means for storing and managing suspicious person information.

[0531] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving video data from a security camera, means for analyzing the video data and detecting suspicious behavior, means for analyzing a user's emotion from the video data, means for outputting an audio warning when a suspicious person is detected, means for transmitting suspicious person information to security staff or residents, means for saving the suspicious person information in a database, means for dividing the video data into frames and performing preprocessing, means for detecting suspicious behavior using a machine learning algorithm, means for estimating a user's emotion using an emotion engine, and means for adjusting the tone of the audio warning as appropriate. This enables the security system to comprehensively analyze the behavior and emotion of a suspicious person and take prompt and appropriate action.

[0532] A "security camera" is a device that monitors a specific area and records and transmits the video data.

[0533] "Video data" is digital data consisting of a series of image frames captured by a security camera.

[0534] A "server" is a computer system for storing, analyzing, and processing data.

[0535] "Analysis" is the process of extracting and evaluating specific information from acquired video data.

[0536] "Suspicious activity" refers to specific unusual or suspicious movements or behaviors in security camera footage.

[0537] "User" refers to a general user, including a person who is the subject of surveillance by a security camera and an administrator who uses the system.

[0538] "Voice warning" is a voice message played through a speaker to alert those around you.

[0539] "Security Staff" means the person or group responsible for monitoring and managing the security system.

[0540] "Resident" means a person who lives in or uses the area in which the security system is installed.

[0541] A "database" is a system for efficiently storing, managing, and retrieving information.

[0542] "Frame segmentation" is the process of breaking down video data into individual still images.

[0543] "Preprocessing" is the process of improving the image quality and removing noise from video data before analysis.

[0544] A "machine learning algorithm" is a learning program that analyzes data and recognizes patterns.

[0545] The "emotion engine" is software that analyzes a user's facial expressions from video data and estimates their emotions.

[0546] "Tone adjustment" is the process of changing the tone and content of an audio alert depending on the situation.

[0547] MODE FOR CARRYING OUT THE INVENTION

[0548] The present invention provides a system that combines security cameras, generative AI models, and an emotion engine to recognize user emotions, detect suspicious individuals, and prevent crimes before they occur.

[0549] System Configuration

[0550] This invention combines various technologies centered around the server, terminal, and user.

[0551] 1. Security camera and server integration

[0552] The server receives video data from security cameras in real time and uses computer vision techniques, machine learning algorithms, and an emotion engine to process the video data.

[0553] 2. Analysis of video data

[0554] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns and suspicious movements, using technologies such as YOLO (You Only Look Once) and OpenCV.

[0555] 3. Emotional Recognition

[0556] The server analyzes the user's facial expressions from the video data and estimates their emotions. This emotion analysis uses an emotion engine (e.g., Facial Emotion Recognition).

[0557] 4. Identifying suspicious individuals and storing their information

[0558] If a suspicious individual is detected, the server stores the information in a database, including the individual's location, time, behavioral patterns, and perceived emotions.

[0559] 5. Audio warning transmission and adjustment

[0560] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which issues an audio warning such as "Security is being strengthened." The content and tone of the audio warning are adjusted appropriately based on the user's emotions recognized by the emotion engine.

[0561] 6. Sending Alerts

[0562] The server sends alerts to security staff and residents when a suspicious individual is detected, including video captures of the individual, their location, danger level, and perceived emotion.

[0563] Specific examples

[0564] For example, suppose that footage from a security camera installed in a station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server then detects the emotion of fear from Person A's facial expression. Based on this, the server determines that not only is Person A's behavior abnormal, but that fear is also behind that behavior.

[0565] This information is stored in a database, and the server sends a command to the speaker next to the camera, which plays a warning message saying "security is being strengthened" and adjusts the tone of the message to a gentler tone. At the same time, an alert is sent to security staff, informing them that the video capture includes location information, danger level, and fear emotion. This allows security staff to respond to the scene quickly and take into account the underlying fear factor.

[0566] Prompt Sentence Examples

[0567] Based on the system log data, an analysis of security camera footage detected Person A behaving suspiciously near the bench. Furthermore, the emotion of fear was detected. Using this information, what information should be provided to the security staff, and what action should be taken next?

[0568] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[0569] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0570] Step 1:

[0571] The server receives video data from security cameras in real time. Specifically, the server receives video streams through a specific IP address and port. The input is the video data from the security cameras, and the output is the video frames stored in the server.

[0572] Step 2:

[0573] The server divides the received video data into frames, extracts the frames using the OpenCV library, and simultaneously performs preprocessing such as noise reduction and image quality improvement. The input is the raw video stream data, and the output is the processed frame images.

[0574] Step 3:

[0575] The server loads a machine learning algorithm (e.g., YOLO) and performs object recognition within each frame. Specifically, it uses the YOLO model to identify people and objects and analyze specific movement and behavior patterns. The input is the preprocessed frame image, and the output is the identified objects and their locations.

[0576] Step 4:

[0577] The server uses an emotion engine (e.g., Facial Emotion Recognition) to analyze the facial expressions of the detected person and estimate their emotions. This allows us to identify the emotions the user is feeling (e.g., anxiety, fear, anger). The input is a person's facial image, and the output is estimated emotional information.

[0578] Step 5:

[0579] The server stores the information of the detected suspicious person (location, time, behavioral patterns, and recognized emotions) in a database. Specifically, it inserts data into a MySQL database. The input is the suspicious person information, and the output is the success / failure status of saving the data to the database.

[0580] Step 6:

[0581] The server sends instructions to the speaker based on the suspicious person information, and issues an audio warning saying "Security is being strengthened." The tone and content of the audio are adjusted based on the results of emotion analysis. The input is suspicious person information and emotion information, and the output is the execution of the audio warning.

[0582] Step 7:

[0583] The server sends alerts to security staff and residents, including video captures of suspicious individuals, their location, danger level, and perceived emotions. The alerts are sent via email or SMS. The input is information about the suspicious individual and their emotions, and the output is the sending of an alert message.

[0584] (Application example 2)

[0585] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0586] In modern society, detecting crimes and suspicious individuals is a major issue. Although security cameras are commonly installed, their ability to detect suspicious individuals and analyze abnormal behavior in real time is insufficient, and there are few systems that can recognize emotions. As a result, it is difficult for users and security staff to take prompt countermeasures, and effective crime prevention measures are needed.

[0587] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0588] In this invention, the server includes means for receiving video data from security cameras, means for analyzing the video data and detecting suspicious behavior, means for analyzing facial expressions from the video data and recognizing emotions, means for sending a push notification to a smartphone when a suspicious person is detected, means for outputting an audio warning when a suspicious person is detected, and means for sending suspicious person information to security staff and residents. This enables real-time detection of suspicious people and recognition of emotions, enabling rapid and effective response.

[0589] A "security camera" is a device that records and transmits video in real time to monitor for crimes and suspicious activity.

[0590] "Video data" is visual information of a scene or object at a particular time obtained from a security camera.

[0591] "Analysis" is the process of using computer vision techniques and machine learning algorithms to identify specific behavioral patterns and anomalous behavior from video data.

[0592] "Suspicious behavior" is behavior that is abnormal to general patterns of behavior and indicates potential criminal activity or security risks.

[0593] "Facial expression" refers to the expression of emotions that can be read from the movement and arrangement of muscles in a person's face.

[0594] "Emotion recognition" is a technology that analyzes people's facial expressions from video data and estimates emotions such as joy, anger, and fear.

[0595] "Push notification" is a technology that allows specific information to be instantly displayed on the screen of a user's device, such as a smartphone.

[0596] A "voice warning" is a method of issuing an audio warning through a speaker to alert suspicious individuals and those around them.

[0597] "Suspicious person information" is data about a person who has engaged in suspicious behavior, and includes video capture, location information, emotion recognition results, danger level, and the like.

[0598] "Security staff" are people who monitor and patrol buildings and areas to ensure their safety.

[0599] This invention is a system that combines a security camera, generative AI, and an emotion recognition engine, and aims to recognize the user's emotions, detect suspicious individuals, and prevent crimes before they occur.

[0600] System Configuration

[0601] This system consists of the following elements:

[0602] 1. Security camera and server integration: The server receives video data from security cameras in real time. This video data is processed on the server. The hardware used can be a standard IP camera or a high-resolution security camera.

[0603] 2. Video data analysis: The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns and suspicious movements. This analysis uses computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., YOLO).

[0604] 3. Emotion Recognition: The server analyzes the user's facial expressions from the video data and estimates their emotions. For emotion analysis, an emotion recognition engine (e.g., Facial Emotion Recognition) is used. This allows the server to identify whether the user is feeling emotions such as anxiety, fear, or anger.

[0605] 4. Suspicious person detection and alert sending: If the server detects a suspicious person, it stores the information in a database and sends a push notification to the smartphone. It can also output an audio warning through the speaker. This notification includes a video capture of the suspicious person, their location, emotion recognition results, and danger level.

[0606] What the program does

[0607] The server processes video data received from security cameras in real time. First, it uses OpenCV to analyze the video data frame by frame to detect suspicious behavior. Next, it uses YOLO to identify specific behavioral patterns. At the same time, it uses an emotion recognition engine to analyze facial expressions from the video data and classify the user's emotions.

[0608] If a suspicious individual is detected, the information is stored in a database and a push notification is sent to the smartphone application. An audio warning is also issued from the speaker based on instructions from the server. The suspicious individual information includes video capture, location information, emotion recognition results, and danger level, allowing security staff and users to take prompt and appropriate action.

[0609] Specific examples

[0610] For example, if a home security camera detects a person staying in a specific location for a long period of time and recognizes the emotion of fear from the person's facial expression, the server will issue a voice warning in a gentle tone from the speaker saying "Security is being strengthened." At the same time, a push notification will be sent to the smartphone, containing information such as "Location: (104, 230, 150, 276), Emotion: Fear."

[0611] Prompt Sentence Examples

[0612] 1. "Security cameras analyze footage in real time to detect people who stay in the area for long periods of time."

[0613] 2. "Build a system that estimates emotions from facial expressions and issues audio warnings in the event of anxiety or fear."

[0614] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0615] Step 1:

[0616] Receive video data from security cameras

[0617] The server receives video data from security cameras in real time. This video data is the input for the entire system. Image data is sent to the server frame by frame, allowing it to proceed to the next analysis step.

[0618] Step 2:

[0619] Video data analysis

[0620] The server analyzes the received video data frame by frame. This analysis uses computer vision technology (OpenCV) and machine learning algorithms (YOLO). It detects people and objects in the frames and identifies suspicious movements and behavior patterns. The input is the video data for each frame, and the output is information on the location and behavior patterns of detected suspicious individuals.

[0621] Step 3:

[0622] Facial Expression Analysis and Emotion Recognition

[0623] The server analyzes the facial expressions of the detected person and uses an emotion recognition engine to recognize their emotions. Specifically, it cuts out the face region and inputs it into the Facial Emotion Recognition model. The data input used here is the image data of the detected face, and the output is the determined emotion (e.g., fear, anger, joy, etc.).

[0624] Step 4:

[0625] Saving suspicious person information

[0626] The server stores information about the detected suspicious individuals (location, behavioral patterns, emotions, etc.) in a database. This information is later used for analysis and countermeasures. The data input is the analysis result, and the record in the database is the output.

[0627] Step 5:

[0628] Sending push notifications

[0629] If a suspicious person is detected, the server sends a push notification to the smartphone. This notification includes the suspicious person's video capture, location information, emotion recognition results, danger level, etc. The data input is the stored suspicious person information, and the output is the notification sent to the smartphone.

[0630] Step 6:

[0631] Sending a voice warning

[0632] The server sends instructions to the speaker to play an audio warning such as "Security is being strengthened." The content and tone of the audio warning are adjusted based on the recognized emotion. The data input is the analysis results and emotion recognition results, and the output is the audio warning emitted from the speaker.

[0633] Step 7:

[0634] Sending alerts

[0635] The server sends alerts containing suspicious person information to security staff and residents. The alerts include video capture, location information, emotion recognition results, and danger level. The data input is the stored suspicious person information, and the output is the sending of an alert message.

[0636] In this way, the system can detect suspicious individuals in real time and take necessary measures immediately.

[0637] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0638] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0639] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0640] [Third embodiment]

[0641] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0642] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0643] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0644] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0645] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0646] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0647] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0648] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0649] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0650] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0651] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0652] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0653] The present invention provides a system that combines security cameras and generative AI to detect suspicious individuals and prevent crimes before they occur.

[0654] System Configuration

[0655] 1. Security camera and server integration

[0656] The server receives video data from security cameras in real time and uses computer vision techniques and machine learning algorithms to process this video data.

[0657] 2. Analysis of video data

[0658] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns. For example, it detects people who stay in a particular location for a long time or people who leave their luggage behind. This analysis uses technologies such as YOLO (You Only Look Once) and OpenCV.

[0659] 3. Identifying suspicious individuals and storing their information

[0660] When a suspicious person is detected, the server stores the information in a database, including the person's location, time, and behavioral patterns. This information provides information for later use in making appropriate decisions.

[0661] 4. Audio warning

[0662] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which plays an audio warning such as "Security is being strengthened." This audio warning is effective in alerting not only the suspicious person but also other people in the area.

[0663] 5. Sending alerts

[0664] When a suspicious person is detected, the server sends an alert to security staff and residents. The alert includes a video capture of the suspicious person, their location, and a danger level, allowing relevant parties to take immediate action.

[0665] Specific examples

[0666] For example, suppose footage from a security camera installed at a station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server determines that Person A's behavior is abnormal and immediately saves the information. The server then sends a command to the speaker next to the camera, which plays a warning audio message saying "Security is being strengthened." At the same time, an alert is sent to the security staff's device, providing location information and danger level along with the video capture. This allows security staff to respond to the scene quickly.

[0667] This system operates 24 hours a day and can be used in a variety of locations, including stations, airports, stores, apartment buildings, and detached homes. It is expected to make a significant contribution to crime prevention.

[0668] Natural language description of the program

[0669] The system's program analyzes video data sent from security cameras in real time to detect suspicious individuals. Specifically, the server receives the video data and analyzes it using a machine learning algorithm. If suspicious activity is detected, the server issues an audio warning through a speaker and simultaneously sends an alert to security staff and residents. The alert includes a video capture of the suspicious individual, their location, and a danger level.

[0670] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[0671] The processing flow will be explained below.

[0672] Step 1:

[0673] The server receives video data from the security camera in real time, specifically by continuously capturing video frames using a streaming protocol (e.g., RTSP).

[0674] Step 2:

[0675] The server analyzes the received video data frame by frame. Specifically, the video frames are input into a machine learning algorithm (such as YOLO or OpenCV) to identify the position and movement of people.

[0676] Step 3:

[0677] The server detects specific patterns of behavior or suspicious activity by using anomaly detection algorithms to identify behavior that differs from normal patterns (such as prolonged stays, rapid movements, or unnatural changes in location).

[0678] Step 4:

[0679] The server determines who is suspected of being a suspicious person by evaluating the detected behavioral patterns and identifying those who meet pre-defined criteria (e.g., staying for a long time or leaving luggage behind).

[0680] Step 5:

[0681] The server stores the suspect's information in a database, including details of location, time, and behavior, and uses it as data for future responses.

[0682] Step 6:

[0683] The server sends an audio warning to the speaker installed in the location where the suspicious person was detected. Specifically, it connects to the speaker controller and sends an instruction to play the warning audio, "Security being strengthened, security being strengthened."

[0684] Step 7:

[0685] The server then sends alerts to security staff and residents, including video capture of the suspicious individual, their location, and a risk level, allowing the relevant parties to initiate a rapid response.

[0686] Step 8:

[0687] Security staff and residents receive an alert and rush to the scene to identify and respond to suspicious individuals. Specifically, they conduct on-site investigations and take necessary measures based on the information contained in the alert.

[0688] Example 1

[0689] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0690] Conventional security camera systems have difficulty detecting suspicious individuals in real time, making it difficult to take immediate action to prevent crimes. Furthermore, they lack a means to accurately and quickly communicate suspicious individual information to relevant parties, which can result in delays in appropriate response. The present invention aims to solve these problems and provide a system for efficiently and effectively detecting and responding to suspicious individuals.

[0691] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0692] In this invention, the server includes means for receiving video data from security cameras, means for preprocessing the video data, means for analyzing the preprocessed video data and identifying suspicious behavior patterns, means for storing information on a suspicious individual in a database when the suspicious individual is detected, means for outputting an audio warning when the suspicious individual is detected, and means for transmitting information on the suspicious individual to security personnel and users. This enables early detection of suspicious individuals and prompt response, making it possible to prevent crimes from occurring.

[0693] A "security camera" is a device installed to capture video within a surveillance area and deter crimes and suspicious behavior.

[0694] "Video Data" means a digital representation of visual information captured by a security camera.

[0695] "Preprocessing" refers to initial processing such as adjusting resolution and removing noise, which is performed to make video data easier to analyze.

[0696] A "behavioral pattern" is an identified series of actions or characteristics of behavior taken by an object in video data.

[0697] A "database" is a system for efficiently storing, managing, and retrieving information.

[0698] A "voice warning" is a voice message that alerts suspicious people and people in the area.

[0699] An "alert message" is a notification containing information about a suspicious person, and is sent promptly to security personnel and users.

[0700] "Image processing technology" refers to a set of computer techniques for analyzing, transforming, and understanding digital images.

[0701] A "learning algorithm" is an algorithm that learns patterns and knowledge from data and makes predictions and classifications.

[0702] "Suspicious person information" is detailed information such as the suspicious person's behavior, location, time, and danger level.

[0703] "Location data" is information that indicates the geographical location of an object or person.

[0704] The "risk level" is an index that indicates the degree of risk based on the behavior of a suspicious person and the situation.

[0705] MODE FOR CARRYING OUT THE INVENTION

[0706] The purpose of this invention is to detect suspicious individuals and prevent crimes before they occur by using a system that combines security cameras and generative AI. This system consists of security cameras, a server, a speaker, and a terminal.

[0707] The server receives video data from security cameras in real time. It is always connected via the network and receives the video data captured by the security cameras sequentially. A multipurpose digital signal processing device is used to analyze the video data. At this stage, the video data is unprocessed raw data.

[0708] The server then preprocesses the received video data. This preprocessing involves using image processing techniques such as the OpenCV library to adjust the resolution and remove noise. Specifically, the video is converted to grayscale and a noise reduction filter is applied. This makes subsequent analysis easier.

[0709] After preprocessing is complete, the server performs object detection and behavioral pattern analysis using YOLO (You Only Look Once) or similar object detection algorithms. The analysis is performed frame by frame to identify people who linger for long periods of time or who behave abnormally.

[0710] If a suspicious person is detected, the server stores the information in a database, including the person's location, time, and behavioral patterns. For example, it connects to an SQL database and stores the information using an INSERT statement.

[0711] The server also issues an audio warning to detected suspicious individuals. It sends an HTTP request to the speaker and uses a speech synthesis library to output an audio warning such as "Security is being strengthened." This audio warning alerts not only the suspicious individual but also those around them.

[0712] The server then sends an alert message to the security personnel or the user's device, which includes a video capture of the suspicious person, their location, and the danger level. For example, the alert message can be sent via email or SMS API (e.g., Twilio).

[0713] Specific examples

[0714] For example, video data from a security camera installed at a station is sent to a server, and after analysis, person A is detected behaving suspiciously near a bench. The server determines that person A's behavior is abnormal and stores this information in a database. The server then issues an instruction to a speaker to play a warning audio message such as "Security is being strengthened." At the same time, an alert is sent to the security officer's device, providing a video capture, location information, and danger level. This allows security officers to respond to the scene quickly.

[0715] Example prompts for generative AI models

[0716] 1. "How can I analyze security camera footage at a station to detect suspicious individuals?"

[0717] 2. "How can I build an automation system that detects suspicious activity from security cameras and sends audio warnings and alerts?"

[0718] As described above, this invention utilizes security cameras and generative AI to enable early detection of suspicious individuals and rapid response. Using this system can greatly contribute to crime prevention in a variety of locations.

[0719] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0720] Program processing steps

[0721] Step 1:

[0722] The server receives video data from security cameras in real time. It takes the video stream (e.g., RTSP) from the security cameras as input and stores the video data in the server's memory. The output at this stage is the raw video data stored in memory.

[0723] Step 2:

[0724] The server preprocesses the received video data. The input is the video data saved in step 1, and the server uses the OpenCV library to adjust the resolution and remove noise. Specifically, it converts the video to grayscale and applies a noise reduction filter. The output is the preprocessed video data.

[0725] Step 3:

[0726] The server analyzes the preprocessed video data to identify objects and behavioral patterns. The input is the preprocessed video data created in step 2, which is analyzed using an object detection algorithm such as YOLO. Specifically, the video data is analyzed frame by frame to identify the location and movement of people and luggage. The output is data on detected objects and behavioral patterns.

[0727] Step 4:

[0728] When a suspicious person is detected, the server stores the information in a database. The input is the suspicious person information detected in step 3, and the server connects to an SQL database and uses the INSERT statement to store information such as the suspicious person's location, time, and behavioral patterns. The output is a new record stored in the database.

[0729] Step 5:

[0730] The server issues an audio warning if a suspicious person is detected. The input is the suspicious person information detected in step 3, and an HTTP request is sent to the speaker, which uses a Python speech synthesis library to play the audio message "Security is being strengthened." The output is an audio warning emitted from the speaker.

[0731] Step 6:

[0732] The server sends suspicious person information to the device of the security officer or user. The input is the suspicious person information detected in step 3, and an alert message is sent using email or an SMS API (for example, Twilio). The alert includes a video capture image, location information, and danger level. The output is an alert message that is displayed on the device of the security officer or user.

[0733] The above are the specific processing steps of this system. We have explained in detail how the input data is processed and what kind of output is obtained at each step. We have also clearly stated what specific operations are performed at each step. This makes the overall flow and function of this system clear.

[0734] (Application example 1)

[0735] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0736] Ensuring safety in autonomous vehicles is an important issue. In particular, there is a need to detect suspicious or dangerous behavior in the vehicle early and take appropriate measures promptly. However, conventional methods make it difficult to detect suspicious behavior in the vehicle in real time and respond immediately. For this reason, a new system is needed to improve safety in autonomous vehicles.

[0737] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0738] In this invention, the server includes means for receiving video data from security cameras, means for analyzing the video data and detecting suspicious behavior, means for outputting an audio warning when a suspicious individual is detected, means for transmitting suspicious individual information to security personnel and residents, means for receiving video data from an in-vehicle camera of the autonomous vehicle and detecting suspicious behavior of passengers, means for issuing an audio warning from an in-vehicle speaker, and means for transmitting an alert to an operation control center including video capture of the suspicious individual, location information, and danger level. This enables suspicious behavior in the autonomous vehicle to be quickly detected and an immediate response to be taken.

[0739] A "security camera" is a recording device that captures video data within a specific area and is used for safety and crime prevention purposes.

[0740] "Video data" refers to images and video information obtained from security cameras, in-car cameras, etc.

[0741] "Analysis" refers to the methods and techniques used to process video data frame by frame and understand and distinguish its content.

[0742] "Suspicious activity" refers to movements or behavior that deviate from normal patterns of behavior and may indicate dangerous or criminal activity.

[0743] A "voice warning" is a voice message that is emitted from a speaker to alert and warn people when suspicious behavior is detected.

[0744] "Security personnel" refers to professional people with safety and security responsibilities.

[0745] "Residents" refers to people who live in a particular building or area.

[0746] An "in-vehicle camera" is a camera device installed inside an autonomous vehicle to monitor passengers and the situation inside the vehicle.

[0747] "Passenger" refers to people using the self-driving vehicle.

[0748] "Operation control center" means a central control facility or organization for monitoring and managing the operation status of autonomous vehicles.

[0749] "Video capture" refers to obtaining a still image of a video at a specific moment.

[0750] "Location information" is data that indicates the geographic location of a particular object or person.

[0751] The "risk level" is a scale for evaluating the degree of danger of detected suspicious behavior or individuals.

[0752] An "alert" is a notification system for issuing warnings and alerts.

[0753] To realize this invention, the following system configuration and program are required. First, video data from security cameras and in-car cameras is received in real time by a server. The server analyzes this video data and uses machine learning algorithms to detect suspicious behavior or individuals. If suspicious behavior or individuals are detected, the server immediately and automatically takes appropriate action.

[0754] The system uses the following hardware and software:

[0755] In-car cameras and security cameras (any camera device)

[0756] Server (high performance computer)

[0757] OpenCV (image processing library)

[0758] YOLO (You Only Look Once) (machine learning model)

[0759] Warning audio playback system (speaker)

[0760] Alert sending system (network connection)

[0761] The server analyzes the video data sent from the camera in real time, analyzing each frame using the YOLO model, which is a technology that detects people and other objects in images quickly and accurately. Using this model, specific behavioral patterns and abnormal behavior can be quickly identified.

[0762] If suspicious behavior is detected, the server plays an audio warning through the speaker and simultaneously sends an alert to the operation control center. The alert includes a video capture of the suspicious person, their location, and a danger level, allowing the operation control center staff to immediately check the situation and take appropriate measures.

[0763] As a concrete example, consider its use in an autonomous vehicle. For example, if an in-car camera detects suspicious activity over a long period of time near the driver's seat, the server analyzes the information and issues an audio warning such as "Security is being strengthened," while simultaneously sending an alert to the operation control center containing details of the suspicious activity. This alert includes a video capture of the scene, the vehicle's location, and a danger level, allowing the administrator to take prompt action.

[0764] In addition, specific examples of prompts for system design using generative AI models are as follows:

[0765] "We need to analyze the in-car footage captured by the camera and detect suspicious individuals. Please help us develop a system that analyzes passenger behavior using the YOLO model and maintains safety."

[0766] As described above, the present invention provides an effective means for enhancing passenger safety in autonomous vehicles, enabling early detection of suspicious behavior and rapid response, thereby ensuring safety inside the vehicle.

[0767] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0768] Step 1:

[0769] The server receives video data from security cameras and in-car cameras in real time.

[0770] Input: Video stream from a camera device.

[0771] Output: Real-time video data.

[0772] How it works: The camera device captures video and sends the data over the network to a server, which receives it and prepares it for analysis.

[0773] Step 2:

[0774] The server analyzes the received video data frame by frame to detect suspicious behavior or people.

[0775] Input: Real-time video data.

[0776] Output: Detection results of suspicious behavior or people.

[0777] How it works: The server uses the YOLO model to analyze video frames, identify people and objects in each frame, and record any specific behavioral patterns or abnormal behaviors.

[0778] Step 3:

[0779] If suspicious behavior is detected, the server obtains video capture and location information to form suspicious person information.

[0780] Input: Suspicious behavior or person detection results.

[0781] Output: Video capture, location information, suspicious person information.

[0782] Specific operation: The server captures frames of any detected suspicious behavior or person and obtains location information such as GPS data. This information is then integrated to generate information about the suspicious person.

[0783] Step 4:

[0784] When a suspicious person is detected, the server plays a warning sound from the car's speakers.

[0785] Input: Suspicious person information.

[0786] Output: Warning audio.

[0787] Specific operation: The server sends instructions to the speaker to play a pre-prepared warning sound such as "Security is being strengthened," thereby alerting passengers and suspicious people inside the vehicle.

[0788] Step 5:

[0789] The server sends an alert containing information about the suspicious person to the operation control center.

[0790] Input: Suspicious person information.

[0791] Output: The alert message.

[0792] Specific operation: The server sends an alert message to the operation control center, including a video capture of the suspicious person, their location, and the danger level. Operation control personnel can use this information to take prompt action.

[0793] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0794] The present invention provides a system that combines security cameras, generative AI, and an emotion engine to recognize the user's emotions, detect suspicious individuals, and prevent crimes before they occur.

[0795] System Configuration

[0796] 1. Security camera and server integration

[0797] The server receives video data from security cameras in real time and uses computer vision techniques, machine learning algorithms, and an emotion engine to process the video data.

[0798] 2. Analysis of video data

[0799] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns or suspicious movements. For example, it can detect people who stay in a particular location for a long time or people who leave their luggage behind. This analysis uses technologies such as YOLO (You Only Look Once) and OpenCV.

[0800] 3. Emotional Recognition

[0801] The server analyzes the user's facial expressions from the video data and estimates their emotions. This emotion analysis uses an emotion engine (e.g., Facial Emotion Recognition). This allows it to identify whether the user is feeling anxiety, fear, anger, or other emotions.

[0802] 4. Identifying suspicious individuals and storing their information

[0803] When a suspicious individual is detected, the server stores the information in a database, including the individual's location, time, behavioral patterns, and perceived emotions, providing information for later appropriate response.

[0804] 5. Audio warning transmission and adjustment

[0805] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which plays an audio warning such as "Security is being strengthened." This warning not only alerts the suspicious person, but also other people in the area. Furthermore, the content and tone of the audio warning are adjusted appropriately based on the user's emotions recognized by the emotion engine.

[0806] 6. Sending Alerts

[0807] The server sends alerts to security staff and residents when a suspicious individual is detected, including video captures of the individual, their location, danger level, and perceived emotion, allowing relevant parties to take immediate action.

[0808] Specific examples

[0809] For example, suppose that footage from a security camera installed at a train station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server then detects the emotion of fear from Person A's facial expression. Based on this, the server determines that not only is Person A's behavior abnormal, but that fear is also behind that behavior.

[0810] This information is stored in a database, and the server sends a command to a speaker next to the camera, playing a warning message saying "security is being strengthened" and adjusting the tone of the message to a gentler tone. At the same time, an alert is sent to security staff, informing them that the video capture includes location information, danger level, and fear emotion. This allows security staff to respond to the scene quickly and with the underlying fear factor in mind.

[0811] In this way, by combining emotion engines, it is possible to create a system that can prevent crimes while also responding appropriately to the situation.

[0812] Natural language description of the program

[0813] The system's program analyzes video data sent from security cameras in real time to not only detect suspicious individuals but also analyze user emotions. Specifically, the server receives the video data and analyzes it using machine learning algorithms and an emotion engine. If suspicious behavior or a specific emotion is detected, the server plays an audio warning through a speaker and simultaneously sends an alert to security staff and residents. The alert includes the suspicious individual's video capture, location information, danger level, and recognized emotion.

[0814] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[0815] The processing flow will be explained below.

[0816] Step 1:

[0817] The server receives video data from the security camera in real time, specifically by continuously capturing video frames using a streaming protocol (e.g., RTSP).

[0818] Step 2:

[0819] The server analyzes the received video data frame by frame. Specifically, the video frames are input into a machine learning algorithm (such as YOLO or OpenCV) to identify the position and movement of people.

[0820] Step 3:

[0821] The server detects specific patterns of behavior or suspicious activity by using anomaly detection algorithms to identify behavior that differs from normal patterns (such as prolonged stays, rapid movements, or unnatural changes in location).

[0822] Step 4:

[0823] The server analyzes the user's facial expressions from the video data and estimates their emotions. Specifically, it uses an emotion engine to identify emotions such as anxiety, fear, and anger from the user's facial expressions.

[0824] Step 5:

[0825] The server determines whether a person is suspicious. Specifically, it evaluates both behavioral patterns and emotions to identify suspicious individuals who meet pre-defined criteria (e.g., long stay + emotion of fear).

[0826] Step 6:

[0827] The server stores the suspect's information in a database, including location, time, behavioral details, and perceived emotions, for use as data for future responses.

[0828] Step 7:

[0829] The server then sends an audio warning to a speaker installed in the location where a suspicious person was detected. Specifically, it connects to a speaker controller and sends instructions to play a warning voice saying, "Security being strengthened, security being strengthened." The tone may be adjusted based on emotion.

[0830] Step 8:

[0831] The server sends alerts to security staff and residents, including video capture of the suspicious individual, their location, danger level, and perceived emotion, allowing relevant parties to initiate a rapid response.

[0832] Step 9:

[0833] Security staff and residents receive an alert and rush to the scene to identify and respond to suspicious individuals. Specifically, they conduct on-site investigations and take necessary measures based on the information contained in the alert.

[0834] Example 2

[0835] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0836] While conventional security systems can detect suspicious behavior by analyzing video data, they face the challenge of being unable to analyze user emotions and respond appropriately. Furthermore, there are insufficient means for notifying security staff and residents of suspicious person information in real time, making it difficult to respond quickly. Furthermore, there is a lack of efficient and accurate means for storing and managing suspicious person information.

[0837] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving video data from a security camera, means for analyzing the video data and detecting suspicious behavior, means for analyzing a user's emotion from the video data, means for outputting an audio warning when a suspicious person is detected, means for transmitting suspicious person information to security staff or residents, means for saving the suspicious person information in a database, means for dividing the video data into frames and performing preprocessing, means for detecting suspicious behavior using a machine learning algorithm, means for estimating a user's emotion using an emotion engine, and means for adjusting the tone of the audio warning as appropriate. This enables the security system to comprehensively analyze the behavior and emotion of a suspicious person and take prompt and appropriate action.

[0838] A "security camera" is a device that monitors a specific area and records and transmits the video data.

[0839] "Video data" is digital data consisting of a series of image frames captured by a security camera.

[0840] A "server" is a computer system for storing, analyzing, and processing data.

[0841] "Analysis" is the process of extracting and evaluating specific information from acquired video data.

[0842] "Suspicious activity" refers to specific unusual or suspicious movements or behaviors in security camera footage.

[0843] "User" refers to a general user, including a person who is the subject of surveillance by a security camera and an administrator who uses the system.

[0844] "Voice warning" is a voice message played through a speaker to alert those around you.

[0845] "Security Staff" means the person or group responsible for monitoring and managing the security system.

[0846] "Resident" means a person who lives in or uses the area in which the security system is installed.

[0847] A "database" is a system for efficiently storing, managing, and retrieving information.

[0848] "Frame segmentation" is the process of breaking down video data into individual still images.

[0849] "Preprocessing" is the process of improving the image quality and removing noise from video data before analysis.

[0850] A "machine learning algorithm" is a learning program that analyzes data and recognizes patterns.

[0851] The "emotion engine" is software that analyzes a user's facial expressions from video data and estimates their emotions.

[0852] "Tone adjustment" is the process of changing the tone and content of an audio alert depending on the situation.

[0853] MODE FOR CARRYING OUT THE INVENTION

[0854] The present invention provides a system that combines security cameras, generative AI models, and an emotion engine to recognize user emotions, detect suspicious individuals, and prevent crimes before they occur.

[0855] System Configuration

[0856] This invention combines various technologies centered around the server, terminal, and user.

[0857] 1. Security camera and server integration

[0858] The server receives video data from security cameras in real time and uses computer vision techniques, machine learning algorithms, and an emotion engine to process the video data.

[0859] 2. Analysis of video data

[0860] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns and suspicious movements, using technologies such as YOLO (You Only Look Once) and OpenCV.

[0861] 3. Emotional Recognition

[0862] The server analyzes the user's facial expressions from the video data and estimates their emotions. This emotion analysis uses an emotion engine (e.g., Facial Emotion Recognition).

[0863] 4. Identifying suspicious individuals and storing their information

[0864] If a suspicious individual is detected, the server stores the information in a database, including the individual's location, time, behavioral patterns, and perceived emotions.

[0865] 5. Audio warning transmission and adjustment

[0866] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which issues an audio warning such as "Security is being strengthened." The content and tone of the audio warning are adjusted appropriately based on the user's emotions recognized by the emotion engine.

[0867] 6. Sending Alerts

[0868] The server sends alerts to security staff and residents when a suspicious individual is detected, including video captures of the individual, their location, danger level, and perceived emotion.

[0869] Specific examples

[0870] For example, suppose that footage from a security camera installed in a station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server then detects the emotion of fear from Person A's facial expression. Based on this, the server determines that not only is Person A's behavior abnormal, but that fear is also behind that behavior.

[0871] This information is stored in a database, and the server sends a command to the speaker next to the camera, which plays a warning message saying "security is being strengthened" and adjusts the tone of the message to a gentler tone. At the same time, an alert is sent to security staff, informing them that the video capture includes location information, danger level, and fear emotion. This allows security staff to respond to the scene quickly and take into account the underlying fear factor.

[0872] Prompt Sentence Examples

[0873] Based on the system log data, an analysis of security camera footage detected Person A behaving suspiciously near the bench. Furthermore, the emotion of fear was detected. Using this information, what information should be provided to the security staff, and what action should be taken next?

[0874] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[0875] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0876] Step 1:

[0877] The server receives video data from security cameras in real time. Specifically, the server receives video streams through a specific IP address and port. The input is the video data from the security cameras, and the output is the video frames stored in the server.

[0878] Step 2:

[0879] The server divides the received video data into frames, extracts the frames using the OpenCV library, and simultaneously performs preprocessing such as noise reduction and image quality improvement. The input is the raw video stream data, and the output is the processed frame images.

[0880] Step 3:

[0881] The server loads a machine learning algorithm (e.g., YOLO) and performs object recognition within each frame. Specifically, it uses the YOLO model to identify people and objects and analyze specific movement and behavior patterns. The input is the preprocessed frame image, and the output is the identified objects and their locations.

[0882] Step 4:

[0883] The server uses an emotion engine (e.g., Facial Emotion Recognition) to analyze the facial expressions of the detected person and estimate their emotions. This allows us to identify the emotions the user is feeling (e.g., anxiety, fear, anger). The input is a person's facial image, and the output is estimated emotional information.

[0884] Step 5:

[0885] The server stores the information of the detected suspicious person (location, time, behavioral patterns, and recognized emotions) in a database. Specifically, it inserts data into a MySQL database. The input is the suspicious person information, and the output is the success / failure status of saving the data to the database.

[0886] Step 6:

[0887] The server sends instructions to the speaker based on the suspicious person information, and issues an audio warning saying "Security is being strengthened." The tone and content of the audio are adjusted based on the results of emotion analysis. The input is suspicious person information and emotion information, and the output is the execution of the audio warning.

[0888] Step 7:

[0889] The server sends alerts to security staff and residents, including video captures of suspicious individuals, their location, danger level, and perceived emotions. The alerts are sent via email or SMS. The input is information about the suspicious individual and their emotions, and the output is the sending of an alert message.

[0890] (Application example 2)

[0891] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0892] In modern society, detecting crimes and suspicious individuals is a major issue. Although security cameras are commonly installed, their ability to detect suspicious individuals and analyze abnormal behavior in real time is insufficient, and there are few systems that can recognize emotions. As a result, it is difficult for users and security staff to take prompt countermeasures, and effective crime prevention measures are needed.

[0893] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0894] In this invention, the server includes means for receiving video data from security cameras, means for analyzing the video data and detecting suspicious behavior, means for analyzing facial expressions from the video data and recognizing emotions, means for sending a push notification to a smartphone when a suspicious person is detected, means for outputting an audio warning when a suspicious person is detected, and means for sending suspicious person information to security staff and residents. This enables real-time detection of suspicious people and recognition of emotions, enabling rapid and effective response.

[0895] A "security camera" is a device that records and transmits video in real time to monitor for crimes and suspicious activity.

[0896] "Video data" is visual information of a scene or object at a particular time obtained from a security camera.

[0897] "Analysis" is the process of using computer vision techniques and machine learning algorithms to identify specific behavioral patterns and anomalous behavior from video data.

[0898] "Suspicious behavior" is behavior that is abnormal to general patterns of behavior and indicates potential criminal activity or security risks.

[0899] "Facial expression" refers to the expression of emotions that can be read from the movement and arrangement of muscles in a person's face.

[0900] "Emotion recognition" is a technology that analyzes people's facial expressions from video data and estimates emotions such as joy, anger, and fear.

[0901] "Push notification" is a technology that allows specific information to be instantly displayed on the screen of a user's device, such as a smartphone.

[0902] A "voice warning" is a method of issuing an audio warning through a speaker to alert suspicious individuals and those around them.

[0903] "Suspicious person information" is data about a person who has engaged in suspicious behavior, and includes video capture, location information, emotion recognition results, danger level, and the like.

[0904] "Security staff" are people who monitor and patrol buildings and areas to ensure their safety.

[0905] This invention is a system that combines a security camera, generative AI, and an emotion recognition engine, and aims to recognize the user's emotions, detect suspicious individuals, and prevent crimes before they occur.

[0906] System Configuration

[0907] This system consists of the following elements:

[0908] 1. Security camera and server integration: The server receives video data from security cameras in real time. This video data is processed on the server. The hardware used can be a standard IP camera or a high-resolution security camera.

[0909] 2. Video data analysis: The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns and suspicious movements. This analysis uses computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., YOLO).

[0910] 3. Emotion Recognition: The server analyzes the user's facial expressions from the video data and estimates their emotions. For emotion analysis, an emotion recognition engine (e.g., Facial Emotion Recognition) is used. This allows the server to identify whether the user is feeling emotions such as anxiety, fear, or anger.

[0911] 4. Suspicious person detection and alert sending: If the server detects a suspicious person, it stores the information in a database and sends a push notification to the smartphone. It can also output an audio warning through the speaker. This notification includes a video capture of the suspicious person, their location, emotion recognition results, and danger level.

[0912] What the program does

[0913] The server processes video data received from security cameras in real time. First, it uses OpenCV to analyze the video data frame by frame to detect suspicious behavior. Next, it uses YOLO to identify specific behavioral patterns. At the same time, it uses an emotion recognition engine to analyze facial expressions from the video data and classify the user's emotions.

[0914] If a suspicious individual is detected, the information is stored in a database and a push notification is sent to the smartphone application. An audio warning is also issued from the speaker based on instructions from the server. The suspicious individual information includes video capture, location information, emotion recognition results, and danger level, allowing security staff and users to take prompt and appropriate action.

[0915] Specific examples

[0916] For example, if a home security camera detects a person staying in a specific location for a long period of time and recognizes the emotion of fear from the person's facial expression, the server will issue a voice warning in a gentle tone from the speaker saying "Security is being strengthened." At the same time, a push notification will be sent to the smartphone, containing information such as "Location: (104, 230, 150, 276), Emotion: Fear."

[0917] Prompt Sentence Examples

[0918] 1. "Security cameras analyze footage in real time to detect people who stay in the area for long periods of time."

[0919] 2. "Build a system that estimates emotions from facial expressions and issues audio warnings in the event of anxiety or fear."

[0920] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0921] Step 1:

[0922] Receive video data from security cameras

[0923] The server receives video data from security cameras in real time. This video data is the input for the entire system. Image data is sent to the server frame by frame, allowing it to proceed to the next analysis step.

[0924] Step 2:

[0925] Video data analysis

[0926] The server analyzes the received video data frame by frame. This analysis uses computer vision technology (OpenCV) and machine learning algorithms (YOLO). It detects people and objects in the frames and identifies suspicious movements and behavior patterns. The input is the video data for each frame, and the output is information on the location and behavior patterns of detected suspicious individuals.

[0927] Step 3:

[0928] Facial Expression Analysis and Emotion Recognition

[0929] The server analyzes the facial expressions of the detected person and uses an emotion recognition engine to recognize their emotions. Specifically, it cuts out the face region and inputs it into the Facial Emotion Recognition model. The data input used here is the image data of the detected face, and the output is the determined emotion (e.g., fear, anger, joy, etc.).

[0930] Step 4:

[0931] Saving suspicious person information

[0932] The server stores information about the detected suspicious individuals (location, behavioral patterns, emotions, etc.) in a database. This information is later used for analysis and countermeasures. The data input is the analysis result, and the record in the database is the output.

[0933] Step 5:

[0934] Sending push notifications

[0935] If a suspicious person is detected, the server sends a push notification to the smartphone. This notification includes the suspicious person's video capture, location information, emotion recognition results, danger level, etc. The data input is the stored suspicious person information, and the output is the notification sent to the smartphone.

[0936] Step 6:

[0937] Sending a voice warning

[0938] The server sends instructions to the speaker to play an audio warning such as "Security is being strengthened." The content and tone of the audio warning are adjusted based on the recognized emotion. The data input is the analysis results and emotion recognition results, and the output is the audio warning emitted from the speaker.

[0939] Step 7:

[0940] Sending alerts

[0941] The server sends alerts containing suspicious person information to security staff and residents. The alerts include video capture, location information, emotion recognition results, and danger level. The data input is the stored suspicious person information, and the output is the sending of an alert message.

[0942] In this way, the system can detect suspicious individuals in real time and take necessary measures immediately.

[0943] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0944] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0945] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0946] [Fourth embodiment]

[0947] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0948] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0949] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0950] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0951] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0952] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0953] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0954] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0955] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0956] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0957] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0958] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0959] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0960] The present invention provides a system that combines security cameras and generative AI to detect suspicious individuals and prevent crimes before they occur.

[0961] System Configuration

[0962] 1. Security camera and server integration

[0963] The server receives video data from security cameras in real time and uses computer vision techniques and machine learning algorithms to process this video data.

[0964] 2. Analysis of video data

[0965] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns. For example, it detects people who stay in a particular location for a long time or people who leave their luggage behind. This analysis uses technologies such as YOLO (You Only Look Once) and OpenCV.

[0966] 3. Identifying suspicious individuals and storing their information

[0967] When a suspicious person is detected, the server stores the information in a database, including the person's location, time, and behavioral patterns. This information provides information for later use in making appropriate decisions.

[0968] 4. Audio warning

[0969] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which plays an audio warning such as "Security is being strengthened." This audio warning is effective in alerting not only the suspicious person but also other people in the area.

[0970] 5. Sending alerts

[0971] When a suspicious person is detected, the server sends an alert to security staff and residents. The alert includes a video capture of the suspicious person, their location, and a danger level, allowing relevant parties to take immediate action.

[0972] Specific examples

[0973] For example, suppose footage from a security camera installed at a station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server determines that Person A's behavior is abnormal and immediately saves the information. The server then sends a command to the speaker next to the camera, which plays a warning audio message saying "Security is being strengthened." At the same time, an alert is sent to the security staff's device, providing location information and danger level along with the video capture. This allows security staff to respond to the scene quickly.

[0974] This system operates 24 hours a day and can be used in a variety of locations, including stations, airports, stores, apartment buildings, and detached homes. It is expected to make a significant contribution to crime prevention.

[0975] Natural language description of the program

[0976] The system's program analyzes video data sent from security cameras in real time to detect suspicious individuals. Specifically, the server receives the video data and analyzes it using a machine learning algorithm. If suspicious activity is detected, the server issues an audio warning through a speaker and simultaneously sends an alert to security staff and residents. The alert includes a video capture of the suspicious individual, their location, and a danger level.

[0977] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[0978] The processing flow will be explained below.

[0979] Step 1:

[0980] The server receives video data from the security camera in real time, specifically by continuously capturing video frames using a streaming protocol (e.g., RTSP).

[0981] Step 2:

[0982] The server analyzes the received video data frame by frame. Specifically, the video frames are input into a machine learning algorithm (such as YOLO or OpenCV) to identify the position and movement of people.

[0983] Step 3:

[0984] The server detects specific patterns of behavior or suspicious activity by using anomaly detection algorithms to identify behavior that differs from normal patterns (such as prolonged stays, rapid movements, or unnatural changes in location).

[0985] Step 4:

[0986] The server determines who is suspected of being a suspicious person by evaluating the detected behavioral patterns and identifying those who meet pre-defined criteria (e.g., staying for a long time or leaving luggage behind).

[0987] Step 5:

[0988] The server stores the suspect's information in a database, including details of location, time, and behavior, and uses it as data for future responses.

[0989] Step 6:

[0990] The server sends an audio warning to the speaker installed in the location where the suspicious person was detected. Specifically, it connects to the speaker controller and sends an instruction to play the warning audio, "Security being strengthened, security being strengthened."

[0991] Step 7:

[0992] The server then sends alerts to security staff and residents, including video capture of the suspicious individual, their location, and a risk level, allowing the relevant parties to initiate a rapid response.

[0993] Step 8:

[0994] Security staff and residents receive an alert and rush to the scene to identify and respond to suspicious individuals. Specifically, they conduct on-site investigations and take necessary measures based on the information contained in the alert.

[0995] Example 1

[0996] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0997] Conventional security camera systems have difficulty detecting suspicious individuals in real time, making it difficult to take immediate action to prevent crimes. Furthermore, they lack a means to accurately and quickly communicate suspicious individual information to relevant parties, which can result in delays in appropriate response. The present invention aims to solve these problems and provide a system for efficiently and effectively detecting and responding to suspicious individuals.

[0998] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0999] In this invention, the server includes means for receiving video data from security cameras, means for preprocessing the video data, means for analyzing the preprocessed video data and identifying suspicious behavior patterns, means for storing information on a suspicious individual in a database when the suspicious individual is detected, means for outputting an audio warning when the suspicious individual is detected, and means for transmitting information on the suspicious individual to security personnel and users. This enables early detection of suspicious individuals and prompt response, making it possible to prevent crimes from occurring.

[1000] A "security camera" is a device installed to capture video within a surveillance area and deter crimes and suspicious behavior.

[1001] "Video Data" means a digital representation of visual information captured by a security camera.

[1002] "Preprocessing" refers to initial processing such as adjusting resolution and removing noise, which is performed to make video data easier to analyze.

[1003] A "behavioral pattern" is an identified series of actions or characteristics of behavior taken by an object in video data.

[1004] A "database" is a system for efficiently storing, managing, and retrieving information.

[1005] A "voice warning" is a voice message that alerts suspicious people and people in the area.

[1006] An "alert message" is a notification containing information about a suspicious person, and is sent promptly to security personnel and users.

[1007] "Image processing technology" refers to a set of computer techniques for analyzing, transforming, and understanding digital images.

[1008] A "learning algorithm" is an algorithm that learns patterns and knowledge from data and makes predictions and classifications.

[1009] "Suspicious person information" is detailed information such as the suspicious person's behavior, location, time, and danger level.

[1010] "Location data" is information that indicates the geographical location of an object or person.

[1011] The "risk level" is an index that indicates the degree of risk based on the behavior of a suspicious person and the situation.

[1012] MODE FOR CARRYING OUT THE INVENTION

[1013] The purpose of this invention is to detect suspicious individuals and prevent crimes before they occur by using a system that combines security cameras and generative AI. This system consists of security cameras, a server, a speaker, and a terminal.

[1014] The server receives video data from security cameras in real time. It is always connected via the network and receives the video data captured by the security cameras sequentially. A multipurpose digital signal processing device is used to analyze the video data. At this stage, the video data is unprocessed raw data.

[1015] The server then preprocesses the received video data. This preprocessing involves using image processing techniques such as the OpenCV library to adjust the resolution and remove noise. Specifically, the video is converted to grayscale and a noise reduction filter is applied. This makes subsequent analysis easier.

[1016] After preprocessing is complete, the server performs object detection and behavioral pattern analysis using YOLO (You Only Look Once) or similar object detection algorithms. The analysis is performed frame by frame to identify people who linger for long periods of time or who behave abnormally.

[1017] If a suspicious person is detected, the server stores the information in a database, including the person's location, time, and behavioral patterns. For example, it connects to an SQL database and stores the information using an INSERT statement.

[1018] The server also issues an audio warning to detected suspicious individuals. It sends an HTTP request to the speaker and uses a speech synthesis library to output an audio warning such as "Security is being strengthened." This audio warning alerts not only the suspicious individual but also those around them.

[1019] The server then sends an alert message to the security personnel or the user's device, which includes a video capture of the suspicious person, their location, and the danger level. For example, the alert message can be sent via email or SMS API (e.g., Twilio).

[1020] Specific examples

[1021] For example, video data from a security camera installed at a station is sent to a server, and after analysis, person A is detected behaving suspiciously near a bench. The server determines that person A's behavior is abnormal and stores this information in a database. The server then issues an instruction to a speaker to play a warning audio message such as "Security is being strengthened." At the same time, an alert is sent to the security officer's device, providing a video capture, location information, and danger level. This allows security officers to respond to the scene quickly.

[1022] Example prompts for generative AI models

[1023] 1. "How can I analyze security camera footage at a station to detect suspicious individuals?"

[1024] 2. "How can I build an automation system that detects suspicious activity from security cameras and sends audio warnings and alerts?"

[1025] As described above, this invention utilizes security cameras and generative AI to enable early detection of suspicious individuals and rapid response. Using this system can greatly contribute to crime prevention in a variety of locations.

[1026] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1027] Program processing steps

[1028] Step 1:

[1029] The server receives video data from security cameras in real time. It takes the video stream (e.g., RTSP) from the security cameras as input and stores the video data in the server's memory. The output at this stage is the raw video data stored in memory.

[1030] Step 2:

[1031] The server preprocesses the received video data. The input is the video data saved in step 1, and the server uses the OpenCV library to adjust the resolution and remove noise. Specifically, it converts the video to grayscale and applies a noise reduction filter. The output is the preprocessed video data.

[1032] Step 3:

[1033] The server analyzes the preprocessed video data to identify objects and behavioral patterns. The input is the preprocessed video data created in step 2, which is analyzed using an object detection algorithm such as YOLO. Specifically, the video data is analyzed frame by frame to identify the location and movement of people and luggage. The output is data on detected objects and behavioral patterns.

[1034] Step 4:

[1035] When a suspicious person is detected, the server stores the information in a database. The input is the suspicious person information detected in step 3, and the server connects to an SQL database and uses the INSERT statement to store information such as the suspicious person's location, time, and behavioral patterns. The output is a new record stored in the database.

[1036] Step 5:

[1037] The server issues an audio warning if a suspicious person is detected. The input is the suspicious person information detected in step 3, and an HTTP request is sent to the speaker, which uses a Python speech synthesis library to play the audio message "Security is being strengthened." The output is an audio warning emitted from the speaker.

[1038] Step 6:

[1039] The server sends suspicious person information to the device of the security officer or user. The input is the suspicious person information detected in step 3, and an alert message is sent using email or an SMS API (for example, Twilio). The alert includes a video capture image, location information, and danger level. The output is an alert message that is displayed on the device of the security officer or user.

[1040] The above are the specific processing steps of this system. We have explained in detail how the input data is processed and what kind of output is obtained at each step. We have also clearly stated what specific operations are performed at each step. This makes the overall flow and function of this system clear.

[1041] (Application example 1)

[1042] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1043] Ensuring safety in autonomous vehicles is an important issue. In particular, there is a need to detect suspicious or dangerous behavior in the vehicle early and take appropriate measures promptly. However, conventional methods make it difficult to detect suspicious behavior in the vehicle in real time and respond immediately. For this reason, a new system is needed to improve safety in autonomous vehicles.

[1044] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1045] In this invention, the server includes means for receiving video data from security cameras, means for analyzing the video data and detecting suspicious behavior, means for outputting an audio warning when a suspicious individual is detected, means for transmitting suspicious individual information to security personnel and residents, means for receiving video data from an in-vehicle camera of the autonomous vehicle and detecting suspicious behavior of passengers, means for issuing an audio warning from an in-vehicle speaker, and means for transmitting an alert to an operation control center including video capture of the suspicious individual, location information, and danger level. This enables suspicious behavior in the autonomous vehicle to be quickly detected and an immediate response to be taken.

[1046] A "security camera" is a recording device that captures video data within a specific area and is used for safety and crime prevention purposes.

[1047] "Video data" refers to images and video information obtained from security cameras, in-car cameras, etc.

[1048] "Analysis" refers to the methods and techniques used to process video data frame by frame and understand and distinguish its content.

[1049] "Suspicious activity" refers to movements or behavior that deviate from normal patterns of behavior and may indicate dangerous or criminal activity.

[1050] A "voice warning" is a voice message that is emitted from a speaker to alert and warn people when suspicious behavior is detected.

[1051] "Security personnel" refers to professional people with safety and security responsibilities.

[1052] "Residents" refers to people who live in a particular building or area.

[1053] An "in-vehicle camera" is a camera device installed inside an autonomous vehicle to monitor passengers and the situation inside the vehicle.

[1054] "Passenger" refers to people using the self-driving vehicle.

[1055] "Operation control center" means a central control facility or organization for monitoring and managing the operation status of autonomous vehicles.

[1056] "Video capture" refers to obtaining a still image of a video at a specific moment.

[1057] "Location information" is data that indicates the geographic location of a particular object or person.

[1058] The "risk level" is a scale for evaluating the degree of danger of detected suspicious behavior or individuals.

[1059] An "alert" is a notification system for issuing warnings and alerts.

[1060] To realize this invention, the following system configuration and program are required. First, video data from security cameras and in-car cameras is received in real time by a server. The server analyzes this video data and uses machine learning algorithms to detect suspicious behavior or individuals. If suspicious behavior or individuals are detected, the server immediately and automatically takes appropriate action.

[1061] The system uses the following hardware and software:

[1062] In-car cameras and security cameras (any camera device)

[1063] Server (high performance computer)

[1064] OpenCV (image processing library)

[1065] YOLO (You Only Look Once) (machine learning model)

[1066] Warning audio playback system (speaker)

[1067] Alert sending system (network connection)

[1068] The server analyzes the video data sent from the camera in real time, analyzing each frame using the YOLO model, which is a technology that detects people and other objects in images quickly and accurately. Using this model, specific behavioral patterns and abnormal behavior can be quickly identified.

[1069] If suspicious behavior is detected, the server plays an audio warning through the speaker and simultaneously sends an alert to the operation control center. The alert includes a video capture of the suspicious person, their location, and a danger level, allowing the operation control center staff to immediately check the situation and take appropriate measures.

[1070] As a concrete example, consider its use in an autonomous vehicle. For example, if an in-car camera detects suspicious activity over a long period of time near the driver's seat, the server analyzes the information and issues an audio warning such as "Security is being strengthened," while simultaneously sending an alert to the operation control center containing details of the suspicious activity. This alert includes a video capture of the scene, the vehicle's location, and a danger level, allowing the administrator to take prompt action.

[1071] In addition, specific examples of prompts for system design using generative AI models are as follows:

[1072] "We need to analyze the in-car footage captured by the camera and detect suspicious individuals. Please help us develop a system that analyzes passenger behavior using the YOLO model and maintains safety."

[1073] As described above, the present invention provides an effective means for enhancing passenger safety in autonomous vehicles, enabling early detection of suspicious behavior and rapid response, thereby ensuring safety inside the vehicle.

[1074] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1075] Step 1:

[1076] The server receives video data from security cameras and in-car cameras in real time.

[1077] Input: Video stream from a camera device.

[1078] Output: Real-time video data.

[1079] How it works: The camera device captures video and sends the data over the network to a server, which receives it and prepares it for analysis.

[1080] Step 2:

[1081] The server analyzes the received video data frame by frame to detect suspicious behavior or people.

[1082] Input: Real-time video data.

[1083] Output: Detection results of suspicious behavior or people.

[1084] How it works: The server uses the YOLO model to analyze video frames, identify people and objects in each frame, and record any specific behavioral patterns or abnormal behaviors.

[1085] Step 3:

[1086] If suspicious behavior is detected, the server obtains video capture and location information to form suspicious person information.

[1087] Input: Suspicious behavior or person detection results.

[1088] Output: Video capture, location information, suspicious person information.

[1089] Specific operation: The server captures frames of any detected suspicious behavior or person and obtains location information such as GPS data. This information is then integrated to generate information about the suspicious person.

[1090] Step 4:

[1091] When a suspicious person is detected, the server plays a warning sound from the car's speakers.

[1092] Input: Suspicious person information.

[1093] Output: Warning audio.

[1094] Specific operation: The server sends instructions to the speaker to play a pre-prepared warning sound such as "Security is being strengthened," thereby alerting passengers and suspicious people inside the vehicle.

[1095] Step 5:

[1096] The server sends an alert containing information about the suspicious person to the operation control center.

[1097] Input: Suspicious person information.

[1098] Output: The alert message.

[1099] Specific operation: The server sends an alert message to the operation control center, including a video capture of the suspicious person, their location, and the danger level. Operation control personnel can use this information to take prompt action.

[1100] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1101] The present invention provides a system that combines security cameras, generative AI, and an emotion engine to recognize the user's emotions, detect suspicious individuals, and prevent crimes before they occur.

[1102] System Configuration

[1103] 1. Security camera and server integration

[1104] The server receives video data from security cameras in real time and uses computer vision techniques, machine learning algorithms, and an emotion engine to process the video data.

[1105] 2. Analysis of video data

[1106] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns or suspicious movements. For example, it can detect people who stay in a particular location for a long time or people who leave their luggage behind. This analysis uses technologies such as YOLO (You Only Look Once) and OpenCV.

[1107] 3. Emotional Recognition

[1108] The server analyzes the user's facial expressions from the video data and estimates their emotions. This emotion analysis uses an emotion engine (e.g., Facial Emotion Recognition). This allows it to identify whether the user is feeling anxiety, fear, anger, or other emotions.

[1109] 4. Identifying suspicious individuals and storing their information

[1110] When a suspicious individual is detected, the server stores the information in a database, including the individual's location, time, behavioral patterns, and perceived emotions, providing information for later appropriate response.

[1111] 5. Audio warning transmission and adjustment

[1112] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which plays an audio warning such as "Security is being strengthened." This warning not only alerts the suspicious person, but also other people in the area. Furthermore, the content and tone of the audio warning are adjusted appropriately based on the user's emotions recognized by the emotion engine.

[1113] 6. Sending Alerts

[1114] The server sends alerts to security staff and residents when a suspicious individual is detected, including video captures of the individual, their location, danger level, and perceived emotion, allowing relevant parties to take immediate action.

[1115] Specific examples

[1116] For example, suppose that footage from a security camera installed at a train station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server then detects the emotion of fear from Person A's facial expression. Based on this, the server determines that not only is Person A's behavior abnormal, but that fear is also behind that behavior.

[1117] This information is stored in a database, and the server sends a command to a speaker next to the camera, playing a warning message saying "security is being strengthened" and adjusting the tone of the message to a gentler tone. At the same time, an alert is sent to security staff, informing them that the video capture includes location information, danger level, and fear emotion. This allows security staff to respond to the scene quickly and with the underlying fear factor in mind.

[1118] In this way, by combining emotion engines, it is possible to create a system that can prevent crimes while also responding appropriately to the situation.

[1119] Natural language description of the program

[1120] The system's program analyzes video data sent from security cameras in real time to not only detect suspicious individuals but also analyze user emotions. Specifically, the server receives the video data and analyzes it using machine learning algorithms and an emotion engine. If suspicious behavior or a specific emotion is detected, the server plays an audio warning through a speaker and simultaneously sends an alert to security staff and residents. The alert includes the suspicious individual's video capture, location information, danger level, and recognized emotion.

[1121] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[1122] The processing flow will be explained below.

[1123] Step 1:

[1124] The server receives video data from the security camera in real time, specifically by continuously capturing video frames using a streaming protocol (e.g., RTSP).

[1125] Step 2:

[1126] The server analyzes the received video data frame by frame. Specifically, the video frames are input into a machine learning algorithm (such as YOLO or OpenCV) to identify the position and movement of people.

[1127] Step 3:

[1128] The server detects specific patterns of behavior or suspicious activity by using anomaly detection algorithms to identify behavior that differs from normal patterns (such as prolonged stays, rapid movements, or unnatural changes in location).

[1129] Step 4:

[1130] The server analyzes the user's facial expressions from the video data and estimates their emotions. Specifically, it uses an emotion engine to identify emotions such as anxiety, fear, and anger from the user's facial expressions.

[1131] Step 5:

[1132] The server determines whether a person is suspicious. Specifically, it evaluates both behavioral patterns and emotions to identify suspicious individuals who meet pre-defined criteria (e.g., long stay + emotion of fear).

[1133] Step 6:

[1134] The server stores the suspect's information in a database, including location, time, behavioral details, and perceived emotions, for use as data for future responses.

[1135] Step 7:

[1136] The server then sends an audio warning to a speaker installed in the location where a suspicious person was detected. Specifically, it connects to a speaker controller and sends instructions to play a warning voice saying, "Security being strengthened, security being strengthened." The tone may be adjusted based on emotion.

[1137] Step 8:

[1138] The server sends alerts to security staff and residents, including video capture of the suspicious individual, their location, danger level, and perceived emotion, allowing relevant parties to initiate a rapid response.

[1139] Step 9:

[1140] Security staff and residents receive an alert and rush to the scene to identify and respond to suspicious individuals. Specifically, they conduct on-site investigations and take necessary measures based on the information contained in the alert.

[1141] Example 2

[1142] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1143] While conventional security systems can detect suspicious behavior by analyzing video data, they face the challenge of being unable to analyze user emotions and respond appropriately. Furthermore, there are insufficient means for notifying security staff and residents of suspicious person information in real time, making it difficult to respond quickly. Furthermore, there is a lack of efficient and accurate means for storing and managing suspicious person information.

[1144] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving video data from a security camera, means for analyzing the video data and detecting suspicious behavior, means for analyzing a user's emotion from the video data, means for outputting an audio warning when a suspicious person is detected, means for transmitting suspicious person information to security staff or residents, means for saving the suspicious person information in a database, means for dividing the video data into frames and performing preprocessing, means for detecting suspicious behavior using a machine learning algorithm, means for estimating a user's emotion using an emotion engine, and means for adjusting the tone of the audio warning as appropriate. This enables the security system to comprehensively analyze the behavior and emotion of a suspicious person and take prompt and appropriate action.

[1145] A "security camera" is a device that monitors a specific area and records and transmits the video data.

[1146] "Video data" is digital data consisting of a series of image frames captured by a security camera.

[1147] A "server" is a computer system for storing, analyzing, and processing data.

[1148] "Analysis" is the process of extracting and evaluating specific information from acquired video data.

[1149] "Suspicious activity" refers to specific unusual or suspicious movements or behaviors in security camera footage.

[1150] "User" refers to a general user, including a person who is the subject of surveillance by a security camera and an administrator who uses the system.

[1151] "Voice warning" is a voice message played through a speaker to alert those around you.

[1152] "Security Staff" means the person or group responsible for monitoring and managing the security system.

[1153] "Resident" means a person who lives in or uses the area in which the security system is installed.

[1154] A "database" is a system for efficiently storing, managing, and retrieving information.

[1155] "Frame segmentation" is the process of breaking down video data into individual still images.

[1156] "Preprocessing" is the process of improving the image quality and removing noise from video data before analysis.

[1157] A "machine learning algorithm" is a learning program that analyzes data and recognizes patterns.

[1158] The "emotion engine" is software that analyzes a user's facial expressions from video data and estimates their emotions.

[1159] "Tone adjustment" is the process of changing the tone and content of an audio alert depending on the situation.

[1160] MODE FOR CARRYING OUT THE INVENTION

[1161] The present invention provides a system that combines security cameras, generative AI models, and an emotion engine to recognize user emotions, detect suspicious individuals, and prevent crimes before they occur.

[1162] System Configuration

[1163] This invention combines various technologies centered around the server, terminal, and user.

[1164] 1. Security camera and server integration

[1165] The server receives video data from security cameras in real time and uses computer vision techniques, machine learning algorithms, and an emotion engine to process the video data.

[1166] 2. Analysis of video data

[1167] The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns and suspicious movements, using technologies such as YOLO (You Only Look Once) and OpenCV.

[1168] 3. Emotional Recognition

[1169] The server analyzes the user's facial expressions from the video data and estimates their emotions. This emotion analysis uses an emotion engine (e.g., Facial Emotion Recognition).

[1170] 4. Identifying suspicious individuals and storing their information

[1171] If a suspicious individual is detected, the server stores the information in a database, including the individual's location, time, behavioral patterns, and perceived emotions.

[1172] 5. Audio warning transmission and adjustment

[1173] When a suspicious person is identified, the server uses that information to send instructions to a speaker installed next to the camera, which issues an audio warning such as "Security is being strengthened." The content and tone of the audio warning are adjusted appropriately based on the user's emotions recognized by the emotion engine.

[1174] 6. Sending Alerts

[1175] The server sends alerts to security staff and residents when a suspicious individual is detected, including video captures of the individual, their location, danger level, and perceived emotion.

[1176] Specific examples

[1177] For example, suppose that footage from a security camera installed in a station is sent to a server, and after analysis, Person A is detected behaving suspiciously near a bench. The server then detects the emotion of fear from Person A's facial expression. Based on this, the server determines that not only is Person A's behavior abnormal, but that fear is also behind that behavior.

[1178] This information is stored in a database, and the server sends a command to the speaker next to the camera, which plays a warning message saying "security is being strengthened" and adjusts the tone of the message to a gentler tone. At the same time, an alert is sent to security staff, informing them that the video capture includes location information, danger level, and fear emotion. This allows security staff to respond to the scene quickly and take into account the underlying fear factor.

[1179] Prompt Sentence Examples

[1180] Based on the system log data, an analysis of security camera footage detected Person A behaving suspiciously near the bench. Furthermore, the emotion of fear was detected. Using this information, what information should be provided to the security staff, and what action should be taken next?

[1181] The above is the "Mode for Carrying Out the Invention" of the present invention. How this system operates has been explained using specific examples.

[1182] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1183] Step 1:

[1184] The server receives video data from security cameras in real time. Specifically, the server receives video streams through a specific IP address and port. The input is the video data from the security cameras, and the output is the video frames stored in the server.

[1185] Step 2:

[1186] The server divides the received video data into frames, extracts the frames using the OpenCV library, and simultaneously performs preprocessing such as noise reduction and image quality improvement. The input is the raw video stream data, and the output is the processed frame images.

[1187] Step 3:

[1188] The server loads a machine learning algorithm (e.g., YOLO) and performs object recognition within each frame. Specifically, it uses the YOLO model to identify people and objects and analyze specific movement and behavior patterns. The input is the preprocessed frame image, and the output is the identified objects and their locations.

[1189] Step 4:

[1190] The server uses an emotion engine (e.g., Facial Emotion Recognition) to analyze the facial expressions of the detected person and estimate their emotions. This allows us to identify the emotions the user is feeling (e.g., anxiety, fear, anger). The input is a person's facial image, and the output is estimated emotional information.

[1191] Step 5:

[1192] The server stores the information of the detected suspicious person (location, time, behavioral patterns, and recognized emotions) in a database. Specifically, it inserts data into a MySQL database. The input is the suspicious person information, and the output is the success / failure status of saving the data to the database.

[1193] Step 6:

[1194] The server sends instructions to the speaker based on the suspicious person information, and issues an audio warning saying "Security is being strengthened." The tone and content of the audio are adjusted based on the results of emotion analysis. The input is suspicious person information and emotion information, and the output is the execution of the audio warning.

[1195] Step 7:

[1196] The server sends alerts to security staff and residents, including video captures of suspicious individuals, their location, danger level, and perceived emotions. The alerts are sent via email or SMS. The input is information about the suspicious individual and their emotions, and the output is the sending of an alert message.

[1197] (Application example 2)

[1198] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1199] In modern society, detecting crimes and suspicious individuals is a major issue. Although security cameras are commonly installed, their ability to detect suspicious individuals and analyze abnormal behavior in real time is insufficient, and there are few systems that can recognize emotions. As a result, it is difficult for users and security staff to take prompt countermeasures, and effective crime prevention measures are needed.

[1200] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1201] In this invention, the server includes means for receiving video data from security cameras, means for analyzing the video data and detecting suspicious behavior, means for analyzing facial expressions from the video data and recognizing emotions, means for sending a push notification to a smartphone when a suspicious person is detected, means for outputting an audio warning when a suspicious person is detected, and means for sending suspicious person information to security staff and residents. This enables real-time detection of suspicious people and recognition of emotions, enabling rapid and effective response.

[1202] A "security camera" is a device that records and transmits video in real time to monitor for crimes and suspicious activity.

[1203] "Video data" is visual information of a scene or object at a particular time obtained from a security camera.

[1204] "Analysis" is the process of using computer vision techniques and machine learning algorithms to identify specific behavioral patterns and anomalous behavior from video data.

[1205] "Suspicious behavior" is behavior that is abnormal to general patterns of behavior and indicates potential criminal activity or security risks.

[1206] "Facial expression" refers to the expression of emotions that can be read from the movement and arrangement of muscles in a person's face.

[1207] "Emotion recognition" is a technology that analyzes people's facial expressions from video data and estimates emotions such as joy, anger, and fear.

[1208] "Push notification" is a technology that allows specific information to be instantly displayed on the screen of a user's device, such as a smartphone.

[1209] A "voice warning" is a method of issuing an audio warning through a speaker to alert suspicious individuals and those around them.

[1210] "Suspicious person information" is data about a person who has engaged in suspicious behavior, and includes video capture, location information, emotion recognition results, danger level, and the like.

[1211] "Security staff" are people who monitor and patrol buildings and areas to ensure their safety.

[1212] This invention is a system that combines a security camera, generative AI, and an emotion recognition engine, and aims to recognize the user's emotions, detect suspicious individuals, and prevent crimes before they occur.

[1213] System Configuration

[1214] This system consists of the following elements:

[1215] 1. Security camera and server integration: The server receives video data from security cameras in real time. This video data is processed on the server. The hardware used can be a standard IP camera or a high-resolution security camera.

[1216] 2. Video data analysis: The server analyzes the video data received from the security cameras frame by frame to identify specific behavioral patterns and suspicious movements. This analysis uses computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., YOLO).

[1217] 3. Emotion Recognition: The server analyzes the user's facial expressions from the video data and estimates their emotions. For emotion analysis, an emotion recognition engine (e.g., Facial Emotion Recognition) is used. This allows the server to identify whether the user is feeling emotions such as anxiety, fear, or anger.

[1218] 4. Suspicious person detection and alert sending: If the server detects a suspicious person, it stores the information in a database and sends a push notification to the smartphone. It can also output an audio warning through the speaker. This notification includes a video capture of the suspicious person, their location, emotion recognition results, and danger level.

[1219] What the program does

[1220] The server processes video data received from security cameras in real time. First, it uses OpenCV to analyze the video data frame by frame to detect suspicious behavior. Next, it uses YOLO to identify specific behavioral patterns. At the same time, it uses an emotion recognition engine to analyze facial expressions from the video data and classify the user's emotions.

[1221] If a suspicious individual is detected, the information is stored in a database and a push notification is sent to the smartphone application. An audio warning is also issued from the speaker based on instructions from the server. The suspicious individual information includes video capture, location information, emotion recognition results, and danger level, allowing security staff and users to take prompt and appropriate action.

[1222] Specific examples

[1223] For example, if a home security camera detects a person staying in a specific location for a long period of time and recognizes the emotion of fear from the person's facial expression, the server will issue a voice warning in a gentle tone from the speaker saying "Security is being strengthened." At the same time, a push notification will be sent to the smartphone, containing information such as "Location: (104, 230, 150, 276), Emotion: Fear."

[1224] Prompt Sentence Examples

[1225] 1. "Security cameras analyze footage in real time to detect people who stay in the area for long periods of time."

[1226] 2. "Build a system that estimates emotions from facial expressions and issues audio warnings in the event of anxiety or fear."

[1227] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1228] Step 1:

[1229] Receive video data from security cameras

[1230] The server receives video data from security cameras in real time. This video data is the input for the entire system. Image data is sent to the server frame by frame, allowing it to proceed to the next analysis step.

[1231] Step 2:

[1232] Video data analysis

[1233] The server analyzes the received video data frame by frame. This analysis uses computer vision technology (OpenCV) and machine learning algorithms (YOLO). It detects people and objects in the frames and identifies suspicious movements and behavior patterns. The input is the video data for each frame, and the output is information on the location and behavior patterns of detected suspicious individuals.

[1234] Step 3:

[1235] Facial Expression Analysis and Emotion Recognition

[1236] The server analyzes the facial expressions of the detected person and uses an emotion recognition engine to recognize their emotions. Specifically, it cuts out the face region and inputs it into the Facial Emotion Recognition model. The data input used here is the image data of the detected face, and the output is the determined emotion (e.g., fear, anger, joy, etc.).

[1237] Step 4:

[1238] Saving suspicious person information

[1239] The server stores information about the detected suspicious individuals (location, behavioral patterns, emotions, etc.) in a database. This information is later used for analysis and countermeasures. The data input is the analysis result, and the record in the database is the output.

[1240] Step 5:

[1241] Sending push notifications

[1242] If a suspicious person is detected, the server sends a push notification to the smartphone. This notification includes the suspicious person's video capture, location information, emotion recognition results, danger level, etc. The data input is the stored suspicious person information, and the output is the notification sent to the smartphone.

[1243] Step 6:

[1244] Sending a voice warning

[1245] The server sends instructions to the speaker to play an audio warning such as "Security is being strengthened." The content and tone of the audio warning are adjusted based on the recognized emotion. The data input is the analysis results and emotion recognition results, and the output is the audio warning emitted from the speaker.

[1246] Step 7:

[1247] Sending alerts

[1248] The server sends alerts containing suspicious person information to security staff and residents. The alerts include video capture, location information, emotion recognition results, and danger level. The data input is the stored suspicious person information, and the output is the sending of an alert message.

[1249] In this way, the system can detect suspicious individuals in real time and take necessary measures immediately.

[1250] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1251] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1252] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1253] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1254] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1255] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1256] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1257] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1258] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1259] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1260] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1261] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1262] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1263] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1264] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1265] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1266] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1267] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1268] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1269] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1270] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1271] The following is further disclosed regarding the above embodiment.

[1272] (Claim 1)

[1273] means for receiving video data from a security camera;

[1274] A means for analyzing video data and detecting suspicious behavior;

[1275] means for outputting an audio warning when a suspicious person is detected;

[1276] A system that includes a means of transmitting suspicious person information to security staff and residents.

[1277] (Claim 2)

[1278] 10. The system of claim 1, comprising computer vision techniques and machine learning algorithms for detecting specific behavioral patterns from video data.

[1279] (Claim 3)

[1280] The system of claim 1, wherein the suspicious person information is transmitted in the form of an alert message including a video capture, location information, and a danger level.

[1281] "Example 1"

[1282] (Claim 1)

[1283] means for receiving video data from a security camera;

[1284] means for preprocessing the video data;

[1285] means for analyzing the pre-processed video data to identify suspicious behavioral patterns;

[1286] A means for storing information about a suspicious person in a database when that person is detected;

[1287] means for outputting an audio warning when a suspicious person is detected;

[1288] A means for transmitting suspicious person information to security personnel and users;

[1289] A system including:

[1290] (Claim 2)

[1291] 10. The system of claim 1, including image processing techniques and learning algorithms for detecting specific behavioral patterns from video data.

[1292] (Claim 3)

[1293] The system of claim 1, wherein the suspicious person information is transmitted in the form of an alert message including a video capture, location data, and a danger level.

[1294] "Application Example 1"

[1295] (Claim 1)

[1296] means for receiving video data from a security camera;

[1297] A means for analyzing video data and detecting suspicious behavior;

[1298] means for outputting an audio warning when a suspicious person is detected;

[1299] A means for transmitting suspicious person information to security personnel and residents;

[1300] A means for receiving video data from an in-vehicle camera of an autonomous vehicle and detecting suspicious passenger behavior;

[1301] means for issuing an audio warning through an in-vehicle speaker;

[1302] A system that includes a means to capture video of suspicious individuals, send location information, and alerts including danger levels to the operation control center.

[1303] (Claim 2)

[1304] 10. The system of claim 1, comprising computer vision techniques and machine learning algorithms for detecting specific behavioral patterns from video data.

[1305] (Claim 3)

[1306] The system of claim 1, wherein the suspicious person information is transmitted in the form of an alert message including a video capture, location information, and a danger level.

[1307] "Example 2: Combining Emotion Engines"

[1308] (Claim 1)

[1309] means for receiving video data from a security camera;

[1310] A means for analyzing video data and detecting suspicious behavior;

[1311] A means for analyzing user emotions from video data;

[1312] means for outputting an audio warning when a suspicious person is detected;

[1313] A means of transmitting suspicious person information to security staff and residents;

[1314] A means for storing suspicious person information in a database;

[1315] A means for dividing video data into frames and performing preprocessing;

[1316] a means for detecting suspicious activity using machine learning algorithms;

[1317] means for estimating a user's emotion using an emotion engine;

[1318] The system includes a means for adjusting the tone of the audio warning at appropriate times.

[1319] (Claim 2)

[1320] 10. The system of claim 1, including image processing techniques and machine learning algorithms for detecting specific behavioral patterns from video data.

[1321] (Claim 3)

[1322] The system of claim 1, wherein the suspicious person information is transmitted in the form of an alert message including a video capture, location information, danger level, and recognized emotion.

[1323] "Application example 2 when combining emotion engines"

[1324] Rewritten claims

[1325] (Claim 1)

[1326] means for receiving video data from a security camera;

[1327] A means for analyzing video data and detecting suspicious behavior;

[1328] A means of analyzing facial expressions from video data and recognizing emotions,

[1329] A means of sending a push notification to a smartphone when a suspicious person is detected;

[1330] means for outputting an audio warning when a suspicious person is detected;

[1331] A system that includes a means of transmitting suspicious person information to security staff and residents.

[1332] (Claim 2)

[1333] 10. The system of claim 1, comprising computer vision techniques and machine learning algorithms for detecting specific behavioral patterns from video data.

[1334] (Claim 3)

[1335] The system of claim 1, wherein the suspicious person information is transmitted in the form of an alert message including a video capture, location information, emotion recognition results, and a danger level. [Explanation of symbols]

[1336] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving video data from a security camera; A means for analyzing video data and detecting suspicious behavior; means for outputting an audio warning when a suspicious person is detected; A system that includes a means of transmitting suspicious person information to security staff and residents.

2. 10. The system of claim 1, comprising computer vision techniques and machine learning algorithms for detecting specific behavioral patterns from video data.

3. The system according to claim 1, wherein the suspicious person information is transmitted in the form of an alert message including a video capture, location information, and a danger level.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A