System
The system addresses theft detection in unmanned stores by analyzing surveillance camera data for prohibited behavior, issuing warnings, and capturing facial images, enhancing security and response efficiency.
Patent Information
- Application Number
- JP2024130348
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-19
AI Technical Summary
Conventional measures are inadequate for real-time theft detection and prevention in unmanned stores, lacking efficient automated video analysis and timely warnings, leading to inefficiencies and increased security risks.
A system that acquires video data from surveillance cameras, analyzes customer behavior into text data, issues warnings for prohibited actions, activates facial recognition cameras, and stores evidence for later access by store managers.
Enables real-time theft detection and rapid response, improving store safety by automating security measures and providing reliable evidence for store managers.
Smart Images

Figure 2026028050000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional measures are not sufficient to prevent theft in unmanned stores. In particular, it is currently difficult to detect theft in real time, issue prompt warnings, and report the incident to the police early. Furthermore, with conventional security camera systems, video analysis and management is done manually, which is inefficient and requires a lot of human resources. This leaves unmanned store owners constantly worried about theft, and creates the problem of insuring the safety of their store operations. [Means for solving the problem]
[0005] In order to solve the above-mentioned problems, the present invention provides the following means. A system is provided that includes means for acquiring video data from surveillance cameras installed in a store and means for analyzing the acquired video data and converting customer behavior into text data. The system also includes means for issuing a warning if specific prohibited behavior is detected based on the text data, and means for notifying the warning to a communication terminal of a store manager. The system further includes means for activating a facial recognition camera installed near the entrance and acquiring facial video of the customer simultaneously with the issuance of the warning. The system also includes means for saving the video data acquired from the surveillance cameras and facial recognition cameras, and means for saving the text data and video data in storage accessible to the store manager, thereby significantly improving the efficiency and effectiveness of security measures in unmanned stores.
[0006] A "surveillance camera" is a photographing device installed in a store to capture video data.
[0007] "Video data" refers to video footage captured by surveillance cameras or facial recognition cameras.
[0008] "Analysis" is the application of algorithms to process captured video data and automatically understand customer behavior.
[0009] "Customer behavior" refers to a series of actions that indicate the customer's movements and operations within the store.
[0010] "Text data" is data that expresses the analyzed customer behavior in text form.
[0011] "Prohibited Behavior" refers to certain pre-defined behaviors, including misconduct.
[0012] "Warning" refers to a notification or alert issued when prohibited behavior is detected.
[0013] A "communication terminal" is a device used to receive information via communication means such as LINE or telephone, and is usually owned by the store manager.
[0014] A "face recognition camera" is a photographic device installed near the entrance to capture images of customers' faces.
[0015] "Acquisition" refers to the act of collecting video data using a device such as a camera.
[0016] "Saving" means storing the acquired video data and text data in a storage device.
[0017] "Storage" refers to the hardware and software mechanisms for storing data. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention is a security system designed to prevent theft in unmanned stores. This system acquires video data from surveillance cameras installed in the store, analyzes the video data, and converts customer behavior into text data. If a specific prohibited behavior is detected, the system issues a warning and notifies the store manager's communication terminal. The system also has the function of activating a face recognition camera installed near the entrance and capturing a video of the customer's face at the same time as issuing the warning.
[0040] Program processing
[0041] Video data acquisition and analysis
[0042] The server collects video data in real time from surveillance cameras installed in the store, then passes the collected video data to a generative AI model, which analyzes customer behavior and converts it into text data.
[0043] Examples:
[0044] The camera captures users entering the store and browsing the shelves.
[0045] The server inputs the video into a generative AI model and outputs the text data, "Customer browsing the shelves."
[0046] NG word detection and notification
[0047] The server analyzes the text data created by the generative AI model and detects whether it contains specific prohibited behavior (e.g., "no money was inserted"). If prohibited behavior is detected, the server issues an alert and sends a notification to the store manager's communication device (e.g., LINE or phone).
[0048] Examples:
[0049] If the text data records that "the customer picked up the product and headed for the entrance without going through the cash register."
[0050] The server determines that this behavior is prohibited and sends a warning notification via LINE or phone stating, "The customer may have stolen a product."
[0051] Activating the face recognition camera and acquiring video
[0052] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[0053] Examples:
[0054] The server sends a warning notification and at the same time sends a command to activate the facial recognition camera.
[0055] The camera captures the customer's face and sends the image to a server.
[0056] The server stores the acquired facial images.
[0057] Information storage and access
[0058] The server stores all captured video data, generated text data, and log data when prohibited behavior is detected. This data can be accessed later by store managers and provided to external agencies such as the police.
[0059] Examples:
[0060] The server stores the video data and text data when prohibited behavior is detected in dedicated storage.
[0061] Store managers can later access the storage and download the data if necessary, and provide it to police.
[0062] This security system will enable the prevention of theft in unmanned stores and a rapid response, significantly improving the safety of store operations.
[0063] The processing flow will be explained below.
[0064] Step 1:
[0065] The server receives video data in real time from surveillance cameras installed in the store, allowing the situation inside the store to be constantly monitored.
[0066] Specific behavior:
[0067] Surveillance cameras capture footage.
[0068] The server receives video data from the surveillance camera and stores it in a buffer.
[0069] Step 2:
[0070] The server passes the captured video data to a generative AI model, which analyzes customer behavior in the video. This analysis converts the video into text data.
[0071] Specific behavior:
[0072] The server inputs the video data frame by frame into the generative AI model.
[0073] The generative AI model analyzes the frames and outputs text data such as "A customer is browsing the shelves."
[0074] The server stores the text data.
[0075] Step 3:
[0076] The server analyzes the text data generated by the generative AI model to detect whether it contains certain prohibited behaviors, which triggers the next action.
[0077] Specific behavior:
[0078] The server analyzes the text data sentence by sentence.
[0079] Based on the analysis results, it is checked whether prohibited actions (e.g., "did not put money in") are included.
[0080] If a prohibited action is detected, the information is recorded in an internal flag.
[0081] Step 4:
[0082] If a prohibited activity is detected, the server issues an alert and sends a notification to the store manager's communication terminal, which includes details of the prohibited activity.
[0083] Specific behavior:
[0084] The server generates a warning message based on an internal flag.
[0085] Using the LINE API, a message is sent to the store manager's LINE account stating, "A customer may have stolen an item."
[0086] In the case of telephone notification, an automated voice warning message will be sent.
[0087] Step 5:
[0088] The server automatically activates a facial recognition camera installed near the entrance to capture images of the customer's face, which can then be used as evidence.
[0089] Specific behavior:
[0090] The server sends a start command to the face recognition camera.
[0091] A facial recognition camera captures the customer's face and sends the video data to a server.
[0092] The server stores the acquired facial images.
[0093] Step 6:
[0094] The server stores the captured video data, text data, and log data when prohibited behavior is detected, which will be used for later analysis and reporting to the police.
[0095] Specific behavior:
[0096] The server stores the video data and text data when prohibited behavior is detected in dedicated storage.
[0097] Video data obtained from facial recognition cameras will also be stored in the same way.
[0098] The stored data is accessible to store managers.
[0099] Step 7:
[0100] The user (store manager) receives the notification from the server and reports it to the police if necessary. The user accesses the data on the server and provides the necessary video and text data to the police.
[0101] Specific behavior:
[0102] The store manager will check the warning notification received via LINE or phone.
[0103] The store manager accesses the data stored on the server and downloads the video and text data of the theft.
[0104] If necessary, provide this data to the police.
[0105] The above is a specific processing flow of the system of the present invention.
[0106] Example 1
[0107] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0108] In recent years, the risk of theft has increased with the increase in unmanned stores. However, current surveillance systems have difficulty detecting theft in real time and responding quickly. Furthermore, conventional systems lack a means to reliably and quickly notify store managers, making it difficult to prevent theft. This poses a risk to the safety of store operations.
[0109] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0110] In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning if specific prohibited behavior is detected based on the text data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance simultaneously with the issuance of the warning to acquire a face image of the customer, means for inputting the acquired video data and text data into a generative AI model, and means for detecting whether specific prohibited behavior is included based on the analysis results of the generative AI model. This enables real-time detection of theft in unmanned stores and rapid response, significantly improving the safety of store operations.
[0111] A "surveillance camera" is a device that is installed for the purpose of monitoring a specific location and acquires video data.
[0112] "Video data" refers to image information and video information captured by a surveillance camera.
[0113] A "generative AI model" is a model that uses artificial intelligence to analyze input data and generate text or other information as output.
[0114] "Text data" refers to character information obtained by analyzing video data.
[0115] "Prohibited behavior" refers to specific behaviors that are performed within a store and that violate predetermined rules.
[0116] A "warning" is a notification issued when a prohibited action is detected.
[0117] A "communication terminal" is a device for sending and receiving information, and includes smartphones, tablets, etc.
[0118] A "face recognition camera" is a device that detects an individual's face and acquires its image data.
[0119] "Storage" refers to a storage device or storage service for storing data.
[0120] "Log data" is data that records historical information about events and operations that occur within the system.
[0121] This invention is a security system for preventing theft in unmanned stores, and operates in the following procedure, centered around a server.
[0122] First, the server acquires video data in real time from the surveillance cameras installed in the store. These cameras are often standard IP cameras, such as Hikvision cameras. The server then stores this video data in a specific buffer.
[0123] Next, the server inputs the captured video data frame by frame into a generative AI model, such as OpenAI's GPT-4, which has image analysis capabilities. The server sends the following prompt to the generative AI model:
[0124] "Analyze the image below and explain the customer behavior. Image: [Image data]"
[0125] The generative AI model then analyzes the video data and outputs the customer's behavior as text data, such as "The customer is browsing the shelves."
[0126] The server then analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behavior. Prohibited behavior includes "taking away merchandise without going through the cash register." If prohibited behavior is detected, the server issues a warning and sends a notification to the store manager's communication device. Notifications can be sent using the LINE API or SMS API, and a specific example would be a warning message stating, "A customer may have stolen an item."
[0127] Furthermore, at the same time as issuing the warning, the server activates a facial recognition camera installed near the entrance. This facial recognition camera is equipped with facial recognition software such as Face++ or Amazon Rekognition. The server then sends a capture command to the camera, which captures and stores the customer's facial image.
[0128] Finally, the server stores all captured video data, generated text data, and log data when prohibited behavior is detected in a secure storage location, using Amazon S3 or Google Cloud Storage. Store managers can access this storage and download the data as needed, providing it to external agencies (such as the police).
[0129] In this way, it becomes possible to detect theft in unmanned stores in real time and respond quickly, greatly improving the safety of store operations.
[0130] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0131] Step 1: Acquire video data
[0132] The server acquires video data from surveillance cameras installed in the store. The server periodically accesses the surveillance cameras to acquire new video data. General IP cameras are used for this purpose. The input is the video data acquired from the surveillance cameras, and the output is the video data stored in the temporary storage buffer within the server.
[0133] Example: A server captures video streams from a Hikvision IP camera once per second and stores them in a buffer for analysis.
[0134] Step 2: Analyzing the video data
[0135] The server passes the acquired video data to a generative AI model and converts customer behavior into text data. The server then divides the video data into frames and inputs each frame image into the generative AI model. OpenAI's GPT-4 and other models are used as generative AI models. The input is each frame of video data, and the output is text data.
[0136] Example: The server sends an "image recognition prompt" to OpenAI's GPT-4 and saves the returned text data in an analysis folder.
[0137] Example prompt: "Analyze the image below and describe the customer's behavior. Image: [image data]"
[0138] Step 3: Analyzing the text data
[0139] The server analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behavior. The input is the text data output by the generative AI model, and the output is a flag indicating whether or not prohibited behavior exists. Analysis is performed using natural language processing technology based on the definition of prohibited behavior.
[0140] Example: The server detects text that describes prohibited actions, such as "taking merchandise without going through the cash register."
[0141] Step 4: Issuing warnings and notifications
[0142] If the server detects a prohibited behavior, it issues a warning and sends a notification to the store manager's communication device. The server sends the notification using a pre-configured communication method (e.g., LINE API or SMS API). The input is a flag for the prohibited behavior, and the output is the warning notification sent.
[0143] Example: The server uses the LINE API to send a notification to the store manager saying, "A customer may have stolen an item."
[0144] Step 5: Turn on the face recognition camera
[0145] Immediately after issuing the warning, the server activates a facial recognition camera installed near the entrance and captures the customer's facial image. The server accesses the specified facial recognition camera and sends a capture command. The input is a warning issuance flag, and the output is the captured facial image data.
[0146] Example: The server uses the Face++ API to send a "capture start command" and saves the captured facial image on the server.
[0147] Step 6: Store and access your data
[0148] The server securely stores the captured video data, generated text data, and log data when prohibited behavior is detected. This storage uses Amazon S3 or Google Cloud Storage. The input is various types of data, and the output is data stored in the cloud storage.
[0149] Example: The server uploads video data, text data, and warning logs related to prohibited behavior to Amazon S3 and generates a link for later access.
[0150] The terminal (store manager) accesses this data and provides it to external agencies (e.g., police) if necessary.
[0151] Example: A store manager opens the management app, clicks a URL link, and downloads the necessary video and log data.
[0152] (Application example 1)
[0153] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0154] Preventing theft in unmanned stores and responding quickly and reliably are key challenges. Currently, monitoring in unmanned stores is inefficient, which can lead to delayed detection and countermeasures when theft occurs. Another problem is that information on detected theft is not properly recorded and managed, making it impossible to use as evidence later.
[0155] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0156] In this invention, the server includes: means for acquiring video data from surveillance cameras installed in the store; means for analyzing the acquired video data and converting customer behavior into text data; means for issuing a warning if specific prohibited behavior is detected based on the text data; means for using remote communication technology to notify the store manager's communication terminal of the warning; means for using the server to transmit video data in real time and analyze the text data; means for using a generative AI model to analyze the text data; and means for activating a facial recognition camera installed near the entrance and capturing facial images of customers simultaneously with the issuance of the warning. This enables the immediate detection of theft in an unmanned store and the notification of a warning to the store manager. Furthermore, capturing and storing facial images of customers at the same time can be used as evidence at a later date.
[0157] A "surveillance camera" is a device installed in a store to acquire video data.
[0158] "Video data" refers to video information acquired by a surveillance camera.
[0159] "Analysis" is the process of analyzing customer behavior based on the acquired video data.
[0160] "Text data" is character data generated as a result of analyzing video data.
[0161] "Prohibited Behavior" refers to any behavior that is not permitted within the store.
[0162] A "warning" is a notification issued when prohibited behavior is detected.
[0163] A "communication terminal" is an information and communication device such as a smartphone or computer used by a store manager.
[0164] A "facial recognition camera" is a camera that recognizes the facial characteristics of a specific person and acquires video data.
[0165] A "generative AI model" is a model for generating and analyzing data using artificial intelligence.
[0166] "Telecommunications technology" refers to technology for sending and receiving information to and from communication terminals.
[0167] A "server" is a computer system that processes and stores data.
[0168] "Storage" refers to a storage device for saving data.
[0169] "Real-time" refers to processing that is immediate and without delay.
[0170] The system for implementing this invention consists of surveillance cameras installed in a store, a generative AI model for analyzing video data, a server for detecting prohibited behavior and sending warnings, a communication terminal for the store manager, a facial recognition camera, and storage for data storage.
[0171] Specific system configuration and operation
[0172] 1. Acquiring video data
[0173] The server acquires video data in real time from surveillance cameras installed in the store. The surveillance cameras capture the situation inside the store 24 hours a day and send the video to the server.
[0174] 2. Analysis of video data
[0175] The server inputs the acquired video data into a generative AI model to analyze customer behavior. The generative AI model converts the customer behavior into text data based on the video data. This generative AI model incorporates pre-learned behavioral patterns, allowing it to analyze specific behavior in detail.
[0176] An example of a prompt sentence is, "Please describe in text the actions of the person in the video. In particular, please detect the action of picking up a product and heading towards the entrance / exit without going through the cash register."
[0177] 3. Detecting and warning against prohibited behavior
[0178] The server analyzes the text data created by the generative AI model and detects whether it contains certain prohibited behaviors. For example, if the server detects behavior such as "picking up a product and heading to the entrance without going through the cash register," it issues an alert and immediately sends a notification to the store manager's communication terminal. This notification is sent using remote communication technology.
[0179] 4. Activating the face recognition camera and acquiring video
[0180] When the warning is issued, the server activates a facial recognition camera installed near the entrance to capture the facial image of the customer. The facial recognition camera quickly captures the face of a specific person and sends the image data to the server.
[0181] 5. Data storage and management
[0182] The server stores video data acquired from surveillance cameras and facial recognition cameras, as well as text data created by generative AI models, allowing store managers to access the storage later and download the data as needed to provide it to external agencies such as the police.
[0183] Specific examples
[0184] For example, if a surveillance camera captures a customer picking up an item and then heading for the entrance without going through the cash register, the video data is sent to a server. The server uses a generative AI model to generate text data from the video, such as "The customer picks up an item and heads for the entrance without going through the cash register." Based on this text data, the server detects prohibited behavior and issues a warning to the store manager. At the same time, a facial recognition camera captures the customer's face, and the video data is sent to the server for storage.
[0185] In this way, the system of the present invention makes it possible to prevent and quickly respond to theft in unmanned stores.
[0186] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0187] Step 1:
[0188] Acquiring video data
[0189] Subject: Server
[0190] The server acquires video data in real time from the surveillance cameras installed in the store. These surveillance cameras operate 24 hours a day and continuously transmit images from inside the store to the server. The input is live video from the surveillance cameras, and the output is raw data stored on the server.
[0191] Step 2:
[0192] Video data analysis
[0193] Subject: Server
[0194] The server inputs the acquired video data into the generative AI model and analyzes customer behavior. Specifically, the video data sent from the surveillance camera is sent to the generative AI model, which analyzes "what kind of behavior the customer is exhibiting." The prompt text used is "Please describe in text the behavior of the people in the video. In particular, please detect the behavior of someone picking up a product and heading toward the entrance / exit without going through the cash register." The input is the video data from the surveillance camera, and the output is the analyzed text data.
[0195] Step 3:
[0196] Detecting prohibited behavior
[0197] Subject: Server
[0198] The server analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behaviors. During this process, it checks whether the text data contains pre-set prohibited behavior keywords (e.g., heading to the entrance / exit without going through the cash register). The input is the analyzed text data, and the output is the results of the detection of prohibited behaviors. Specific operations involve the use of a keyword matching algorithm.
[0199] Step 4:
[0200] Warnings and Notifications
[0201] Subject: Server
[0202] If a prohibited behavior is detected, the server issues an alert and sends a notification to the store manager's communication device. This notification is sent using remote communication technologies (e.g., push notification, SMS). The input is the result of the detection of the prohibited behavior, and the output is a warning notification sent to the communication device. The device can then analyze the notification and take appropriate action.
[0203] Step 5:
[0204] Activating the face recognition camera and acquiring video
[0205] Subject: Server
[0206] When the server issues a warning, it activates a facial recognition camera installed near the entrance and captures the customer's facial image. This camera quickly captures the face of a specific person and sends the image data to the server. The input is the warning event, and the output is the facial image data from the facial recognition camera.
[0207] Step 6:
[0208] Data storage and access
[0209] Subject: Server
[0210] The server stores video data acquired from surveillance cameras and facial recognition cameras, as well as text data created by the generative AI model. Store managers can access this data later and provide it to external agencies such as the police if necessary. The input is video data and text data, and the output is the data stored in the storage. Specific operations use a database management system (e.g., SQL, NoSQL).
[0211] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0212] This invention is a security system designed to prevent theft and analyze customer emotions in unmanned stores. This system acquires video data from surveillance cameras installed in the store, analyzes the video data, and converts customer behavior into text data. If specific prohibited behavior is detected, the system then issues a warning and notifies the store manager's communication terminal. The system also has the function of simultaneously issuing a warning and activating a facial recognition camera installed near the entrance to capture video of the customer's face. Furthermore, by incorporating an emotion engine that recognizes user emotions and acquiring and analyzing customer emotion data, the system can improve the accuracy of detecting prohibited behavior.
[0213] Program processing
[0214] Video data acquisition and analysis
[0215] The server collects video data in real time from surveillance cameras installed in the store, then passes the video data to a generative AI model and emotion engine to analyze customer behavior and emotions and convert them into text and emotion data.
[0216] Examples:
[0217] The camera captures users entering the store and browsing the shelves.
[0218] The server inputs the video into a generative AI model and emotion engine, generating text data such as "The customer is browsing the shelves" and emotion data such as "The customer is interested."
[0219] NG word detection and notification
[0220] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether it contains certain prohibited behaviors, which triggers the next action.
[0221] Examples:
[0222] If the text data records that "the customer picked up the product and headed for the entrance without going through the cash register."
[0223] Suppose the emotional data contains information that "the customer is nervous."
[0224] The server determines that this behavior and emotion corresponds to prohibited behavior and sends a warning notification via LINE or phone stating, "The customer may have stolen a product."
[0225] Activating the face recognition camera and acquiring video
[0226] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[0227] Examples:
[0228] The server sends a warning notification and at the same time sends a command to activate the facial recognition camera.
[0229] The camera captures the customer's face and sends the image to a server.
[0230] The server stores the acquired facial images.
[0231] Information storage and access
[0232] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. These data can be accessed later by store managers and provided to external agencies such as the police.
[0233] Examples:
[0234] The server stores video and text data in dedicated storage when prohibited behavior is detected.
[0235] Emotion data generated by the emotion engine is also stored in the same way.
[0236] Store managers can later access the storage and download the data if necessary, and provide it to police.
[0237] This security system will enable the prevention of theft in unmanned stores and a rapid response, significantly improving the safety of store operations. In addition, by using an emotion engine, it will be possible to analyze not only customer behavior but also emotions, enabling more accurate detection of prohibited behavior.
[0238] The processing flow will be explained below.
[0239] Step 1:
[0240] The server collects video data in real time from surveillance cameras installed in the store, allowing it to constantly monitor the situation inside the store.
[0241] Specific behavior:
[0242] Surveillance cameras capture footage.
[0243] The server receives video data from the surveillance camera and stores it in a buffer.
[0244] Step 2:
[0245] The server passes the acquired video data to the generative AI model and emotion engine, which analyzes the customer's behavior and emotions in the video. Based on the analysis results, the video data is converted into text data and emotion data.
[0246] Specific behavior:
[0247] The server inputs the video data frame by frame into the generative AI model and emotion engine.
[0248] The generative AI model outputs the text data "A customer is browsing the shelves."
[0249] The emotion engine generates emotion data that indicates "customer interest."
[0250] The server stores the text data and emotion data.
[0251] Step 3:
[0252] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether it contains specific prohibited behaviors. If prohibited behaviors are detected, the server proceeds to the next step.
[0253] Specific behavior:
[0254] The server analyzes the text data and detects the behavior of "a customer picking up a product and heading to the entrance / exit without going through the cash register."
[0255] The server analyzes the emotional data and confirms that the customer is nervous.
[0256] The server determines whether the behavior is prohibited based on the behavior and emotions.
[0257] Step 4:
[0258] If a prohibited activity is detected, the server issues an alert and sends a notification to the store manager's communication terminal, which includes details of the prohibited activity.
[0259] Specific behavior:
[0260] The server generates a warning message based on the result of the detection of the prohibited behavior.
[0261] Using the LINE API, a message is sent to the store manager's LINE account stating, "A customer may have stolen an item."
[0262] In the case of telephone notifications, an automated voice message will warn, "A customer may have stolen your product."
[0263] Step 5:
[0264] At the same time as issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[0265] Specific behavior:
[0266] The server sends a start command to the face recognition camera.
[0267] A facial recognition camera captures the customer's face and sends the video data to a server.
[0268] The server stores the acquired facial images.
[0269] Step 6:
[0270] The server stores all captured video data, generated text data, and log data when prohibited behavior is detected, and these data can be accessed later by store managers.
[0271] Specific behavior:
[0272] The server stores video and text data in dedicated storage when prohibited behavior is detected.
[0273] Emotion data generated by the emotion engine is also stored in the same way.
[0274] The saved data is stored in a format that can be accessed by the store manager.
[0275] Step 7:
[0276] The user (store manager) receives the notification from the server and reports it to the police if necessary. The user accesses the data on the server and provides the necessary video and text data to the police.
[0277] Specific behavior:
[0278] The store manager will check the warning notification received via LINE or phone.
[0279] The store manager accesses the data stored on the server and downloads the video and text data of the theft.
[0280] If necessary, provide this data to the police.
[0281] The above is a specific processing flow of the system of the present invention, which significantly improves the security function of unmanned stores, suppresses theft, and realizes highly accurate analysis based on customer emotions.
[0282] Example 2
[0283] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0284] To effectively prevent theft in unmanned stores, it is necessary not only to analyze customer behavior but also to accurately grasp their emotions and respond promptly and appropriately. Conventional security systems focus on behavioral analysis and do not consider emotion analysis, resulting in false positives and oversights. Thus, there is a need for methods to improve the accuracy and safety of theft prevention in unmanned stores.
[0285] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0286] In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning when specific prohibited behavior is detected based on the text data and emotion data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance at the same time as issuing the warning and acquiring an image of the customer's face, and means incorporating an emotion recognition engine for analyzing the customer's behavior and emotion and converting it into text data and emotion data. This makes it possible to prevent theft in unmanned stores and to perform highly accurate behavioral analysis based on customer emotions.
[0287] A "surveillance camera" is a device that is installed in a store and continuously captures the situation inside the store as video data.
[0288] "Video Data" means a collection of video frames captured by surveillance cameras or facial recognition cameras, and is digital data containing customer behavior and other significant events.
[0289] "Analysis" is the process of analyzing the captured video data using specific algorithms and AI models to extract customer behavior and emotions.
[0290] "Customer" refers to a shopper or visitor to an unmanned store.
[0291] "Text data" refers to data that converts customer behavior and status obtained as analysis results into text information.
[0292] "Emotion data" is data that indicates the psychological state of a customer, obtained from their facial expressions and actions.
[0293] "Prohibited behavior" refers to stealing merchandise or other behavior that violates store rules.
[0294] "Warning" refers to an alert or notification issued when prohibited behavior is detected.
[0295] A "communication terminal" is a digital device such as a smartphone, tablet, or PC owned by a store manager.
[0296] A "facial recognition camera" is a camera installed to identify customers' faces and acquire video data.
[0297] An "emotion recognition engine" is software or an algorithm for analyzing customer emotions from video data.
[0298] "Storage" refers to cloud-based or physical data storage devices for storing captured data (video data, text data, emotion data, etc.).
[0299] "Store Manager" means a person or organization responsible for the operation and management of an unmanned store.
[0300] The present invention is a security system for preventing theft and analyzing customer sentiment in unmanned stores. The system operates using multiple hardware and software components.
[0301] Hardware Configuration
[0302] This system includes the following main hardware components:
[0303] Surveillance cameras: Installed in stores, they capture customer behavior in real time.
[0304] Facial recognition cameras: Installed near the store entrance, they capture images of customers' faces under certain conditions.
[0305] Communication device: Smartphone, tablet, PC, etc. owned by the store manager.
[0306] Software Configuration
[0307] The operation of the system is supported by the following software:
[0308] Generative AI models: For example, using GPT-4 for natural language processing to convert video data into text data.
[0309] Emotion recognition engine: For example, using Microsoft Azure's emotion analysis API to analyze customer emotions from video data.
[0310] Cloud storage: For example, Google Cloud Storage is used to store the acquired data.
[0311] Processing flow
[0312] The operation of the system proceeds as follows.
[0313] Acquisition of video data: The server acquires video data from the surveillance cameras in real time. The video data is sent to the server using a streaming protocol.
[0314] Video data analysis: The acquired video data is input into a generative AI model and an emotion recognition engine to analyze the customer's behavior and emotions. The analysis results are stored on the server as text data and emotion data.
[0315] Detection of prohibited behavior: Based on the generated text data and emotion data, the server detects whether certain prohibited behaviors are included. If prohibited behaviors are detected, a warning is issued.
[0316] Sending warning notifications: Warnings are sent in real time to the store manager's communication device via the LINE API or Twilio API.
[0317] Activating the facial recognition camera and capturing video: At the same time as receiving the warning notification, the server activates the facial recognition camera and captures the customer's facial video. The captured video data is stored in cloud storage as evidence.
[0318] Information storage: All captured data is stored in a dedicated cloud storage. Store managers can access this data and download it to provide to the police if necessary.
[0319] Specific examples
[0320] For example, imagine a situation where a user enters a store and looks at the shelves. In this case, a surveillance camera captures the user's behavior, and the server inputs the video data into a generative AI model and emotion recognition engine. The analysis results in text data such as "The customer is browsing the shelves" and emotion data such as "The customer is interested." If prohibited behavior is detected, an alert is immediately issued and the facial recognition camera is activated.
[0321] Example prompt sentence:
[0322] "Analyze real-time video footage from security cameras to detect customer behavior and emotions. Then, write a program to issue a warning notification, activate a facial recognition camera, and save the footage if prohibited behavior is detected."
[0323] This invention not only prevents theft in unmanned stores and enables highly accurate analysis of customer behavior, but also takes emotional data into account, enabling more accurate detection of prohibited behavior and quicker responses.
[0324] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0325] Step 1:
[0326] Acquiring video data
[0327] The server receives real-time video data from the surveillance cameras installed in the store. The input is the live video stream from the surveillance cameras. This video data is sent to the server using the RTSP protocol and buffered in memory. The output is a series of video frames.
[0328] What happens: The server connects to the video stream using the camera's RTSP URL and captures video frames, each ready to be sent for analysis.
[0329] Step 2:
[0330] Video data analysis
[0331] The server passes the acquired video data to a generative AI model and an emotion recognition engine to analyze customer behavior and emotions. The input is the video frames acquired in step 1. Data processing is performed by inputting the video frames into a generative AI model (e.g., OpenAI's GPT-4) and an emotion recognition engine (e.g., Microsoft Azure's Sentiment Analysis API). The output is text data and emotion data.
[0332] How it works: The server uses a Python script to input each video frame into a generative AI model, obtaining text data such as "A customer is browsing the shelves." At the same time, it sends the frame to an emotion recognition engine to obtain emotion data such as "The customer is interested."
[0333] Step 3:
[0334] Detecting prohibited behavior
[0335] The server analyzes the generated text data and emotion data to detect whether specific prohibited behaviors are included. The inputs are the text data and emotion data generated in step 2. Data calculations are performed by analyzing this data based on specific algorithms and regular expression patterns. The output is a determination result indicating whether prohibited behaviors have been detected.
[0336] Specific operation: The server performs regular expression pattern matching to detect phrases such as "The customer picked up the product and headed for the entrance without going through the cash register" from the text data. If the emotion data contains information such as "The customer is nervous," it records this as a prohibited behavior.
[0337] Step 4:
[0338] Sending warning notifications
[0339] If the server detects a prohibited behavior, it immediately sends a warning to the store manager's communication terminal. The input is the prohibited behavior judgment result obtained in step 3. Data processing involves generating and sending a warning message. The output is a warning notification sent to the store manager's communication terminal.
[0340] Specific operation: The server uses the LINE API to generate a message such as "A customer may have stolen an item" and sends it to the store manager's LINE account. At the same time, it also sends a phone notification using the Twilio API.
[0341] Step 5:
[0342] Activating the face recognition camera and acquiring video
[0343] Immediately after issuing the warning, the server automatically activates the facial recognition camera installed near the entrance and captures an image of the customer's face. The input is the warning notification from step 4. Data processing involves capturing and saving the facial image. The output is the captured facial image data.
[0344] Specific operation: The server sends the specified API request to the facial recognition camera to activate the camera. The camera captures the customer's face and sends the video to the server in real time. The captured facial video is encoded in the specified format (e.g., MPEG-4) and saved in the server's storage.
[0345] Step 6:
[0346] Retention of Information
[0347] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. The input is all the data generated at each step. Data calculation aggregates this data in one place and stores it in cloud storage. The output is all the stored data.
[0348] Specific operation: When prohibited behavior is detected, the server uploads video data, text data, and emotion data to dedicated cloud storage (e.g., Google Cloud Storage). The data is categorized with a timestamp and can be accessed later by store managers.
[0349] (Application example 2)
[0350] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0351] The challenge is to improve the safety and efficiency of store operations by preventing theft in unmanned stores and analyzing customer sentiment.In particular, since unmanned stores do not have permanent staff on-site, immediate response is difficult, and conventional surveillance systems have limitations in the accuracy of detecting prohibited behavior and the speed of response.
[0352] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning when specific prohibited behavior is detected based on the text data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance simultaneously with the issuance of the warning and acquiring a facial image of the customer, and means for analyzing customer emotion data using an emotion analysis engine to improve the accuracy of detecting prohibited behavior. This makes it possible to prevent theft in unmanned stores and detect prohibited behavior with high accuracy.
[0353] A "surveillance camera" is a device installed in a store to capture video data.
[0354] "Video data" refers to data that records the state of customers and the inside of a store captured by surveillance cameras and facial recognition cameras.
[0355] A "server" is a device or system that analyzes acquired video data and generates behavioral data and emotional data.
[0356] "Text data" is the analysis of video data and the representation of customer behavior as text information.
[0357] "Specific prohibited activities" refer to activities that should not be carried out in the store, including, for example, theft.
[0358] An "alert" is a notification that is sent when certain prohibited behavior is detected.
[0359] A "communication terminal" refers to a device owned by a store manager that is capable of receiving notifications such as warnings.
[0360] A "facial recognition camera" is a camera used to capture and recognize facial images of customers.
[0361] "Emotion data" is data generated by analyzing customer emotions using an emotion analysis engine.
[0362] An "emotion analysis engine" is software or algorithms for analyzing a customer's emotional state from acquired video data.
[0363] "Storage" is a storage device that stores data for a long period of time and allows it to be accessed later.
[0364] In this invention, a system is constructed by integrating multiple pieces of hardware and software to improve the monitoring and security of unmanned stores.
[0365] Hardware and software used
[0366] 1. Hardware
[0367] Surveillance camera: Used to capture video data within the store.
[0368] Facial recognition camera: Used to capture facial images of customers when an alert is issued.
[0369] Communication terminal: A device that allows store managers to receive alert notifications.
[0370] Server: A central device that analyzes data and manages the generated behavioral and emotional data.
[0371] 2. Software
[0372] Generative AI models (e.g., OpenAI's GPT series): Analyze captured video data and convert customer behavior into text data.
[0373] Emotion engine (e.g. Microsoft's Emotion API): Analyzes customer emotion data.
[0374] Notification system (e.g., LINE API, Twilio): Sends warnings to the store manager's communication device based on the prohibited behavior detected.
[0375] System Operation Overview
[0376] Video data acquisition and analysis
[0377] Surveillance cameras capture real-time video data of customer behavior in the store, and the server passes this video data to a generative AI model to generate behavioral data.
[0378] Examples:
[0379] Customers are seen on surveillance cameras picking up items from the shelves.
[0380] The server inputs the video data into a generative AI model and generates behavioral data such as "a customer picking up a product from a shelf."
[0381] The server then uses an emotion engine to generate "emotional data" of the customer from the video data, and obtains the result that "the customer is interested."
[0382] Example prompt for a generative AI model:
[0383] Analyze the video data below and convert customer behavior into text data.
[0384] Video footage: A customer picks up an item from the shelf and heads straight for the exit.
[0385] Output of the generative AI model: The customer picks up the item and heads to the exit without going through the checkout.
[0386] Detecting and notifying prohibited behavior
[0387] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether prohibited behavior is included. If prohibited behavior is detected, the server sends a warning to the store manager's communication device via the notification system.
[0388] Examples:
[0389] Behavioral data records that "the customer picked up the product and headed for the exit without going through the cash register."
[0390] The emotional data includes information that "the customer is nervous."
[0391] The server uses this information to detect prohibited behavior and sends a warning notification saying, "A customer may have stolen a product."
[0392] Activating the face recognition camera and acquiring video
[0393] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture the customer's facial image, which is then saved by the server.
[0394] Examples:
[0395] The server will send a warning notification and simultaneously activate the facial recognition camera.
[0396] The camera captures the customer's face and sends the image to a server.
[0397] The server stores the acquired facial image as evidence.
[0398] Example of a prompt for the emotion engine:
[0399] Analyze the video data below and convert customer sentiment into text data.
[0400] Footage: Customers are seen picking up items and heading for the exit.
[0401] Emotion Engine Output: Customer is nervous.
[0402] Information storage and access
[0403] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. These data can be accessed by store managers and later downloaded and provided to external organizations as needed.
[0404] Examples:
[0405] The server stores the video data, text data, and emotional data when prohibited behavior is detected in dedicated storage.
[0406] Store managers can access the storage and download this data as needed, providing it to external agencies such as the police.
[0407] This system will enable the prevention and rapid response of theft in unmanned stores, significantly improving the safety of store operations. In addition, by using an emotion engine, it is possible to analyze not only customer behavior but also emotions, enabling more accurate detection of prohibited behavior.
[0408] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0409] Step 1:
[0410] The server acquires video data in real time from surveillance cameras installed in the store. Specifically, the cameras capture customer behavior in the store and send the video data to the server. The input data is the video data from the surveillance cameras, and the output data is the video data itself.
[0411] Step 2:
[0412] The server passes the acquired video data to a generative AI model and converts the customer's behavior into text data. Specifically, the server inputs the video data into a generative AI model (e.g., OpenAI's GPT series) and generates behavioral data such as "the customer picking up a product." The input data is video data, and the output data is behavioral data in text format.
[0413] Step 3:
[0414] The server then uses an emotion engine to analyze the customer's emotions from the video data. Specifically, the server inputs the video data into an emotion engine (e.g., Microsoft's Emotion API) and generates emotion data indicating that the customer is interested. The input data is video data, and the output data is emotion data in text format.
[0415] Step 4:
[0416] The server analyzes the generated behavioral and emotional data to detect whether it contains specific prohibited behavior. Specifically, it uses an analysis algorithm to identify prohibited behavior, such as "a customer picks up a product and heads toward the exit without going through the cash register," and also analyzes emotional data, such as "the customer is nervous." The input data are behavioral and emotional data, and the output data are the results of the detection of prohibited behavior.
[0417] Step 5:
[0418] If prohibited behavior is detected, the server sends a warning to the store manager's communication device via the notification system. Specifically, the server uses the LINE API or Twilio to send a warning message to the store manager's smartphone, such as "A customer may have stolen an item." The input data is the detection result, and the output data is the warning notification.
[0419] Step 6:
[0420] Immediately after issuing the warning, the server automatically activates the facial recognition camera installed near the entrance. Specifically, the server sends a start command to the facial recognition camera to start capturing facial images. The input data is the start command, and the output data is facial image data.
[0421] Step 7:
[0422] The facial recognition camera captures the customer's facial image and sends it to a server, which then stores the data. Specifically, the server receives the facial image data and stores it in dedicated storage. This storage can be accessed later. The input data is the customer's facial image, and the output data is the data stored in the storage.
[0423] Step 8:
[0424] The server saves the acquired video data, the generated behavioral and emotional data, and the notification log in storage. Specifically, the server stores each data in a dedicated storage device so that it can be preserved for a long period of time. The input data is the video data, behavioral data, emotional data, and notification log, and the output data is the data saved in storage.
[0425] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0426] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0427] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0428] [Second embodiment]
[0429] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0430] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0431] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0432] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0433] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0434] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0435] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0436] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0437] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0438] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0439] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0440] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0441] The present invention is a security system designed to prevent theft in unmanned stores. This system acquires video data from surveillance cameras installed in the store, analyzes the video data, and converts customer behavior into text data. If a specific prohibited behavior is detected, the system issues a warning and notifies the store manager's communication terminal. The system also has the function of activating a face recognition camera installed near the entrance and capturing a video of the customer's face at the same time as issuing the warning.
[0442] Program processing
[0443] Video data acquisition and analysis
[0444] The server collects video data in real time from surveillance cameras installed in the store, then passes the collected video data to a generative AI model, which analyzes customer behavior and converts it into text data.
[0445] Examples:
[0446] The camera captures users entering the store and browsing the shelves.
[0447] The server inputs the video into a generative AI model and outputs the text data, "Customer browsing the shelves."
[0448] NG word detection and notification
[0449] The server analyzes the text data created by the generative AI model and detects whether it contains specific prohibited behavior (e.g., "no money was inserted"). If prohibited behavior is detected, the server issues an alert and sends a notification to the store manager's communication device (e.g., LINE or phone).
[0450] Examples:
[0451] If the text data records that "the customer picked up the product and headed for the entrance without going through the cash register."
[0452] The server determines that this behavior is prohibited and sends a warning notification via LINE or phone stating, "The customer may have stolen a product."
[0453] Activating the face recognition camera and acquiring video
[0454] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[0455] Examples:
[0456] The server sends a warning notification and at the same time sends a command to activate the facial recognition camera.
[0457] The camera captures the customer's face and sends the image to a server.
[0458] The server stores the acquired facial images.
[0459] Information storage and access
[0460] The server stores all captured video data, generated text data, and log data when prohibited behavior is detected. This data can be accessed later by store managers and provided to external agencies such as the police.
[0461] Examples:
[0462] The server stores the video data and text data when prohibited behavior is detected in dedicated storage.
[0463] Store managers can later access the storage and download the data if necessary, and provide it to police.
[0464] This security system will enable the prevention of theft in unmanned stores and a rapid response, significantly improving the safety of store operations.
[0465] The processing flow will be explained below.
[0466] Step 1:
[0467] The server receives video data in real time from surveillance cameras installed in the store, allowing the situation inside the store to be constantly monitored.
[0468] Specific behavior:
[0469] Surveillance cameras capture footage.
[0470] The server receives video data from the surveillance camera and stores it in a buffer.
[0471] Step 2:
[0472] The server passes the captured video data to a generative AI model, which analyzes customer behavior in the video. This analysis converts the video into text data.
[0473] Specific behavior:
[0474] The server inputs the video data frame by frame into the generative AI model.
[0475] The generative AI model analyzes the frames and outputs text data such as "A customer is browsing the shelves."
[0476] The server stores the text data.
[0477] Step 3:
[0478] The server analyzes the text data generated by the generative AI model to detect whether it contains certain prohibited behaviors, which triggers the next action.
[0479] Specific behavior:
[0480] The server analyzes the text data sentence by sentence.
[0481] Based on the analysis results, it is checked whether prohibited actions (e.g., "did not put money in") are included.
[0482] If a prohibited action is detected, the information is recorded in an internal flag.
[0483] Step 4:
[0484] If a prohibited activity is detected, the server issues an alert and sends a notification to the store manager's communication terminal, which includes details of the prohibited activity.
[0485] Specific behavior:
[0486] The server generates a warning message based on an internal flag.
[0487] Using the LINE API, a message is sent to the store manager's LINE account stating, "A customer may have stolen an item."
[0488] In the case of telephone notification, an automated voice warning message will be sent.
[0489] Step 5:
[0490] The server automatically activates a facial recognition camera installed near the entrance to capture images of the customer's face, which can then be used as evidence.
[0491] Specific behavior:
[0492] The server sends a start command to the face recognition camera.
[0493] A facial recognition camera captures the customer's face and sends the video data to a server.
[0494] The server stores the acquired facial images.
[0495] Step 6:
[0496] The server stores the captured video data, text data, and log data when prohibited behavior is detected, which will be used for later analysis and reporting to the police.
[0497] Specific behavior:
[0498] The server stores the video data and text data when prohibited behavior is detected in dedicated storage.
[0499] Video data obtained from facial recognition cameras will also be stored in the same way.
[0500] The stored data is accessible to store managers.
[0501] Step 7:
[0502] The user (store manager) receives the notification from the server and reports it to the police if necessary. The user accesses the data on the server and provides the necessary video and text data to the police.
[0503] Specific behavior:
[0504] The store manager will check the warning notification received via LINE or phone.
[0505] The store manager accesses the data stored on the server and downloads the video and text data of the theft.
[0506] If necessary, provide this data to the police.
[0507] The above is a specific processing flow of the system of the present invention.
[0508] Example 1
[0509] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0510] In recent years, the risk of theft has increased with the increase in unmanned stores. However, current surveillance systems have difficulty detecting theft in real time and responding quickly. Furthermore, conventional systems lack a means to reliably and quickly notify store managers, making it difficult to prevent theft. This poses a risk to the safety of store operations.
[0511] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0512] In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning if specific prohibited behavior is detected based on the text data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance simultaneously with the issuance of the warning to acquire a face image of the customer, means for inputting the acquired video data and text data into a generative AI model, and means for detecting whether specific prohibited behavior is included based on the analysis results of the generative AI model. This enables real-time detection of theft in unmanned stores and rapid response, significantly improving the safety of store operations.
[0513] A "surveillance camera" is a device that is installed for the purpose of monitoring a specific location and acquires video data.
[0514] "Video data" refers to image information and video information captured by a surveillance camera.
[0515] A "generative AI model" is a model that uses artificial intelligence to analyze input data and generate text or other information as output.
[0516] "Text data" refers to character information obtained by analyzing video data.
[0517] "Prohibited behavior" refers to specific behaviors that are performed within a store and that violate predetermined rules.
[0518] A "warning" is a notification issued when a prohibited action is detected.
[0519] A "communication terminal" is a device for sending and receiving information, and includes smartphones, tablets, etc.
[0520] A "face recognition camera" is a device that detects an individual's face and acquires its image data.
[0521] "Storage" refers to a storage device or storage service for storing data.
[0522] "Log data" is data that records historical information about events and operations that occur within the system.
[0523] This invention is a security system for preventing theft in unmanned stores, and operates in the following procedure, centered around a server.
[0524] First, the server acquires video data in real time from the surveillance cameras installed in the store. These cameras are often standard IP cameras, such as Hikvision cameras. The server then stores this video data in a specific buffer.
[0525] Next, the server inputs the captured video data frame by frame into a generative AI model, such as OpenAI's GPT-4, which has image analysis capabilities. The server sends the following prompt to the generative AI model:
[0526] "Analyze the image below and explain the customer behavior. Image: [Image data]"
[0527] The generative AI model then analyzes the video data and outputs the customer's behavior as text data, such as "The customer is browsing the shelves."
[0528] The server then analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behavior. Prohibited behavior includes "taking away merchandise without going through the cash register." If prohibited behavior is detected, the server issues a warning and sends a notification to the store manager's communication device. Notifications can be sent using the LINE API or SMS API, and a specific example would be a warning message stating, "A customer may have stolen an item."
[0529] Furthermore, at the same time as issuing the warning, the server activates a facial recognition camera installed near the entrance. This facial recognition camera is equipped with facial recognition software such as Face++ or Amazon Rekognition. The server then sends a capture command to the camera, which captures and stores the customer's facial image.
[0530] Finally, the server stores all captured video data, generated text data, and log data when prohibited behavior is detected in a secure storage location, using Amazon S3 or Google Cloud Storage. Store managers can access this storage and download the data as needed, providing it to external agencies (such as the police).
[0531] In this way, it becomes possible to detect theft in unmanned stores in real time and respond quickly, greatly improving the safety of store operations.
[0532] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0533] Step 1: Acquire video data
[0534] The server acquires video data from surveillance cameras installed in the store. The server periodically accesses the surveillance cameras to acquire new video data. General IP cameras are used for this purpose. The input is the video data acquired from the surveillance cameras, and the output is the video data stored in the temporary storage buffer within the server.
[0535] Example: A server captures video streams from a Hikvision IP camera once per second and stores them in a buffer for analysis.
[0536] Step 2: Analyzing the video data
[0537] The server passes the acquired video data to a generative AI model and converts customer behavior into text data. The server then divides the video data into frames and inputs each frame image into the generative AI model. OpenAI's GPT-4 and other models are used as generative AI models. The input is each frame of video data, and the output is text data.
[0538] Example: The server sends an "image recognition prompt" to OpenAI's GPT-4 and saves the returned text data in an analysis folder.
[0539] Example prompt: "Analyze the image below and describe the customer's behavior. Image: [image data]"
[0540] Step 3: Analyzing the text data
[0541] The server analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behavior. The input is the text data output by the generative AI model, and the output is a flag indicating whether or not prohibited behavior exists. Analysis is performed using natural language processing technology based on the definition of prohibited behavior.
[0542] Example: The server detects text that describes prohibited actions, such as "taking merchandise without going through the cash register."
[0543] Step 4: Issuing warnings and notifications
[0544] If the server detects a prohibited behavior, it issues a warning and sends a notification to the store manager's communication device. The server sends the notification using a pre-configured communication method (e.g., LINE API or SMS API). The input is a flag for the prohibited behavior, and the output is the warning notification sent.
[0545] Example: The server uses the LINE API to send a notification to the store manager saying, "A customer may have stolen an item."
[0546] Step 5: Turn on the face recognition camera
[0547] Immediately after issuing the warning, the server activates a facial recognition camera installed near the entrance and captures the customer's facial image. The server accesses the specified facial recognition camera and sends a capture command. The input is a warning issuance flag, and the output is the captured facial image data.
[0548] Example: The server uses the Face++ API to send a "capture start command" and saves the captured facial image on the server.
[0549] Step 6: Store and access your data
[0550] The server securely stores the captured video data, generated text data, and log data when prohibited behavior is detected. This storage uses Amazon S3 or Google Cloud Storage. The input is various types of data, and the output is data stored in the cloud storage.
[0551] Example: The server uploads video data, text data, and warning logs related to prohibited behavior to Amazon S3 and generates a link for later access.
[0552] The terminal (store manager) accesses this data and provides it to external agencies (e.g., police) if necessary.
[0553] Example: A store manager opens the management app, clicks a URL link, and downloads the necessary video and log data.
[0554] (Application example 1)
[0555] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0556] Preventing theft in unmanned stores and responding quickly and reliably are key challenges. Currently, monitoring in unmanned stores is inefficient, which can lead to delayed detection and countermeasures when theft occurs. Another problem is that information on detected theft is not properly recorded and managed, making it impossible to use as evidence later.
[0557] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0558] In this invention, the server includes: means for acquiring video data from surveillance cameras installed in the store; means for analyzing the acquired video data and converting customer behavior into text data; means for issuing a warning if specific prohibited behavior is detected based on the text data; means for using remote communication technology to notify the store manager's communication terminal of the warning; means for using the server to transmit video data in real time and analyze the text data; means for using a generative AI model to analyze the text data; and means for activating a facial recognition camera installed near the entrance and capturing facial images of customers simultaneously with the issuance of the warning. This enables the immediate detection of theft in an unmanned store and the notification of a warning to the store manager. Furthermore, capturing and storing facial images of customers at the same time can be used as evidence at a later date.
[0559] A "surveillance camera" is a device installed in a store to acquire video data.
[0560] "Video data" refers to video information acquired by a surveillance camera.
[0561] "Analysis" is the process of analyzing customer behavior based on the acquired video data.
[0562] "Text data" is character data generated as a result of analyzing video data.
[0563] "Prohibited Behavior" refers to any behavior that is not permitted within the store.
[0564] A "warning" is a notification issued when prohibited behavior is detected.
[0565] A "communication terminal" is an information and communication device such as a smartphone or computer used by a store manager.
[0566] A "facial recognition camera" is a camera that recognizes the facial characteristics of a specific person and acquires video data.
[0567] A "generative AI model" is a model for generating and analyzing data using artificial intelligence.
[0568] "Telecommunications technology" refers to technology for sending and receiving information to and from communication terminals.
[0569] A "server" is a computer system that processes and stores data.
[0570] "Storage" refers to a storage device for saving data.
[0571] "Real-time" refers to processing that is immediate and without delay.
[0572] The system for implementing this invention consists of surveillance cameras installed in a store, a generative AI model for analyzing video data, a server for detecting prohibited behavior and sending warnings, a communication terminal for the store manager, a facial recognition camera, and storage for data storage.
[0573] Specific system configuration and operation
[0574] 1. Acquiring video data
[0575] The server acquires video data in real time from surveillance cameras installed in the store. The surveillance cameras capture the situation inside the store 24 hours a day and send the video to the server.
[0576] 2. Analysis of video data
[0577] The server inputs the acquired video data into a generative AI model to analyze customer behavior. The generative AI model converts the customer behavior into text data based on the video data. This generative AI model incorporates pre-learned behavioral patterns, allowing it to analyze specific behavior in detail.
[0578] An example of a prompt sentence is, "Please describe in text the actions of the person in the video. In particular, please detect the action of picking up a product and heading towards the entrance / exit without going through the cash register."
[0579] 3. Detecting and warning against prohibited behavior
[0580] The server analyzes the text data created by the generative AI model and detects whether it contains certain prohibited behaviors. For example, if the server detects behavior such as "picking up a product and heading to the entrance without going through the cash register," it issues an alert and immediately sends a notification to the store manager's communication terminal. This notification is sent using remote communication technology.
[0581] 4. Activating the face recognition camera and acquiring video
[0582] When the warning is issued, the server activates a facial recognition camera installed near the entrance to capture the facial image of the customer. The facial recognition camera quickly captures the face of a specific person and sends the image data to the server.
[0583] 5. Data storage and management
[0584] The server stores video data acquired from surveillance cameras and facial recognition cameras, as well as text data created by generative AI models, allowing store managers to access the storage later and download the data as needed to provide it to external agencies such as the police.
[0585] Specific examples
[0586] For example, if a surveillance camera captures a customer picking up an item and then heading for the entrance without going through the cash register, the video data is sent to a server. The server uses a generative AI model to generate text data from the video, such as "The customer picks up an item and heads for the entrance without going through the cash register." Based on this text data, the server detects prohibited behavior and issues a warning to the store manager. At the same time, a facial recognition camera captures the customer's face, and the video data is sent to the server for storage.
[0587] In this way, the system of the present invention makes it possible to prevent and quickly respond to theft in unmanned stores.
[0588] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0589] Step 1:
[0590] Acquiring video data
[0591] Subject: Server
[0592] The server acquires video data in real time from the surveillance cameras installed in the store. These surveillance cameras operate 24 hours a day and continuously transmit images from inside the store to the server. The input is live video from the surveillance cameras, and the output is raw data stored on the server.
[0593] Step 2:
[0594] Video data analysis
[0595] Subject: Server
[0596] The server inputs the acquired video data into the generative AI model and analyzes customer behavior. Specifically, the video data sent from the surveillance camera is sent to the generative AI model, which analyzes "what kind of behavior the customer is exhibiting." The prompt text used is "Please describe in text the behavior of the people in the video. In particular, please detect the behavior of someone picking up a product and heading toward the entrance / exit without going through the cash register." The input is the video data from the surveillance camera, and the output is the analyzed text data.
[0597] Step 3:
[0598] Detecting prohibited behavior
[0599] Subject: Server
[0600] The server analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behaviors. During this process, it checks whether the text data contains pre-set prohibited behavior keywords (e.g., heading to the entrance / exit without going through the cash register). The input is the analyzed text data, and the output is the results of the detection of prohibited behaviors. Specific operations involve the use of a keyword matching algorithm.
[0601] Step 4:
[0602] Warnings and Notifications
[0603] Subject: Server
[0604] If a prohibited behavior is detected, the server issues an alert and sends a notification to the store manager's communication device. This notification is sent using remote communication technologies (e.g., push notification, SMS). The input is the result of the detection of the prohibited behavior, and the output is a warning notification sent to the communication device. The device can then analyze the notification and take appropriate action.
[0605] Step 5:
[0606] Activating the face recognition camera and acquiring video
[0607] Subject: Server
[0608] When the server issues a warning, it activates a facial recognition camera installed near the entrance and captures the customer's facial image. This camera quickly captures the face of a specific person and sends the image data to the server. The input is the warning event, and the output is the facial image data from the facial recognition camera.
[0609] Step 6:
[0610] Data storage and access
[0611] Subject: Server
[0612] The server stores video data acquired from surveillance cameras and facial recognition cameras, as well as text data created by the generative AI model. Store managers can access this data later and provide it to external agencies such as the police if necessary. The input is video data and text data, and the output is the data stored in the storage. Specific operations use a database management system (e.g., SQL, NoSQL).
[0613] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0614] This invention is a security system designed to prevent theft and analyze customer emotions in unmanned stores. This system acquires video data from surveillance cameras installed in the store, analyzes the video data, and converts customer behavior into text data. If specific prohibited behavior is detected, the system then issues a warning and notifies the store manager's communication terminal. The system also has the function of simultaneously issuing a warning and activating a facial recognition camera installed near the entrance to capture video of the customer's face. Furthermore, by incorporating an emotion engine that recognizes user emotions and acquiring and analyzing customer emotion data, the system can improve the accuracy of detecting prohibited behavior.
[0615] Program processing
[0616] Video data acquisition and analysis
[0617] The server collects video data in real time from surveillance cameras installed in the store, then passes the video data to a generative AI model and emotion engine to analyze customer behavior and emotions and convert them into text and emotion data.
[0618] Examples:
[0619] The camera captures users entering the store and browsing the shelves.
[0620] The server inputs the video into a generative AI model and emotion engine, generating text data such as "The customer is browsing the shelves" and emotion data such as "The customer is interested."
[0621] NG word detection and notification
[0622] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether it contains certain prohibited behaviors, which triggers the next action.
[0623] Examples:
[0624] If the text data records that "the customer picked up the product and headed for the entrance without going through the cash register."
[0625] Suppose the emotional data contains information that "the customer is nervous."
[0626] The server determines that this behavior and emotion corresponds to prohibited behavior and sends a warning notification via LINE or phone stating, "The customer may have stolen a product."
[0627] Activating the face recognition camera and acquiring video
[0628] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[0629] Examples:
[0630] The server sends a warning notification and at the same time sends a command to activate the facial recognition camera.
[0631] The camera captures the customer's face and sends the image to a server.
[0632] The server stores the acquired facial images.
[0633] Information storage and access
[0634] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. These data can be accessed later by store managers and provided to external agencies such as the police.
[0635] Examples:
[0636] The server stores video and text data in dedicated storage when prohibited behavior is detected.
[0637] Emotion data generated by the emotion engine is also stored in the same way.
[0638] Store managers can later access the storage and download the data if necessary, and provide it to police.
[0639] This security system will enable the prevention of theft in unmanned stores and a rapid response, significantly improving the safety of store operations. In addition, by using an emotion engine, it will be possible to analyze not only customer behavior but also emotions, enabling more accurate detection of prohibited behavior.
[0640] The processing flow will be explained below.
[0641] Step 1:
[0642] The server collects video data in real time from surveillance cameras installed in the store, allowing it to constantly monitor the situation inside the store.
[0643] Specific behavior:
[0644] Surveillance cameras capture footage.
[0645] The server receives video data from the surveillance camera and stores it in a buffer.
[0646] Step 2:
[0647] The server passes the acquired video data to the generative AI model and emotion engine, which analyzes the customer's behavior and emotions in the video. Based on the analysis results, the video data is converted into text data and emotion data.
[0648] Specific behavior:
[0649] The server inputs the video data frame by frame into the generative AI model and emotion engine.
[0650] The generative AI model outputs the text data "A customer is browsing the shelves."
[0651] The emotion engine generates emotion data that indicates "customer interest."
[0652] The server stores the text data and emotion data.
[0653] Step 3:
[0654] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether it contains specific prohibited behaviors. If prohibited behaviors are detected, the server proceeds to the next step.
[0655] Specific behavior:
[0656] The server analyzes the text data and detects the behavior of "a customer picking up a product and heading to the entrance / exit without going through the cash register."
[0657] The server analyzes the emotional data and confirms that the customer is nervous.
[0658] The server determines whether the behavior is prohibited based on the behavior and emotions.
[0659] Step 4:
[0660] If a prohibited activity is detected, the server issues an alert and sends a notification to the store manager's communication terminal, which includes details of the prohibited activity.
[0661] Specific behavior:
[0662] The server generates a warning message based on the result of the detection of the prohibited behavior.
[0663] Using the LINE API, a message is sent to the store manager's LINE account stating, "A customer may have stolen an item."
[0664] In the case of telephone notifications, an automated voice message will warn, "A customer may have stolen your product."
[0665] Step 5:
[0666] At the same time as issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[0667] Specific behavior:
[0668] The server sends a start command to the face recognition camera.
[0669] A facial recognition camera captures the customer's face and sends the video data to a server.
[0670] The server stores the acquired facial images.
[0671] Step 6:
[0672] The server stores all captured video data, generated text data, and log data when prohibited behavior is detected, and these data can be accessed later by store managers.
[0673] Specific behavior:
[0674] The server stores video and text data in dedicated storage when prohibited behavior is detected.
[0675] Emotion data generated by the emotion engine is also stored in the same way.
[0676] The saved data is stored in a format that can be accessed by the store manager.
[0677] Step 7:
[0678] The user (store manager) receives the notification from the server and reports it to the police if necessary. The user accesses the data on the server and provides the necessary video and text data to the police.
[0679] Specific behavior:
[0680] The store manager will check the warning notification received via LINE or phone.
[0681] The store manager accesses the data stored on the server and downloads the video and text data of the theft.
[0682] If necessary, provide this data to the police.
[0683] The above is a specific processing flow of the system of the present invention, which significantly improves the security function of unmanned stores, suppresses theft, and realizes highly accurate analysis based on customer emotions.
[0684] Example 2
[0685] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0686] To effectively prevent theft in unmanned stores, it is necessary not only to analyze customer behavior but also to accurately grasp their emotions and respond promptly and appropriately. Conventional security systems focus on behavioral analysis and do not consider emotion analysis, resulting in false positives and oversights. Thus, there is a need for methods to improve the accuracy and safety of theft prevention in unmanned stores.
[0687] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0688] In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning when specific prohibited behavior is detected based on the text data and emotion data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance at the same time as issuing the warning and acquiring an image of the customer's face, and means incorporating an emotion recognition engine for analyzing the customer's behavior and emotion and converting it into text data and emotion data. This makes it possible to prevent theft in unmanned stores and to perform highly accurate behavioral analysis based on customer emotions.
[0689] A "surveillance camera" is a device that is installed in a store and continuously captures the situation inside the store as video data.
[0690] "Video Data" means a collection of video frames captured by surveillance cameras or facial recognition cameras, and is digital data containing customer behavior and other significant events.
[0691] "Analysis" is the process of analyzing the captured video data using specific algorithms and AI models to extract customer behavior and emotions.
[0692] "Customer" refers to a shopper or visitor to an unmanned store.
[0693] "Text data" refers to data that converts customer behavior and status obtained as analysis results into text information.
[0694] "Emotion data" is data that indicates the psychological state of a customer, obtained from their facial expressions and actions.
[0695] "Prohibited behavior" refers to stealing merchandise or other behavior that violates store rules.
[0696] "Warning" refers to an alert or notification issued when prohibited behavior is detected.
[0697] A "communication terminal" is a digital device such as a smartphone, tablet, or PC owned by a store manager.
[0698] A "facial recognition camera" is a camera installed to identify customers' faces and acquire video data.
[0699] An "emotion recognition engine" is software or an algorithm for analyzing customer emotions from video data.
[0700] "Storage" refers to cloud-based or physical data storage devices for storing captured data (video data, text data, emotion data, etc.).
[0701] "Store Manager" means a person or organization responsible for the operation and management of an unmanned store.
[0702] The present invention is a security system for preventing theft and analyzing customer sentiment in unmanned stores. The system operates using multiple hardware and software components.
[0703] Hardware Configuration
[0704] This system includes the following main hardware components:
[0705] Surveillance cameras: Installed in stores, they capture customer behavior in real time.
[0706] Facial recognition cameras: Installed near the store entrance, they capture images of customers' faces under certain conditions.
[0707] Communication device: Smartphone, tablet, PC, etc. owned by the store manager.
[0708] Software Configuration
[0709] The operation of the system is supported by the following software:
[0710] Generative AI models: For example, using GPT-4 for natural language processing to convert video data into text data.
[0711] Emotion recognition engine: For example, using Microsoft Azure's emotion analysis API to analyze customer emotions from video data.
[0712] Cloud storage: For example, Google Cloud Storage is used to store the acquired data.
[0713] Processing flow
[0714] The operation of the system proceeds as follows.
[0715] Acquisition of video data: The server acquires video data from the surveillance cameras in real time. The video data is sent to the server using a streaming protocol.
[0716] Video data analysis: The acquired video data is input into a generative AI model and an emotion recognition engine to analyze the customer's behavior and emotions. The analysis results are stored on the server as text data and emotion data.
[0717] Detection of prohibited behavior: Based on the generated text data and emotion data, the server detects whether certain prohibited behaviors are included. If prohibited behaviors are detected, a warning is issued.
[0718] Sending warning notifications: Warnings are sent in real time to the store manager's communication device via the LINE API or Twilio API.
[0719] Activating the facial recognition camera and capturing video: At the same time as receiving the warning notification, the server activates the facial recognition camera and captures the customer's facial video. The captured video data is stored in cloud storage as evidence.
[0720] Information storage: All captured data is stored in a dedicated cloud storage. Store managers can access this data and download it to provide to the police if necessary.
[0721] Specific examples
[0722] For example, imagine a situation where a user enters a store and looks at the shelves. In this case, a surveillance camera captures the user's behavior, and the server inputs the video data into a generative AI model and emotion recognition engine. The analysis results in text data such as "The customer is browsing the shelves" and emotion data such as "The customer is interested." If prohibited behavior is detected, an alert is immediately issued and the facial recognition camera is activated.
[0723] Example prompt sentence:
[0724] "Analyze real-time video footage from security cameras to detect customer behavior and emotions. Then, write a program to issue a warning notification, activate a facial recognition camera, and save the footage if prohibited behavior is detected."
[0725] This invention not only prevents theft in unmanned stores and enables highly accurate analysis of customer behavior, but also takes emotional data into account, enabling more accurate detection of prohibited behavior and quicker responses.
[0726] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0727] Step 1:
[0728] Acquiring video data
[0729] The server receives real-time video data from the surveillance cameras installed in the store. The input is the live video stream from the surveillance cameras. This video data is sent to the server using the RTSP protocol and buffered in memory. The output is a series of video frames.
[0730] What happens: The server connects to the video stream using the camera's RTSP URL and captures video frames, each ready to be sent for analysis.
[0731] Step 2:
[0732] Video data analysis
[0733] The server passes the acquired video data to a generative AI model and an emotion recognition engine to analyze customer behavior and emotions. The input is the video frames acquired in step 1. Data processing is performed by inputting the video frames into a generative AI model (e.g., OpenAI's GPT-4) and an emotion recognition engine (e.g., Microsoft Azure's Sentiment Analysis API). The output is text data and emotion data.
[0734] How it works: The server uses a Python script to input each video frame into a generative AI model, obtaining text data such as "A customer is browsing the shelves." At the same time, it sends the frame to an emotion recognition engine to obtain emotion data such as "The customer is interested."
[0735] Step 3:
[0736] Detecting prohibited behavior
[0737] The server analyzes the generated text data and emotion data to detect whether specific prohibited behaviors are included. The inputs are the text data and emotion data generated in step 2. Data calculations are performed by analyzing this data based on specific algorithms and regular expression patterns. The output is a determination result indicating whether prohibited behaviors have been detected.
[0738] Specific operation: The server performs regular expression pattern matching to detect phrases such as "The customer picked up the product and headed for the entrance without going through the cash register" from the text data. If the emotion data contains information such as "The customer is nervous," it records this as a prohibited behavior.
[0739] Step 4:
[0740] Sending warning notifications
[0741] If the server detects a prohibited behavior, it immediately sends a warning to the store manager's communication terminal. The input is the prohibited behavior judgment result obtained in step 3. Data processing involves generating and sending a warning message. The output is a warning notification sent to the store manager's communication terminal.
[0742] Specific operation: The server uses the LINE API to generate a message such as "A customer may have stolen an item" and sends it to the store manager's LINE account. At the same time, it also sends a phone notification using the Twilio API.
[0743] Step 5:
[0744] Activating the face recognition camera and acquiring video
[0745] Immediately after issuing the warning, the server automatically activates the facial recognition camera installed near the entrance and captures an image of the customer's face. The input is the warning notification from step 4. Data processing involves capturing and saving the facial image. The output is the captured facial image data.
[0746] Specific operation: The server sends the specified API request to the facial recognition camera to activate the camera. The camera captures the customer's face and sends the video to the server in real time. The captured facial video is encoded in the specified format (e.g., MPEG-4) and saved in the server's storage.
[0747] Step 6:
[0748] Retention of Information
[0749] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. The input is all the data generated at each step. Data calculation aggregates this data in one place and stores it in cloud storage. The output is all the stored data.
[0750] Specific operation: When prohibited behavior is detected, the server uploads video data, text data, and emotion data to dedicated cloud storage (e.g., Google Cloud Storage). The data is categorized with a timestamp and can be accessed later by store managers.
[0751] (Application example 2)
[0752] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0753] The challenge is to improve the safety and efficiency of store operations by preventing theft in unmanned stores and analyzing customer sentiment.In particular, since unmanned stores do not have permanent staff on-site, immediate response is difficult, and conventional surveillance systems have limitations in the accuracy of detecting prohibited behavior and the speed of response.
[0754] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning when specific prohibited behavior is detected based on the text data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance simultaneously with the issuance of the warning and acquiring a facial image of the customer, and means for analyzing customer emotion data using an emotion analysis engine to improve the accuracy of detecting prohibited behavior. This makes it possible to prevent theft in unmanned stores and detect prohibited behavior with high accuracy.
[0755] A "surveillance camera" is a device installed in a store to capture video data.
[0756] "Video data" refers to data that records the state of customers and the inside of a store captured by surveillance cameras and facial recognition cameras.
[0757] A "server" is a device or system that analyzes acquired video data and generates behavioral data and emotional data.
[0758] "Text data" is the analysis of video data and the representation of customer behavior as text information.
[0759] "Specific prohibited activities" refer to activities that should not be carried out in the store, including, for example, theft.
[0760] An "alert" is a notification that is sent when certain prohibited behavior is detected.
[0761] A "communication terminal" refers to a device owned by a store manager that is capable of receiving notifications such as warnings.
[0762] A "facial recognition camera" is a camera used to capture and recognize facial images of customers.
[0763] "Emotion data" is data generated by analyzing customer emotions using an emotion analysis engine.
[0764] An "emotion analysis engine" is software or algorithms for analyzing a customer's emotional state from acquired video data.
[0765] "Storage" is a storage device that stores data for a long period of time and allows it to be accessed later.
[0766] In this invention, a system is constructed by integrating multiple pieces of hardware and software to improve the monitoring and security of unmanned stores.
[0767] Hardware and software used
[0768] 1. Hardware
[0769] Surveillance camera: Used to capture video data within the store.
[0770] Facial recognition camera: Used to capture facial images of customers when an alert is issued.
[0771] Communication terminal: A device that allows store managers to receive alert notifications.
[0772] Server: A central device that analyzes data and manages the generated behavioral and emotional data.
[0773] 2. Software
[0774] Generative AI models (e.g., OpenAI's GPT series): Analyze captured video data and convert customer behavior into text data.
[0775] Emotion engine (e.g. Microsoft's Emotion API): Analyzes customer emotion data.
[0776] Notification system (e.g., LINE API, Twilio): Sends warnings to the store manager's communication device based on the prohibited behavior detected.
[0777] System Operation Overview
[0778] Video data acquisition and analysis
[0779] Surveillance cameras capture real-time video data of customer behavior in the store, and the server passes this video data to a generative AI model to generate behavioral data.
[0780] Examples:
[0781] Customers are seen on surveillance cameras picking up items from the shelves.
[0782] The server inputs the video data into a generative AI model and generates behavioral data such as "a customer picking up a product from a shelf."
[0783] The server then uses an emotion engine to generate "emotional data" of the customer from the video data, and obtains the result that "the customer is interested."
[0784] Example prompt for a generative AI model:
[0785] Analyze the video data below and convert customer behavior into text data.
[0786] Video footage: A customer picks up an item from the shelf and heads straight for the exit.
[0787] Output of the generative AI model: The customer picks up the item and heads to the exit without going through the checkout.
[0788] Detecting and notifying prohibited behavior
[0789] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether prohibited behavior is included. If prohibited behavior is detected, the server sends a warning to the store manager's communication device via the notification system.
[0790] Examples:
[0791] Behavioral data records that "the customer picked up the product and headed for the exit without going through the cash register."
[0792] The emotional data includes information that "the customer is nervous."
[0793] The server uses this information to detect prohibited behavior and sends a warning notification saying, "A customer may have stolen a product."
[0794] Activating the face recognition camera and acquiring video
[0795] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture the customer's facial image, which is then saved by the server.
[0796] Examples:
[0797] The server will send a warning notification and simultaneously activate the facial recognition camera.
[0798] The camera captures the customer's face and sends the image to a server.
[0799] The server stores the acquired facial image as evidence.
[0800] Example of a prompt for the emotion engine:
[0801] Analyze the video data below and convert customer sentiment into text data.
[0802] Footage: Customers are seen picking up items and heading for the exit.
[0803] Emotion Engine Output: Customer is nervous.
[0804] Information storage and access
[0805] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. These data can be accessed by store managers and later downloaded and provided to external organizations as needed.
[0806] Examples:
[0807] The server stores the video data, text data, and emotional data when prohibited behavior is detected in dedicated storage.
[0808] Store managers can access the storage and download this data as needed, providing it to external agencies such as the police.
[0809] This system will enable the prevention and rapid response of theft in unmanned stores, significantly improving the safety of store operations. In addition, by using an emotion engine, it is possible to analyze not only customer behavior but also emotions, enabling more accurate detection of prohibited behavior.
[0810] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0811] Step 1:
[0812] The server acquires video data in real time from surveillance cameras installed in the store. Specifically, the cameras capture customer behavior in the store and send the video data to the server. The input data is the video data from the surveillance cameras, and the output data is the video data itself.
[0813] Step 2:
[0814] The server passes the acquired video data to a generative AI model and converts the customer's behavior into text data. Specifically, the server inputs the video data into a generative AI model (e.g., OpenAI's GPT series) and generates behavioral data such as "the customer picking up a product." The input data is video data, and the output data is behavioral data in text format.
[0815] Step 3:
[0816] The server then uses an emotion engine to analyze the customer's emotions from the video data. Specifically, the server inputs the video data into an emotion engine (e.g., Microsoft's Emotion API) and generates emotion data indicating that the customer is interested. The input data is video data, and the output data is emotion data in text format.
[0817] Step 4:
[0818] The server analyzes the generated behavioral and emotional data to detect whether it contains specific prohibited behavior. Specifically, it uses an analysis algorithm to identify prohibited behavior, such as "a customer picks up a product and heads toward the exit without going through the cash register," and also analyzes emotional data, such as "the customer is nervous." The input data are behavioral and emotional data, and the output data are the results of the detection of prohibited behavior.
[0819] Step 5:
[0820] If prohibited behavior is detected, the server sends a warning to the store manager's communication device via the notification system. Specifically, the server uses the LINE API or Twilio to send a warning message to the store manager's smartphone, such as "A customer may have stolen an item." The input data is the detection result, and the output data is the warning notification.
[0821] Step 6:
[0822] Immediately after issuing the warning, the server automatically activates the facial recognition camera installed near the entrance. Specifically, the server sends a start command to the facial recognition camera to start capturing facial images. The input data is the start command, and the output data is facial image data.
[0823] Step 7:
[0824] The facial recognition camera captures the customer's facial image and sends it to a server, which then stores the data. Specifically, the server receives the facial image data and stores it in dedicated storage. This storage can be accessed later. The input data is the customer's facial image, and the output data is the data stored in the storage.
[0825] Step 8:
[0826] The server saves the acquired video data, the generated behavioral and emotional data, and the notification log in storage. Specifically, the server stores each data in a dedicated storage device so that it can be preserved for a long period of time. The input data is the video data, behavioral data, emotional data, and notification log, and the output data is the data saved in storage.
[0827] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0828] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0829] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0830] [Third embodiment]
[0831] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0832] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0833] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0834] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0835] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0836] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0837] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0838] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0839] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0840] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0841] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0842] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0843] The present invention is a security system designed to prevent theft in unmanned stores. This system acquires video data from surveillance cameras installed in the store, analyzes the video data, and converts customer behavior into text data. If a specific prohibited behavior is detected, the system issues a warning and notifies the store manager's communication terminal. The system also has the function of activating a face recognition camera installed near the entrance and capturing a video of the customer's face at the same time as issuing the warning.
[0844] Program processing
[0845] Video data acquisition and analysis
[0846] The server collects video data in real time from surveillance cameras installed in the store, then passes the collected video data to a generative AI model, which analyzes customer behavior and converts it into text data.
[0847] Examples:
[0848] The camera captures users entering the store and browsing the shelves.
[0849] The server inputs the video into a generative AI model and outputs the text data, "Customer browsing the shelves."
[0850] NG word detection and notification
[0851] The server analyzes the text data created by the generative AI model and detects whether it contains specific prohibited behavior (e.g., "no money was inserted"). If prohibited behavior is detected, the server issues an alert and sends a notification to the store manager's communication device (e.g., LINE or phone).
[0852] Examples:
[0853] If the text data records that "the customer picked up the product and headed for the entrance without going through the cash register."
[0854] The server determines that this behavior is prohibited and sends a warning notification via LINE or phone stating, "The customer may have stolen a product."
[0855] Activating the face recognition camera and acquiring video
[0856] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[0857] Examples:
[0858] The server sends a warning notification and at the same time sends a command to activate the facial recognition camera.
[0859] The camera captures the customer's face and sends the image to a server.
[0860] The server stores the acquired facial images.
[0861] Information storage and access
[0862] The server stores all captured video data, generated text data, and log data when prohibited behavior is detected. This data can be accessed later by store managers and provided to external agencies such as the police.
[0863] Examples:
[0864] The server stores the video data and text data when prohibited behavior is detected in dedicated storage.
[0865] Store managers can later access the storage and download the data if necessary, and provide it to police.
[0866] This security system will enable the prevention of theft in unmanned stores and a rapid response, significantly improving the safety of store operations.
[0867] The processing flow will be explained below.
[0868] Step 1:
[0869] The server receives video data in real time from surveillance cameras installed in the store, allowing the situation inside the store to be constantly monitored.
[0870] Specific behavior:
[0871] Surveillance cameras capture footage.
[0872] The server receives video data from the surveillance camera and stores it in a buffer.
[0873] Step 2:
[0874] The server passes the captured video data to a generative AI model, which analyzes customer behavior in the video. This analysis converts the video into text data.
[0875] Specific behavior:
[0876] The server inputs the video data frame by frame into the generative AI model.
[0877] The generative AI model analyzes the frames and outputs text data such as "A customer is browsing the shelves."
[0878] The server stores the text data.
[0879] Step 3:
[0880] The server analyzes the text data generated by the generative AI model to detect whether it contains certain prohibited behaviors, which triggers the next action.
[0881] Specific behavior:
[0882] The server analyzes the text data sentence by sentence.
[0883] Based on the analysis results, it is checked whether prohibited actions (e.g., "did not put money in") are included.
[0884] If a prohibited action is detected, the information is recorded in an internal flag.
[0885] Step 4:
[0886] If a prohibited activity is detected, the server issues an alert and sends a notification to the store manager's communication terminal, which includes details of the prohibited activity.
[0887] Specific behavior:
[0888] The server generates a warning message based on an internal flag.
[0889] Using the LINE API, a message is sent to the store manager's LINE account stating, "A customer may have stolen an item."
[0890] In the case of telephone notification, an automated voice warning message will be sent.
[0891] Step 5:
[0892] The server automatically activates a facial recognition camera installed near the entrance to capture images of the customer's face, which can then be used as evidence.
[0893] Specific behavior:
[0894] The server sends a start command to the face recognition camera.
[0895] A facial recognition camera captures the customer's face and sends the video data to a server.
[0896] The server stores the acquired facial images.
[0897] Step 6:
[0898] The server stores the captured video data, text data, and log data when prohibited behavior is detected, which will be used for later analysis and reporting to the police.
[0899] Specific behavior:
[0900] The server stores the video data and text data when prohibited behavior is detected in dedicated storage.
[0901] Video data obtained from facial recognition cameras will also be stored in the same way.
[0902] The stored data is accessible to store managers.
[0903] Step 7:
[0904] The user (store manager) receives the notification from the server and reports it to the police if necessary. The user accesses the data on the server and provides the necessary video and text data to the police.
[0905] Specific behavior:
[0906] The store manager will check the warning notification received via LINE or phone.
[0907] The store manager accesses the data stored on the server and downloads the video and text data of the theft.
[0908] If necessary, provide this data to the police.
[0909] The above is a specific processing flow of the system of the present invention.
[0910] Example 1
[0911] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0912] In recent years, the risk of theft has increased with the increase in unmanned stores. However, current surveillance systems have difficulty detecting theft in real time and responding quickly. Furthermore, conventional systems lack a means to reliably and quickly notify store managers, making it difficult to prevent theft. This poses a risk to the safety of store operations.
[0913] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0914] In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning if specific prohibited behavior is detected based on the text data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance simultaneously with the issuance of the warning to acquire a face image of the customer, means for inputting the acquired video data and text data into a generative AI model, and means for detecting whether specific prohibited behavior is included based on the analysis results of the generative AI model. This enables real-time detection of theft in unmanned stores and rapid response, significantly improving the safety of store operations.
[0915] A "surveillance camera" is a device that is installed for the purpose of monitoring a specific location and acquires video data.
[0916] "Video data" refers to image information and video information captured by a surveillance camera.
[0917] A "generative AI model" is a model that uses artificial intelligence to analyze input data and generate text or other information as output.
[0918] "Text data" refers to character information obtained by analyzing video data.
[0919] "Prohibited behavior" refers to specific behaviors that are performed within a store and that violate predetermined rules.
[0920] A "warning" is a notification issued when a prohibited action is detected.
[0921] A "communication terminal" is a device for sending and receiving information, and includes smartphones, tablets, etc.
[0922] A "face recognition camera" is a device that detects an individual's face and acquires its image data.
[0923] "Storage" refers to a storage device or storage service for storing data.
[0924] "Log data" is data that records historical information about events and operations that occur within the system.
[0925] This invention is a security system for preventing theft in unmanned stores, and operates in the following procedure, centered around a server.
[0926] First, the server acquires video data in real time from the surveillance cameras installed in the store. These cameras are often standard IP cameras, such as Hikvision cameras. The server then stores this video data in a specific buffer.
[0927] Next, the server inputs the captured video data frame by frame into a generative AI model, such as OpenAI's GPT-4, which has image analysis capabilities. The server sends the following prompt to the generative AI model:
[0928] "Analyze the image below and explain the customer behavior. Image: [Image data]"
[0929] The generative AI model then analyzes the video data and outputs the customer's behavior as text data, such as "The customer is browsing the shelves."
[0930] The server then analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behavior. Prohibited behavior includes "taking away merchandise without going through the cash register." If prohibited behavior is detected, the server issues a warning and sends a notification to the store manager's communication device. Notifications can be sent using the LINE API or SMS API, and a specific example would be a warning message stating, "A customer may have stolen an item."
[0931] Furthermore, at the same time as issuing the warning, the server activates a facial recognition camera installed near the entrance. This facial recognition camera is equipped with facial recognition software such as Face++ or Amazon Rekognition. The server then sends a capture command to the camera, which captures and stores the customer's facial image.
[0932] Finally, the server stores all captured video data, generated text data, and log data when prohibited behavior is detected in a secure storage location, using Amazon S3 or Google Cloud Storage. Store managers can access this storage and download the data as needed, providing it to external agencies (such as the police).
[0933] In this way, it becomes possible to detect theft in unmanned stores in real time and respond quickly, greatly improving the safety of store operations.
[0934] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0935] Step 1: Acquire video data
[0936] The server acquires video data from surveillance cameras installed in the store. The server periodically accesses the surveillance cameras to acquire new video data. General IP cameras are used for this purpose. The input is the video data acquired from the surveillance cameras, and the output is the video data stored in the temporary storage buffer within the server.
[0937] Example: A server captures video streams from a Hikvision IP camera once per second and stores them in a buffer for analysis.
[0938] Step 2: Analyzing the video data
[0939] The server passes the acquired video data to a generative AI model and converts customer behavior into text data. The server then divides the video data into frames and inputs each frame image into the generative AI model. OpenAI's GPT-4 and other models are used as generative AI models. The input is each frame of video data, and the output is text data.
[0940] Example: The server sends an "image recognition prompt" to OpenAI's GPT-4 and saves the returned text data in an analysis folder.
[0941] Example prompt: "Analyze the image below and describe the customer's behavior. Image: [image data]"
[0942] Step 3: Analyzing the text data
[0943] The server analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behavior. The input is the text data output by the generative AI model, and the output is a flag indicating whether or not prohibited behavior exists. Analysis is performed using natural language processing technology based on the definition of prohibited behavior.
[0944] Example: The server detects text that describes prohibited actions, such as "taking merchandise without going through the cash register."
[0945] Step 4: Issuing warnings and notifications
[0946] If the server detects a prohibited behavior, it issues a warning and sends a notification to the store manager's communication device. The server sends the notification using a pre-configured communication method (e.g., LINE API or SMS API). The input is a flag for the prohibited behavior, and the output is the warning notification sent.
[0947] Example: The server uses the LINE API to send a notification to the store manager saying, "A customer may have stolen an item."
[0948] Step 5: Turn on the face recognition camera
[0949] Immediately after issuing the warning, the server activates a facial recognition camera installed near the entrance and captures the customer's facial image. The server accesses the specified facial recognition camera and sends a capture command. The input is a warning issuance flag, and the output is the captured facial image data.
[0950] Example: The server uses the Face++ API to send a "capture start command" and saves the captured facial image on the server.
[0951] Step 6: Store and access your data
[0952] The server securely stores the captured video data, generated text data, and log data when prohibited behavior is detected. This storage uses Amazon S3 or Google Cloud Storage. The input is various types of data, and the output is data stored in the cloud storage.
[0953] Example: The server uploads video data, text data, and warning logs related to prohibited behavior to Amazon S3 and generates a link for later access.
[0954] The terminal (store manager) accesses this data and provides it to external agencies (e.g., police) if necessary.
[0955] Example: A store manager opens the management app, clicks a URL link, and downloads the necessary video and log data.
[0956] (Application example 1)
[0957] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0958] Preventing theft in unmanned stores and responding quickly and reliably are key challenges. Currently, monitoring in unmanned stores is inefficient, which can lead to delayed detection and countermeasures when theft occurs. Another problem is that information on detected theft is not properly recorded and managed, making it impossible to use as evidence later.
[0959] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0960] In this invention, the server includes: means for acquiring video data from surveillance cameras installed in the store; means for analyzing the acquired video data and converting customer behavior into text data; means for issuing a warning if specific prohibited behavior is detected based on the text data; means for using remote communication technology to notify the store manager's communication terminal of the warning; means for using the server to transmit video data in real time and analyze the text data; means for using a generative AI model to analyze the text data; and means for activating a facial recognition camera installed near the entrance and capturing facial images of customers simultaneously with the issuance of the warning. This enables the immediate detection of theft in an unmanned store and the notification of a warning to the store manager. Furthermore, capturing and storing facial images of customers at the same time can be used as evidence at a later date.
[0961] A "surveillance camera" is a device installed in a store to acquire video data.
[0962] "Video data" refers to video information acquired by a surveillance camera.
[0963] "Analysis" is the process of analyzing customer behavior based on the acquired video data.
[0964] "Text data" is character data generated as a result of analyzing video data.
[0965] "Prohibited Behavior" refers to any behavior that is not permitted within the store.
[0966] A "warning" is a notification issued when prohibited behavior is detected.
[0967] A "communication terminal" is an information and communication device such as a smartphone or computer used by a store manager.
[0968] A "facial recognition camera" is a camera that recognizes the facial characteristics of a specific person and acquires video data.
[0969] A "generative AI model" is a model for generating and analyzing data using artificial intelligence.
[0970] "Telecommunications technology" refers to technology for sending and receiving information to and from communication terminals.
[0971] A "server" is a computer system that processes and stores data.
[0972] "Storage" refers to a storage device for saving data.
[0973] "Real-time" refers to processing that is immediate and without delay.
[0974] The system for implementing this invention consists of surveillance cameras installed in a store, a generative AI model for analyzing video data, a server for detecting prohibited behavior and sending warnings, a communication terminal for the store manager, a facial recognition camera, and storage for data storage.
[0975] Specific system configuration and operation
[0976] 1. Acquiring video data
[0977] The server acquires video data in real time from surveillance cameras installed in the store. The surveillance cameras capture the situation inside the store 24 hours a day and send the video to the server.
[0978] 2. Analysis of video data
[0979] The server inputs the acquired video data into a generative AI model to analyze customer behavior. The generative AI model converts the customer behavior into text data based on the video data. This generative AI model incorporates pre-learned behavioral patterns, allowing it to analyze specific behavior in detail.
[0980] An example of a prompt sentence is, "Please describe in text the actions of the person in the video. In particular, please detect the action of picking up a product and heading towards the entrance / exit without going through the cash register."
[0981] 3. Detecting and warning against prohibited behavior
[0982] The server analyzes the text data created by the generative AI model and detects whether it contains certain prohibited behaviors. For example, if the server detects behavior such as "picking up a product and heading to the entrance without going through the cash register," it issues an alert and immediately sends a notification to the store manager's communication terminal. This notification is sent using remote communication technology.
[0983] 4. Activating the face recognition camera and acquiring video
[0984] When the warning is issued, the server activates a facial recognition camera installed near the entrance to capture the facial image of the customer. The facial recognition camera quickly captures the face of a specific person and sends the image data to the server.
[0985] 5. Data storage and management
[0986] The server stores video data acquired from surveillance cameras and facial recognition cameras, as well as text data created by generative AI models, allowing store managers to access the storage later and download the data as needed to provide it to external agencies such as the police.
[0987] Specific examples
[0988] For example, if a surveillance camera captures a customer picking up an item and then heading for the entrance without going through the cash register, the video data is sent to a server. The server uses a generative AI model to generate text data from the video, such as "The customer picks up an item and heads for the entrance without going through the cash register." Based on this text data, the server detects prohibited behavior and issues a warning to the store manager. At the same time, a facial recognition camera captures the customer's face, and the video data is sent to the server for storage.
[0989] In this way, the system of the present invention makes it possible to prevent and quickly respond to theft in unmanned stores.
[0990] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0991] Step 1:
[0992] Acquiring video data
[0993] Subject: Server
[0994] The server acquires video data in real time from the surveillance cameras installed in the store. These surveillance cameras operate 24 hours a day and continuously transmit images from inside the store to the server. The input is live video from the surveillance cameras, and the output is raw data stored on the server.
[0995] Step 2:
[0996] Video data analysis
[0997] Subject: Server
[0998] The server inputs the acquired video data into the generative AI model and analyzes customer behavior. Specifically, the video data sent from the surveillance camera is sent to the generative AI model, which analyzes "what kind of behavior the customer is exhibiting." The prompt text used is "Please describe in text the behavior of the people in the video. In particular, please detect the behavior of someone picking up a product and heading toward the entrance / exit without going through the cash register." The input is the video data from the surveillance camera, and the output is the analyzed text data.
[0999] Step 3:
[1000] Detecting prohibited behavior
[1001] Subject: Server
[1002] The server analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behaviors. During this process, it checks whether the text data contains pre-set prohibited behavior keywords (e.g., heading to the entrance / exit without going through the cash register). The input is the analyzed text data, and the output is the results of the detection of prohibited behaviors. Specific operations involve the use of a keyword matching algorithm.
[1003] Step 4:
[1004] Warnings and Notifications
[1005] Subject: Server
[1006] If a prohibited behavior is detected, the server issues an alert and sends a notification to the store manager's communication device. This notification is sent using remote communication technologies (e.g., push notification, SMS). The input is the result of the detection of the prohibited behavior, and the output is a warning notification sent to the communication device. The device can then analyze the notification and take appropriate action.
[1007] Step 5:
[1008] Activating the face recognition camera and acquiring video
[1009] Subject: Server
[1010] When the server issues a warning, it activates a facial recognition camera installed near the entrance and captures the customer's facial image. This camera quickly captures the face of a specific person and sends the image data to the server. The input is the warning event, and the output is the facial image data from the facial recognition camera.
[1011] Step 6:
[1012] Data storage and access
[1013] Subject: Server
[1014] The server stores video data acquired from surveillance cameras and facial recognition cameras, as well as text data created by the generative AI model. Store managers can access this data later and provide it to external agencies such as the police if necessary. The input is video data and text data, and the output is the data stored in the storage. Specific operations use a database management system (e.g., SQL, NoSQL).
[1015] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1016] This invention is a security system designed to prevent theft and analyze customer emotions in unmanned stores. This system acquires video data from surveillance cameras installed in the store, analyzes the video data, and converts customer behavior into text data. If specific prohibited behavior is detected, the system then issues a warning and notifies the store manager's communication terminal. The system also has the function of simultaneously issuing a warning and activating a facial recognition camera installed near the entrance to capture video of the customer's face. Furthermore, by incorporating an emotion engine that recognizes user emotions and acquiring and analyzing customer emotion data, the system can improve the accuracy of detecting prohibited behavior.
[1017] Program processing
[1018] Video data acquisition and analysis
[1019] The server collects video data in real time from surveillance cameras installed in the store, then passes the video data to a generative AI model and emotion engine to analyze customer behavior and emotions and convert them into text and emotion data.
[1020] Examples:
[1021] The camera captures users entering the store and browsing the shelves.
[1022] The server inputs the video into a generative AI model and emotion engine, generating text data such as "The customer is browsing the shelves" and emotion data such as "The customer is interested."
[1023] NG word detection and notification
[1024] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether it contains certain prohibited behaviors, which triggers the next action.
[1025] Examples:
[1026] If the text data records that "the customer picked up the product and headed for the entrance without going through the cash register."
[1027] Suppose the emotional data contains information that "the customer is nervous."
[1028] The server determines that this behavior and emotion corresponds to prohibited behavior and sends a warning notification via LINE or phone stating, "The customer may have stolen a product."
[1029] Activating the face recognition camera and acquiring video
[1030] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[1031] Examples:
[1032] The server sends a warning notification and at the same time sends a command to activate the facial recognition camera.
[1033] The camera captures the customer's face and sends the image to a server.
[1034] The server stores the acquired facial images.
[1035] Information storage and access
[1036] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. These data can be accessed later by store managers and provided to external agencies such as the police.
[1037] Examples:
[1038] The server stores video and text data in dedicated storage when prohibited behavior is detected.
[1039] Emotion data generated by the emotion engine is also stored in the same way.
[1040] Store managers can later access the storage and download the data if necessary, and provide it to police.
[1041] This security system will enable the prevention of theft in unmanned stores and a rapid response, significantly improving the safety of store operations. In addition, by using an emotion engine, it will be possible to analyze not only customer behavior but also emotions, enabling more accurate detection of prohibited behavior.
[1042] The processing flow will be explained below.
[1043] Step 1:
[1044] The server collects video data in real time from surveillance cameras installed in the store, allowing it to constantly monitor the situation inside the store.
[1045] Specific behavior:
[1046] Surveillance cameras capture footage.
[1047] The server receives video data from the surveillance camera and stores it in a buffer.
[1048] Step 2:
[1049] The server passes the acquired video data to the generative AI model and emotion engine, which analyzes the customer's behavior and emotions in the video. Based on the analysis results, the video data is converted into text data and emotion data.
[1050] Specific behavior:
[1051] The server inputs the video data frame by frame into the generative AI model and emotion engine.
[1052] The generative AI model outputs the text data "A customer is browsing the shelves."
[1053] The emotion engine generates emotion data that indicates "customer interest."
[1054] The server stores the text data and emotion data.
[1055] Step 3:
[1056] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether it contains specific prohibited behaviors. If prohibited behaviors are detected, the server proceeds to the next step.
[1057] Specific behavior:
[1058] The server analyzes the text data and detects the behavior of "a customer picking up a product and heading to the entrance / exit without going through the cash register."
[1059] The server analyzes the emotional data and confirms that the customer is nervous.
[1060] The server determines whether the behavior is prohibited based on the behavior and emotions.
[1061] Step 4:
[1062] If a prohibited activity is detected, the server issues an alert and sends a notification to the store manager's communication terminal, which includes details of the prohibited activity.
[1063] Specific behavior:
[1064] The server generates a warning message based on the result of the detection of the prohibited behavior.
[1065] Using the LINE API, a message is sent to the store manager's LINE account stating, "A customer may have stolen an item."
[1066] In the case of telephone notifications, an automated voice message will warn, "A customer may have stolen your product."
[1067] Step 5:
[1068] At the same time as issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[1069] Specific behavior:
[1070] The server sends a start command to the face recognition camera.
[1071] A facial recognition camera captures the customer's face and sends the video data to a server.
[1072] The server stores the acquired facial images.
[1073] Step 6:
[1074] The server stores all captured video data, generated text data, and log data when prohibited behavior is detected, and these data can be accessed later by store managers.
[1075] Specific behavior:
[1076] The server stores video and text data in dedicated storage when prohibited behavior is detected.
[1077] Emotion data generated by the emotion engine is also stored in the same way.
[1078] The saved data is stored in a format that can be accessed by the store manager.
[1079] Step 7:
[1080] The user (store manager) receives the notification from the server and reports it to the police if necessary. The user accesses the data on the server and provides the necessary video and text data to the police.
[1081] Specific behavior:
[1082] The store manager will check the warning notification received via LINE or phone.
[1083] The store manager accesses the data stored on the server and downloads the video and text data of the theft.
[1084] If necessary, provide this data to the police.
[1085] The above is a specific processing flow of the system of the present invention, which significantly improves the security function of unmanned stores, suppresses theft, and realizes highly accurate analysis based on customer emotions.
[1086] Example 2
[1087] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1088] To effectively prevent theft in unmanned stores, it is necessary not only to analyze customer behavior but also to accurately grasp their emotions and respond promptly and appropriately. Conventional security systems focus on behavioral analysis and do not consider emotion analysis, resulting in false positives and oversights. Thus, there is a need for methods to improve the accuracy and safety of theft prevention in unmanned stores.
[1089] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1090] In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning when specific prohibited behavior is detected based on the text data and emotion data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance at the same time as issuing the warning and acquiring an image of the customer's face, and means incorporating an emotion recognition engine for analyzing the customer's behavior and emotion and converting it into text data and emotion data. This makes it possible to prevent theft in unmanned stores and to perform highly accurate behavioral analysis based on customer emotions.
[1091] A "surveillance camera" is a device that is installed in a store and continuously captures the situation inside the store as video data.
[1092] "Video Data" means a collection of video frames captured by surveillance cameras or facial recognition cameras, and is digital data containing customer behavior and other significant events.
[1093] "Analysis" is the process of analyzing the captured video data using specific algorithms and AI models to extract customer behavior and emotions.
[1094] "Customer" refers to a shopper or visitor to an unmanned store.
[1095] "Text data" refers to data that converts customer behavior and status obtained as analysis results into text information.
[1096] "Emotion data" is data that indicates the psychological state of a customer, obtained from their facial expressions and actions.
[1097] "Prohibited behavior" refers to stealing merchandise or other behavior that violates store rules.
[1098] "Warning" refers to an alert or notification issued when prohibited behavior is detected.
[1099] A "communication terminal" is a digital device such as a smartphone, tablet, or PC owned by a store manager.
[1100] A "facial recognition camera" is a camera installed to identify customers' faces and acquire video data.
[1101] An "emotion recognition engine" is software or an algorithm for analyzing customer emotions from video data.
[1102] "Storage" refers to cloud-based or physical data storage devices for storing captured data (video data, text data, emotion data, etc.).
[1103] "Store Manager" means a person or organization responsible for the operation and management of an unmanned store.
[1104] The present invention is a security system for preventing theft and analyzing customer sentiment in unmanned stores. The system operates using multiple hardware and software components.
[1105] Hardware Configuration
[1106] This system includes the following main hardware components:
[1107] Surveillance cameras: Installed in stores, they capture customer behavior in real time.
[1108] Facial recognition cameras: Installed near the store entrance, they capture images of customers' faces under certain conditions.
[1109] Communication device: Smartphone, tablet, PC, etc. owned by the store manager.
[1110] Software Configuration
[1111] The operation of the system is supported by the following software:
[1112] Generative AI models: For example, using GPT-4 for natural language processing to convert video data into text data.
[1113] Emotion recognition engine: For example, using Microsoft Azure's emotion analysis API to analyze customer emotions from video data.
[1114] Cloud storage: For example, Google Cloud Storage is used to store the acquired data.
[1115] Processing flow
[1116] The operation of the system proceeds as follows.
[1117] Acquisition of video data: The server acquires video data from the surveillance cameras in real time. The video data is sent to the server using a streaming protocol.
[1118] Video data analysis: The acquired video data is input into a generative AI model and an emotion recognition engine to analyze the customer's behavior and emotions. The analysis results are stored on the server as text data and emotion data.
[1119] Detection of prohibited behavior: Based on the generated text data and emotion data, the server detects whether certain prohibited behaviors are included. If prohibited behaviors are detected, a warning is issued.
[1120] Sending warning notifications: Warnings are sent in real time to the store manager's communication device via the LINE API or Twilio API.
[1121] Activating the facial recognition camera and capturing video: At the same time as receiving the warning notification, the server activates the facial recognition camera and captures the customer's facial video. The captured video data is stored in cloud storage as evidence.
[1122] Information storage: All captured data is stored in a dedicated cloud storage. Store managers can access this data and download it to provide to the police if necessary.
[1123] Specific examples
[1124] For example, imagine a situation where a user enters a store and looks at the shelves. In this case, a surveillance camera captures the user's behavior, and the server inputs the video data into a generative AI model and emotion recognition engine. The analysis results in text data such as "The customer is browsing the shelves" and emotion data such as "The customer is interested." If prohibited behavior is detected, an alert is immediately issued and the facial recognition camera is activated.
[1125] Example prompt sentence:
[1126] "Analyze real-time video footage from security cameras to detect customer behavior and emotions. Then, write a program to issue a warning notification, activate a facial recognition camera, and save the footage if prohibited behavior is detected."
[1127] This invention not only prevents theft in unmanned stores and enables highly accurate analysis of customer behavior, but also takes emotional data into account, enabling more accurate detection of prohibited behavior and quicker responses.
[1128] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1129] Step 1:
[1130] Acquiring video data
[1131] The server receives real-time video data from the surveillance cameras installed in the store. The input is the live video stream from the surveillance cameras. This video data is sent to the server using the RTSP protocol and buffered in memory. The output is a series of video frames.
[1132] What happens: The server connects to the video stream using the camera's RTSP URL and captures video frames, each ready to be sent for analysis.
[1133] Step 2:
[1134] Video data analysis
[1135] The server passes the acquired video data to a generative AI model and an emotion recognition engine to analyze customer behavior and emotions. The input is the video frames acquired in step 1. Data processing is performed by inputting the video frames into a generative AI model (e.g., OpenAI's GPT-4) and an emotion recognition engine (e.g., Microsoft Azure's Sentiment Analysis API). The output is text data and emotion data.
[1136] How it works: The server uses a Python script to input each video frame into a generative AI model, obtaining text data such as "A customer is browsing the shelves." At the same time, it sends the frame to an emotion recognition engine to obtain emotion data such as "The customer is interested."
[1137] Step 3:
[1138] Detecting prohibited behavior
[1139] The server analyzes the generated text data and emotion data to detect whether specific prohibited behaviors are included. The inputs are the text data and emotion data generated in step 2. Data calculations are performed by analyzing this data based on specific algorithms and regular expression patterns. The output is a determination result indicating whether prohibited behaviors have been detected.
[1140] Specific operation: The server performs regular expression pattern matching to detect phrases such as "The customer picked up the product and headed for the entrance without going through the cash register" from the text data. If the emotion data contains information such as "The customer is nervous," it records this as a prohibited behavior.
[1141] Step 4:
[1142] Sending warning notifications
[1143] If the server detects a prohibited behavior, it immediately sends a warning to the store manager's communication terminal. The input is the prohibited behavior judgment result obtained in step 3. Data processing involves generating and sending a warning message. The output is a warning notification sent to the store manager's communication terminal.
[1144] Specific operation: The server uses the LINE API to generate a message such as "A customer may have stolen an item" and sends it to the store manager's LINE account. At the same time, it also sends a phone notification using the Twilio API.
[1145] Step 5:
[1146] Activating the face recognition camera and acquiring video
[1147] Immediately after issuing the warning, the server automatically activates the facial recognition camera installed near the entrance and captures an image of the customer's face. The input is the warning notification from step 4. Data processing involves capturing and saving the facial image. The output is the captured facial image data.
[1148] Specific operation: The server sends the specified API request to the facial recognition camera to activate the camera. The camera captures the customer's face and sends the video to the server in real time. The captured facial video is encoded in the specified format (e.g., MPEG-4) and saved in the server's storage.
[1149] Step 6:
[1150] Retention of Information
[1151] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. The input is all the data generated at each step. Data calculation aggregates this data in one place and stores it in cloud storage. The output is all the stored data.
[1152] Specific operation: When prohibited behavior is detected, the server uploads video data, text data, and emotion data to dedicated cloud storage (e.g., Google Cloud Storage). The data is categorized with a timestamp and can be accessed later by store managers.
[1153] (Application example 2)
[1154] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1155] The challenge is to improve the safety and efficiency of store operations by preventing theft in unmanned stores and analyzing customer sentiment.In particular, since unmanned stores do not have permanent staff on-site, immediate response is difficult, and conventional surveillance systems have limitations in the accuracy of detecting prohibited behavior and the speed of response.
[1156] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning when specific prohibited behavior is detected based on the text data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance simultaneously with the issuance of the warning and acquiring a facial image of the customer, and means for analyzing customer emotion data using an emotion analysis engine to improve the accuracy of detecting prohibited behavior. This makes it possible to prevent theft in unmanned stores and detect prohibited behavior with high accuracy.
[1157] A "surveillance camera" is a device installed in a store to capture video data.
[1158] "Video data" refers to data that records the state of customers and the inside of a store captured by surveillance cameras and facial recognition cameras.
[1159] A "server" is a device or system that analyzes acquired video data and generates behavioral data and emotional data.
[1160] "Text data" is the analysis of video data and the representation of customer behavior as text information.
[1161] "Specific prohibited activities" refer to activities that should not be carried out in the store, including, for example, theft.
[1162] An "alert" is a notification that is sent when certain prohibited behavior is detected.
[1163] A "communication terminal" refers to a device owned by a store manager that is capable of receiving notifications such as warnings.
[1164] A "facial recognition camera" is a camera used to capture and recognize facial images of customers.
[1165] "Emotion data" is data generated by analyzing customer emotions using an emotion analysis engine.
[1166] An "emotion analysis engine" is software or algorithms for analyzing a customer's emotional state from acquired video data.
[1167] "Storage" is a storage device that stores data for a long period of time and allows it to be accessed later.
[1168] In this invention, a system is constructed by integrating multiple pieces of hardware and software to improve the monitoring and security of unmanned stores.
[1169] Hardware and software used
[1170] 1. Hardware
[1171] Surveillance camera: Used to capture video data within the store.
[1172] Facial recognition camera: Used to capture facial images of customers when an alert is issued.
[1173] Communication terminal: A device that allows store managers to receive alert notifications.
[1174] Server: A central device that analyzes data and manages the generated behavioral and emotional data.
[1175] 2. Software
[1176] Generative AI models (e.g., OpenAI's GPT series): Analyze captured video data and convert customer behavior into text data.
[1177] Emotion engine (e.g. Microsoft's Emotion API): Analyzes customer emotion data.
[1178] Notification system (e.g., LINE API, Twilio): Sends warnings to the store manager's communication device based on the prohibited behavior detected.
[1179] System Operation Overview
[1180] Video data acquisition and analysis
[1181] Surveillance cameras capture real-time video data of customer behavior in the store, and the server passes this video data to a generative AI model to generate behavioral data.
[1182] Examples:
[1183] Customers are seen on surveillance cameras picking up items from the shelves.
[1184] The server inputs the video data into a generative AI model and generates behavioral data such as "a customer picking up a product from a shelf."
[1185] The server then uses an emotion engine to generate "emotional data" of the customer from the video data, and obtains the result that "the customer is interested."
[1186] Example prompt for a generative AI model:
[1187] Analyze the video data below and convert customer behavior into text data.
[1188] Video footage: A customer picks up an item from the shelf and heads straight for the exit.
[1189] Output of the generative AI model: The customer picks up the item and heads to the exit without going through the checkout.
[1190] Detecting and notifying prohibited behavior
[1191] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether prohibited behavior is included. If prohibited behavior is detected, the server sends a warning to the store manager's communication device via the notification system.
[1192] Examples:
[1193] Behavioral data records that "the customer picked up the product and headed for the exit without going through the cash register."
[1194] The emotional data includes information that "the customer is nervous."
[1195] The server uses this information to detect prohibited behavior and sends a warning notification saying, "A customer may have stolen a product."
[1196] Activating the face recognition camera and acquiring video
[1197] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture the customer's facial image, which is then saved by the server.
[1198] Examples:
[1199] The server will send a warning notification and simultaneously activate the facial recognition camera.
[1200] The camera captures the customer's face and sends the image to a server.
[1201] The server stores the acquired facial image as evidence.
[1202] Example of a prompt for the emotion engine:
[1203] Analyze the video data below and convert customer sentiment into text data.
[1204] Footage: Customers are seen picking up items and heading for the exit.
[1205] Emotion Engine Output: Customer is nervous.
[1206] Information storage and access
[1207] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. These data can be accessed by store managers and later downloaded and provided to external organizations as needed.
[1208] Examples:
[1209] The server stores the video data, text data, and emotional data when prohibited behavior is detected in dedicated storage.
[1210] Store managers can access the storage and download this data as needed, providing it to external agencies such as the police.
[1211] This system will enable the prevention and rapid response of theft in unmanned stores, significantly improving the safety of store operations. In addition, by using an emotion engine, it is possible to analyze not only customer behavior but also emotions, enabling more accurate detection of prohibited behavior.
[1212] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1213] Step 1:
[1214] The server acquires video data in real time from surveillance cameras installed in the store. Specifically, the cameras capture customer behavior in the store and send the video data to the server. The input data is the video data from the surveillance cameras, and the output data is the video data itself.
[1215] Step 2:
[1216] The server passes the acquired video data to a generative AI model and converts the customer's behavior into text data. Specifically, the server inputs the video data into a generative AI model (e.g., OpenAI's GPT series) and generates behavioral data such as "the customer picking up a product." The input data is video data, and the output data is behavioral data in text format.
[1217] Step 3:
[1218] The server then uses an emotion engine to analyze the customer's emotions from the video data. Specifically, the server inputs the video data into an emotion engine (e.g., Microsoft's Emotion API) and generates emotion data indicating that the customer is interested. The input data is video data, and the output data is emotion data in text format.
[1219] Step 4:
[1220] The server analyzes the generated behavioral and emotional data to detect whether it contains specific prohibited behavior. Specifically, it uses an analysis algorithm to identify prohibited behavior, such as "a customer picks up a product and heads toward the exit without going through the cash register," and also analyzes emotional data, such as "the customer is nervous." The input data are behavioral and emotional data, and the output data are the results of the detection of prohibited behavior.
[1221] Step 5:
[1222] If prohibited behavior is detected, the server sends a warning to the store manager's communication device via the notification system. Specifically, the server uses the LINE API or Twilio to send a warning message to the store manager's smartphone, such as "A customer may have stolen an item." The input data is the detection result, and the output data is the warning notification.
[1223] Step 6:
[1224] Immediately after issuing the warning, the server automatically activates the facial recognition camera installed near the entrance. Specifically, the server sends a start command to the facial recognition camera to start capturing facial images. The input data is the start command, and the output data is facial image data.
[1225] Step 7:
[1226] The facial recognition camera captures the customer's facial image and sends it to a server, which then stores the data. Specifically, the server receives the facial image data and stores it in dedicated storage. This storage can be accessed later. The input data is the customer's facial image, and the output data is the data stored in the storage.
[1227] Step 8:
[1228] The server saves the acquired video data, the generated behavioral and emotional data, and the notification log in storage. Specifically, the server stores each data in a dedicated storage device so that it can be preserved for a long period of time. The input data is the video data, behavioral data, emotional data, and notification log, and the output data is the data saved in storage.
[1229] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1230] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1231] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1232] [Fourth embodiment]
[1233] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1234] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1235] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1236] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1237] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1238] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1239] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1240] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1241] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1242] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1243] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1244] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1245] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1246] The present invention is a security system designed to prevent theft in unmanned stores. This system acquires video data from surveillance cameras installed in the store, analyzes the video data, and converts customer behavior into text data. If a specific prohibited behavior is detected, the system issues a warning and notifies the store manager's communication terminal. The system also has the function of activating a face recognition camera installed near the entrance and capturing a video of the customer's face at the same time as issuing the warning.
[1247] Program processing
[1248] Video data acquisition and analysis
[1249] The server collects video data in real time from surveillance cameras installed in the store, then passes the collected video data to a generative AI model, which analyzes customer behavior and converts it into text data.
[1250] Examples:
[1251] The camera captures users entering the store and browsing the shelves.
[1252] The server inputs the video into a generative AI model and outputs the text data, "Customer browsing the shelves."
[1253] NG word detection and notification
[1254] The server analyzes the text data created by the generative AI model and detects whether it contains specific prohibited behavior (e.g., "no money was inserted"). If prohibited behavior is detected, the server issues an alert and sends a notification to the store manager's communication device (e.g., LINE or phone).
[1255] Examples:
[1256] If the text data records that "the customer picked up the product and headed for the entrance without going through the cash register."
[1257] The server determines that this behavior is prohibited and sends a warning notification via LINE or phone stating, "The customer may have stolen a product."
[1258] Activating the face recognition camera and acquiring video
[1259] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[1260] Examples:
[1261] The server sends a warning notification and at the same time sends a command to activate the facial recognition camera.
[1262] The camera captures the customer's face and sends the image to a server.
[1263] The server stores the acquired facial images.
[1264] Information storage and access
[1265] The server stores all captured video data, generated text data, and log data when prohibited behavior is detected. This data can be accessed later by store managers and provided to external agencies such as the police.
[1266] Examples:
[1267] The server stores the video data and text data when prohibited behavior is detected in dedicated storage.
[1268] Store managers can later access the storage and download the data if necessary, and provide it to police.
[1269] This security system will enable the prevention of theft in unmanned stores and a rapid response, significantly improving the safety of store operations.
[1270] The processing flow will be explained below.
[1271] Step 1:
[1272] The server receives video data in real time from surveillance cameras installed in the store, allowing the situation inside the store to be constantly monitored.
[1273] Specific behavior:
[1274] Surveillance cameras capture footage.
[1275] The server receives video data from the surveillance camera and stores it in a buffer.
[1276] Step 2:
[1277] The server passes the captured video data to a generative AI model, which analyzes customer behavior in the video. This analysis converts the video into text data.
[1278] Specific behavior:
[1279] The server inputs the video data frame by frame into the generative AI model.
[1280] The generative AI model analyzes the frames and outputs text data such as "A customer is browsing the shelves."
[1281] The server stores the text data.
[1282] Step 3:
[1283] The server analyzes the text data generated by the generative AI model to detect whether it contains certain prohibited behaviors, which triggers the next action.
[1284] Specific behavior:
[1285] The server analyzes the text data sentence by sentence.
[1286] Based on the analysis results, it is checked whether prohibited actions (e.g., "did not put money in") are included.
[1287] If a prohibited action is detected, the information is recorded in an internal flag.
[1288] Step 4:
[1289] If a prohibited activity is detected, the server issues an alert and sends a notification to the store manager's communication terminal, which includes details of the prohibited activity.
[1290] Specific behavior:
[1291] The server generates a warning message based on an internal flag.
[1292] Using the LINE API, a message is sent to the store manager's LINE account stating, "A customer may have stolen an item."
[1293] In the case of telephone notification, an automated voice warning message will be sent.
[1294] Step 5:
[1295] The server automatically activates a facial recognition camera installed near the entrance to capture images of the customer's face, which can then be used as evidence.
[1296] Specific behavior:
[1297] The server sends a start command to the face recognition camera.
[1298] A facial recognition camera captures the customer's face and sends the video data to a server.
[1299] The server stores the acquired facial images.
[1300] Step 6:
[1301] The server stores the captured video data, text data, and log data when prohibited behavior is detected, which will be used for later analysis and reporting to the police.
[1302] Specific behavior:
[1303] The server stores the video data and text data when prohibited behavior is detected in dedicated storage.
[1304] Video data obtained from facial recognition cameras will also be stored in the same way.
[1305] The stored data is accessible to store managers.
[1306] Step 7:
[1307] The user (store manager) receives the notification from the server and reports it to the police if necessary. The user accesses the data on the server and provides the necessary video and text data to the police.
[1308] Specific behavior:
[1309] The store manager will check the warning notification received via LINE or phone.
[1310] The store manager accesses the data stored on the server and downloads the video and text data of the theft.
[1311] If necessary, provide this data to the police.
[1312] The above is a specific processing flow of the system of the present invention.
[1313] Example 1
[1314] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1315] In recent years, the risk of theft has increased with the increase in unmanned stores. However, current surveillance systems have difficulty detecting theft in real time and responding quickly. Furthermore, conventional systems lack a means to reliably and quickly notify store managers, making it difficult to prevent theft. This poses a risk to the safety of store operations.
[1316] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1317] In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning if specific prohibited behavior is detected based on the text data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance simultaneously with the issuance of the warning to acquire a face image of the customer, means for inputting the acquired video data and text data into a generative AI model, and means for detecting whether specific prohibited behavior is included based on the analysis results of the generative AI model. This enables real-time detection of theft in unmanned stores and rapid response, significantly improving the safety of store operations.
[1318] A "surveillance camera" is a device that is installed for the purpose of monitoring a specific location and acquires video data.
[1319] "Video data" refers to image information and video information captured by a surveillance camera.
[1320] A "generative AI model" is a model that uses artificial intelligence to analyze input data and generate text or other information as output.
[1321] "Text data" refers to character information obtained by analyzing video data.
[1322] "Prohibited behavior" refers to specific behaviors that are performed within a store and that violate predetermined rules.
[1323] A "warning" is a notification issued when a prohibited action is detected.
[1324] A "communication terminal" is a device for sending and receiving information, and includes smartphones, tablets, etc.
[1325] A "face recognition camera" is a device that detects an individual's face and acquires its image data.
[1326] "Storage" refers to a storage device or storage service for storing data.
[1327] "Log data" is data that records historical information about events and operations that occur within the system.
[1328] This invention is a security system for preventing theft in unmanned stores, and operates in the following procedure, centered around a server.
[1329] First, the server acquires video data in real time from the surveillance cameras installed in the store. These cameras are often standard IP cameras, such as Hikvision cameras. The server then stores this video data in a specific buffer.
[1330] Next, the server inputs the captured video data frame by frame into a generative AI model, such as OpenAI's GPT-4, which has image analysis capabilities. The server sends the following prompt to the generative AI model:
[1331] "Analyze the image below and explain the customer behavior. Image: [Image data]"
[1332] The generative AI model then analyzes the video data and outputs the customer's behavior as text data, such as "The customer is browsing the shelves."
[1333] The server then analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behavior. Prohibited behavior includes "taking away merchandise without going through the cash register." If prohibited behavior is detected, the server issues a warning and sends a notification to the store manager's communication device. Notifications can be sent using the LINE API or SMS API, and a specific example would be a warning message stating, "A customer may have stolen an item."
[1334] Furthermore, at the same time as issuing the warning, the server activates a facial recognition camera installed near the entrance. This facial recognition camera is equipped with facial recognition software such as Face++ or Amazon Rekognition. The server then sends a capture command to the camera, which captures and stores the customer's facial image.
[1335] Finally, the server stores all captured video data, generated text data, and log data when prohibited behavior is detected in a secure storage location, using Amazon S3 or Google Cloud Storage. Store managers can access this storage and download the data as needed, providing it to external agencies (such as the police).
[1336] In this way, it becomes possible to detect theft in unmanned stores in real time and respond quickly, greatly improving the safety of store operations.
[1337] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1338] Step 1: Acquire video data
[1339] The server acquires video data from surveillance cameras installed in the store. The server periodically accesses the surveillance cameras to acquire new video data. General IP cameras are used for this purpose. The input is the video data acquired from the surveillance cameras, and the output is the video data stored in the temporary storage buffer within the server.
[1340] Example: A server captures video streams from a Hikvision IP camera once per second and stores them in a buffer for analysis.
[1341] Step 2: Analyzing the video data
[1342] The server passes the acquired video data to a generative AI model and converts customer behavior into text data. The server then divides the video data into frames and inputs each frame image into the generative AI model. OpenAI's GPT-4 and other models are used as generative AI models. The input is each frame of video data, and the output is text data.
[1343] Example: The server sends an "image recognition prompt" to OpenAI's GPT-4 and saves the returned text data in an analysis folder.
[1344] Example prompt: "Analyze the image below and describe the customer's behavior. Image: [image data]"
[1345] Step 3: Analyzing the text data
[1346] The server analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behavior. The input is the text data output by the generative AI model, and the output is a flag indicating whether or not prohibited behavior exists. Analysis is performed using natural language processing technology based on the definition of prohibited behavior.
[1347] Example: The server detects text that describes prohibited actions, such as "taking merchandise without going through the cash register."
[1348] Step 4: Issuing warnings and notifications
[1349] If the server detects a prohibited behavior, it issues a warning and sends a notification to the store manager's communication device. The server sends the notification using a pre-configured communication method (e.g., LINE API or SMS API). The input is a flag for the prohibited behavior, and the output is the warning notification sent.
[1350] Example: The server uses the LINE API to send a notification to the store manager saying, "A customer may have stolen an item."
[1351] Step 5: Turn on the face recognition camera
[1352] Immediately after issuing the warning, the server activates a facial recognition camera installed near the entrance and captures the customer's facial image. The server accesses the specified facial recognition camera and sends a capture command. The input is a warning issuance flag, and the output is the captured facial image data.
[1353] Example: The server uses the Face++ API to send a "capture start command" and saves the captured facial image on the server.
[1354] Step 6: Store and access your data
[1355] The server securely stores the captured video data, generated text data, and log data when prohibited behavior is detected. This storage uses Amazon S3 or Google Cloud Storage. The input is various types of data, and the output is data stored in the cloud storage.
[1356] Example: The server uploads video data, text data, and warning logs related to prohibited behavior to Amazon S3 and generates a link for later access.
[1357] The terminal (store manager) accesses this data and provides it to external agencies (e.g., police) if necessary.
[1358] Example: A store manager opens the management app, clicks a URL link, and downloads the necessary video and log data.
[1359] (Application example 1)
[1360] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1361] Preventing theft in unmanned stores and responding quickly and reliably are key challenges. Currently, monitoring in unmanned stores is inefficient, which can lead to delayed detection and countermeasures when theft occurs. Another problem is that information on detected theft is not properly recorded and managed, making it impossible to use as evidence later.
[1362] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1363] In this invention, the server includes: means for acquiring video data from surveillance cameras installed in the store; means for analyzing the acquired video data and converting customer behavior into text data; means for issuing a warning if specific prohibited behavior is detected based on the text data; means for using remote communication technology to notify the store manager's communication terminal of the warning; means for using the server to transmit video data in real time and analyze the text data; means for using a generative AI model to analyze the text data; and means for activating a facial recognition camera installed near the entrance and capturing facial images of customers simultaneously with the issuance of the warning. This enables the immediate detection of theft in an unmanned store and the notification of a warning to the store manager. Furthermore, capturing and storing facial images of customers at the same time can be used as evidence at a later date.
[1364] A "surveillance camera" is a device installed in a store to acquire video data.
[1365] "Video data" refers to video information acquired by a surveillance camera.
[1366] "Analysis" is the process of analyzing customer behavior based on the acquired video data.
[1367] "Text data" is character data generated as a result of analyzing video data.
[1368] "Prohibited Behavior" refers to any behavior that is not permitted within the store.
[1369] A "warning" is a notification issued when prohibited behavior is detected.
[1370] A "communication terminal" is an information and communication device such as a smartphone or computer used by a store manager.
[1371] A "facial recognition camera" is a camera that recognizes the facial characteristics of a specific person and acquires video data.
[1372] A "generative AI model" is a model for generating and analyzing data using artificial intelligence.
[1373] "Telecommunications technology" refers to technology for sending and receiving information to and from communication terminals.
[1374] A "server" is a computer system that processes and stores data.
[1375] "Storage" refers to a storage device for saving data.
[1376] "Real-time" refers to processing that is immediate and without delay.
[1377] The system for implementing this invention consists of surveillance cameras installed in a store, a generative AI model for analyzing video data, a server for detecting prohibited behavior and sending warnings, a communication terminal for the store manager, a facial recognition camera, and storage for data storage.
[1378] Specific system configuration and operation
[1379] 1. Acquiring video data
[1380] The server acquires video data in real time from surveillance cameras installed in the store. The surveillance cameras capture the situation inside the store 24 hours a day and send the video to the server.
[1381] 2. Analysis of video data
[1382] The server inputs the acquired video data into a generative AI model to analyze customer behavior. The generative AI model converts the customer behavior into text data based on the video data. This generative AI model incorporates pre-learned behavioral patterns, allowing it to analyze specific behavior in detail.
[1383] An example of a prompt sentence is, "Please describe in text the actions of the person in the video. In particular, please detect the action of picking up a product and heading towards the entrance / exit without going through the cash register."
[1384] 3. Detecting and warning against prohibited behavior
[1385] The server analyzes the text data created by the generative AI model and detects whether it contains certain prohibited behaviors. For example, if the server detects behavior such as "picking up a product and heading to the entrance without going through the cash register," it issues an alert and immediately sends a notification to the store manager's communication terminal. This notification is sent using remote communication technology.
[1386] 4. Activating the face recognition camera and acquiring video
[1387] When the warning is issued, the server activates a facial recognition camera installed near the entrance to capture the facial image of the customer. The facial recognition camera quickly captures the face of a specific person and sends the image data to the server.
[1388] 5. Data storage and management
[1389] The server stores video data acquired from surveillance cameras and facial recognition cameras, as well as text data created by generative AI models, allowing store managers to access the storage later and download the data as needed to provide it to external agencies such as the police.
[1390] Specific examples
[1391] For example, if a surveillance camera captures a customer picking up an item and then heading for the entrance without going through the cash register, the video data is sent to a server. The server uses a generative AI model to generate text data from the video, such as "The customer picks up an item and heads for the entrance without going through the cash register." Based on this text data, the server detects prohibited behavior and issues a warning to the store manager. At the same time, a facial recognition camera captures the customer's face, and the video data is sent to the server for storage.
[1392] In this way, the system of the present invention makes it possible to prevent and quickly respond to theft in unmanned stores.
[1393] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1394] Step 1:
[1395] Acquiring video data
[1396] Subject: Server
[1397] The server acquires video data in real time from the surveillance cameras installed in the store. These surveillance cameras operate 24 hours a day and continuously transmit images from inside the store to the server. The input is live video from the surveillance cameras, and the output is raw data stored on the server.
[1398] Step 2:
[1399] Video data analysis
[1400] Subject: Server
[1401] The server inputs the acquired video data into the generative AI model and analyzes customer behavior. Specifically, the video data sent from the surveillance camera is sent to the generative AI model, which analyzes "what kind of behavior the customer is exhibiting." The prompt text used is "Please describe in text the behavior of the people in the video. In particular, please detect the behavior of someone picking up a product and heading toward the entrance / exit without going through the cash register." The input is the video data from the surveillance camera, and the output is the analyzed text data.
[1402] Step 3:
[1403] Detecting prohibited behavior
[1404] Subject: Server
[1405] The server analyzes the text data created by the generative AI model to detect whether it contains specific prohibited behaviors. During this process, it checks whether the text data contains pre-set prohibited behavior keywords (e.g., heading to the entrance / exit without going through the cash register). The input is the analyzed text data, and the output is the results of the detection of prohibited behaviors. Specific operations involve the use of a keyword matching algorithm.
[1406] Step 4:
[1407] Warnings and Notifications
[1408] Subject: Server
[1409] If a prohibited behavior is detected, the server issues an alert and sends a notification to the store manager's communication device. This notification is sent using remote communication technologies (e.g., push notification, SMS). The input is the result of the detection of the prohibited behavior, and the output is a warning notification sent to the communication device. The device can then analyze the notification and take appropriate action.
[1410] Step 5:
[1411] Activating the face recognition camera and acquiring video
[1412] Subject: Server
[1413] When the server issues a warning, it activates a facial recognition camera installed near the entrance and captures the customer's facial image. This camera quickly captures the face of a specific person and sends the image data to the server. The input is the warning event, and the output is the facial image data from the facial recognition camera.
[1414] Step 6:
[1415] Data storage and access
[1416] Subject: Server
[1417] The server stores video data acquired from surveillance cameras and facial recognition cameras, as well as text data created by the generative AI model. Store managers can access this data later and provide it to external agencies such as the police if necessary. The input is video data and text data, and the output is the data stored in the storage. Specific operations use a database management system (e.g., SQL, NoSQL).
[1418] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1419] This invention is a security system designed to prevent theft and analyze customer emotions in unmanned stores. This system acquires video data from surveillance cameras installed in the store, analyzes the video data, and converts customer behavior into text data. If specific prohibited behavior is detected, the system then issues a warning and notifies the store manager's communication terminal. The system also has the function of simultaneously issuing a warning and activating a facial recognition camera installed near the entrance to capture video of the customer's face. Furthermore, by incorporating an emotion engine that recognizes user emotions and acquiring and analyzing customer emotion data, the system can improve the accuracy of detecting prohibited behavior.
[1420] Program processing
[1421] Video data acquisition and analysis
[1422] The server collects video data in real time from surveillance cameras installed in the store, then passes the video data to a generative AI model and emotion engine to analyze customer behavior and emotions and convert them into text and emotion data.
[1423] Examples:
[1424] The camera captures users entering the store and browsing the shelves.
[1425] The server inputs the video into a generative AI model and emotion engine, generating text data such as "The customer is browsing the shelves" and emotion data such as "The customer is interested."
[1426] NG word detection and notification
[1427] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether it contains certain prohibited behaviors, which triggers the next action.
[1428] Examples:
[1429] If the text data records that "the customer picked up the product and headed for the entrance without going through the cash register."
[1430] Suppose the emotional data contains information that "the customer is nervous."
[1431] The server determines that this behavior and emotion corresponds to prohibited behavior and sends a warning notification via LINE or phone stating, "The customer may have stolen a product."
[1432] Activating the face recognition camera and acquiring video
[1433] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[1434] Examples:
[1435] The server sends a warning notification and at the same time sends a command to activate the facial recognition camera.
[1436] The camera captures the customer's face and sends the image to a server.
[1437] The server stores the acquired facial images.
[1438] Information storage and access
[1439] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. These data can be accessed later by store managers and provided to external agencies such as the police.
[1440] Examples:
[1441] The server stores video and text data in dedicated storage when prohibited behavior is detected.
[1442] Emotion data generated by the emotion engine is also stored in the same way.
[1443] Store managers can later access the storage and download the data if necessary, and provide it to police.
[1444] This security system will enable the prevention of theft in unmanned stores and a rapid response, significantly improving the safety of store operations. In addition, by using an emotion engine, it will be possible to analyze not only customer behavior but also emotions, enabling more accurate detection of prohibited behavior.
[1445] The processing flow will be explained below.
[1446] Step 1:
[1447] The server collects video data in real time from surveillance cameras installed in the store, allowing it to constantly monitor the situation inside the store.
[1448] Specific behavior:
[1449] Surveillance cameras capture footage.
[1450] The server receives video data from the surveillance camera and stores it in a buffer.
[1451] Step 2:
[1452] The server passes the acquired video data to the generative AI model and emotion engine, which analyzes the customer's behavior and emotions in the video. Based on the analysis results, the video data is converted into text data and emotion data.
[1453] Specific behavior:
[1454] The server inputs the video data frame by frame into the generative AI model and emotion engine.
[1455] The generative AI model outputs the text data "A customer is browsing the shelves."
[1456] The emotion engine generates emotion data that indicates "customer interest."
[1457] The server stores the text data and emotion data.
[1458] Step 3:
[1459] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether it contains specific prohibited behaviors. If prohibited behaviors are detected, the server proceeds to the next step.
[1460] Specific behavior:
[1461] The server analyzes the text data and detects the behavior of "a customer picking up a product and heading to the entrance / exit without going through the cash register."
[1462] The server analyzes the emotional data and confirms that the customer is nervous.
[1463] The server determines whether the behavior is prohibited based on the behavior and emotions.
[1464] Step 4:
[1465] If a prohibited activity is detected, the server issues an alert and sends a notification to the store manager's communication terminal, which includes details of the prohibited activity.
[1466] Specific behavior:
[1467] The server generates a warning message based on the result of the detection of the prohibited behavior.
[1468] Using the LINE API, a message is sent to the store manager's LINE account stating, "A customer may have stolen an item."
[1469] In the case of telephone notifications, an automated voice message will warn, "A customer may have stolen your product."
[1470] Step 5:
[1471] At the same time as issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture an image of the customer's face, which is then saved as evidence.
[1472] Specific behavior:
[1473] The server sends a start command to the face recognition camera.
[1474] A facial recognition camera captures the customer's face and sends the video data to a server.
[1475] The server stores the acquired facial images.
[1476] Step 6:
[1477] The server stores all captured video data, generated text data, and log data when prohibited behavior is detected, and these data can be accessed later by store managers.
[1478] Specific behavior:
[1479] The server stores video and text data in dedicated storage when prohibited behavior is detected.
[1480] Emotion data generated by the emotion engine is also stored in the same way.
[1481] The saved data is stored in a format that can be accessed by the store manager.
[1482] Step 7:
[1483] The user (store manager) receives the notification from the server and reports it to the police if necessary. The user accesses the data on the server and provides the necessary video and text data to the police.
[1484] Specific behavior:
[1485] The store manager will check the warning notification received via LINE or phone.
[1486] The store manager accesses the data stored on the server and downloads the video and text data of the theft.
[1487] If necessary, provide this data to the police.
[1488] The above is a specific processing flow of the system of the present invention, which significantly improves the security function of unmanned stores, suppresses theft, and realizes highly accurate analysis based on customer emotions.
[1489] Example 2
[1490] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1491] To effectively prevent theft in unmanned stores, it is necessary not only to analyze customer behavior but also to accurately grasp their emotions and respond promptly and appropriately. Conventional security systems focus on behavioral analysis and do not consider emotion analysis, resulting in false positives and oversights. Thus, there is a need for methods to improve the accuracy and safety of theft prevention in unmanned stores.
[1492] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1493] In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning when specific prohibited behavior is detected based on the text data and emotion data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance at the same time as issuing the warning and acquiring an image of the customer's face, and means incorporating an emotion recognition engine for analyzing the customer's behavior and emotion and converting it into text data and emotion data. This makes it possible to prevent theft in unmanned stores and to perform highly accurate behavioral analysis based on customer emotions.
[1494] A "surveillance camera" is a device that is installed in a store and continuously captures the situation inside the store as video data.
[1495] "Video Data" means a collection of video frames captured by surveillance cameras or facial recognition cameras, and is digital data containing customer behavior and other significant events.
[1496] "Analysis" is the process of analyzing the captured video data using specific algorithms and AI models to extract customer behavior and emotions.
[1497] "Customer" refers to a shopper or visitor to an unmanned store.
[1498] "Text data" refers to data that converts customer behavior and status obtained as analysis results into text information.
[1499] "Emotion data" is data that indicates the psychological state of a customer, obtained from their facial expressions and actions.
[1500] "Prohibited behavior" refers to stealing merchandise or other behavior that violates store rules.
[1501] "Warning" refers to an alert or notification issued when prohibited behavior is detected.
[1502] A "communication terminal" is a digital device such as a smartphone, tablet, or PC owned by a store manager.
[1503] A "facial recognition camera" is a camera installed to identify customers' faces and acquire video data.
[1504] An "emotion recognition engine" is software or an algorithm for analyzing customer emotions from video data.
[1505] "Storage" refers to cloud-based or physical data storage devices for storing captured data (video data, text data, emotion data, etc.).
[1506] "Store Manager" means a person or organization responsible for the operation and management of an unmanned store.
[1507] The present invention is a security system for preventing theft and analyzing customer sentiment in unmanned stores. The system operates using multiple hardware and software components.
[1508] Hardware Configuration
[1509] This system includes the following main hardware components:
[1510] Surveillance cameras: Installed in stores, they capture customer behavior in real time.
[1511] Facial recognition cameras: Installed near the store entrance, they capture images of customers' faces under certain conditions.
[1512] Communication device: Smartphone, tablet, PC, etc. owned by the store manager.
[1513] Software Configuration
[1514] The operation of the system is supported by the following software:
[1515] Generative AI models: For example, using GPT-4 for natural language processing to convert video data into text data.
[1516] Emotion recognition engine: For example, using Microsoft Azure's emotion analysis API to analyze customer emotions from video data.
[1517] Cloud storage: For example, Google Cloud Storage is used to store the acquired data.
[1518] Processing flow
[1519] The operation of the system proceeds as follows.
[1520] Acquisition of video data: The server acquires video data from the surveillance cameras in real time. The video data is sent to the server using a streaming protocol.
[1521] Video data analysis: The acquired video data is input into a generative AI model and an emotion recognition engine to analyze the customer's behavior and emotions. The analysis results are stored on the server as text data and emotion data.
[1522] Detection of prohibited behavior: Based on the generated text data and emotion data, the server detects whether certain prohibited behaviors are included. If prohibited behaviors are detected, a warning is issued.
[1523] Sending warning notifications: Warnings are sent in real time to the store manager's communication device via the LINE API or Twilio API.
[1524] Activating the facial recognition camera and capturing video: At the same time as receiving the warning notification, the server activates the facial recognition camera and captures the customer's facial video. The captured video data is stored in cloud storage as evidence.
[1525] Information storage: All captured data is stored in a dedicated cloud storage. Store managers can access this data and download it to provide to the police if necessary.
[1526] Specific examples
[1527] For example, imagine a situation where a user enters a store and looks at the shelves. In this case, a surveillance camera captures the user's behavior, and the server inputs the video data into a generative AI model and emotion recognition engine. The analysis results in text data such as "The customer is browsing the shelves" and emotion data such as "The customer is interested." If prohibited behavior is detected, an alert is immediately issued and the facial recognition camera is activated.
[1528] Example prompt sentence:
[1529] "Analyze real-time video footage from security cameras to detect customer behavior and emotions. Then, write a program to issue a warning notification, activate a facial recognition camera, and save the footage if prohibited behavior is detected."
[1530] This invention not only prevents theft in unmanned stores and enables highly accurate analysis of customer behavior, but also takes emotional data into account, enabling more accurate detection of prohibited behavior and quicker responses.
[1531] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1532] Step 1:
[1533] Acquiring video data
[1534] The server receives real-time video data from the surveillance cameras installed in the store. The input is the live video stream from the surveillance cameras. This video data is sent to the server using the RTSP protocol and buffered in memory. The output is a series of video frames.
[1535] What happens: The server connects to the video stream using the camera's RTSP URL and captures video frames, each ready to be sent for analysis.
[1536] Step 2:
[1537] Video data analysis
[1538] The server passes the acquired video data to a generative AI model and an emotion recognition engine to analyze customer behavior and emotions. The input is the video frames acquired in step 1. Data processing is performed by inputting the video frames into a generative AI model (e.g., OpenAI's GPT-4) and an emotion recognition engine (e.g., Microsoft Azure's Sentiment Analysis API). The output is text data and emotion data.
[1539] How it works: The server uses a Python script to input each video frame into a generative AI model, obtaining text data such as "A customer is browsing the shelves." At the same time, it sends the frame to an emotion recognition engine to obtain emotion data such as "The customer is interested."
[1540] Step 3:
[1541] Detecting prohibited behavior
[1542] The server analyzes the generated text data and emotion data to detect whether specific prohibited behaviors are included. The inputs are the text data and emotion data generated in step 2. Data calculations are performed by analyzing this data based on specific algorithms and regular expression patterns. The output is a determination result indicating whether prohibited behaviors have been detected.
[1543] Specific operation: The server performs regular expression pattern matching to detect phrases such as "The customer picked up the product and headed for the entrance without going through the cash register" from the text data. If the emotion data contains information such as "The customer is nervous," it records this as a prohibited behavior.
[1544] Step 4:
[1545] Sending warning notifications
[1546] If the server detects a prohibited behavior, it immediately sends a warning to the store manager's communication terminal. The input is the prohibited behavior judgment result obtained in step 3. Data processing involves generating and sending a warning message. The output is a warning notification sent to the store manager's communication terminal.
[1547] Specific operation: The server uses the LINE API to generate a message such as "A customer may have stolen an item" and sends it to the store manager's LINE account. At the same time, it also sends a phone notification using the Twilio API.
[1548] Step 5:
[1549] Activating the face recognition camera and acquiring video
[1550] Immediately after issuing the warning, the server automatically activates the facial recognition camera installed near the entrance and captures an image of the customer's face. The input is the warning notification from step 4. Data processing involves capturing and saving the facial image. The output is the captured facial image data.
[1551] Specific operation: The server sends the specified API request to the facial recognition camera to activate the camera. The camera captures the customer's face and sends the video to the server in real time. The captured facial video is encoded in the specified format (e.g., MPEG-4) and saved in the server's storage.
[1552] Step 6:
[1553] Retention of Information
[1554] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. The input is all the data generated at each step. Data calculation aggregates this data in one place and stores it in cloud storage. The output is all the stored data.
[1555] Specific operation: When prohibited behavior is detected, the server uploads video data, text data, and emotion data to dedicated cloud storage (e.g., Google Cloud Storage). The data is categorized with a timestamp and can be accessed later by store managers.
[1556] (Application example 2)
[1557] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1558] The challenge is to improve the safety and efficiency of store operations by preventing theft in unmanned stores and analyzing customer sentiment.In particular, since unmanned stores do not have permanent staff on-site, immediate response is difficult, and conventional surveillance systems have limitations in the accuracy of detecting prohibited behavior and the speed of response.
[1559] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from surveillance cameras installed in the store, means for analyzing the acquired video data and converting customer behavior into text data, means for issuing a warning when specific prohibited behavior is detected based on the text data, means for notifying the warning to a communication terminal of a store manager, means for activating a face recognition camera installed near the entrance simultaneously with the issuance of the warning and acquiring a facial image of the customer, and means for analyzing customer emotion data using an emotion analysis engine to improve the accuracy of detecting prohibited behavior. This makes it possible to prevent theft in unmanned stores and detect prohibited behavior with high accuracy.
[1560] A "surveillance camera" is a device installed in a store to capture video data.
[1561] "Video data" refers to data that records the state of customers and the inside of a store captured by surveillance cameras and facial recognition cameras.
[1562] A "server" is a device or system that analyzes acquired video data and generates behavioral data and emotional data.
[1563] "Text data" is the analysis of video data and the representation of customer behavior as text information.
[1564] "Specific prohibited activities" refer to activities that should not be carried out in the store, including, for example, theft.
[1565] An "alert" is a notification that is sent when certain prohibited behavior is detected.
[1566] A "communication terminal" refers to a device owned by a store manager that is capable of receiving notifications such as warnings.
[1567] A "facial recognition camera" is a camera used to capture and recognize facial images of customers.
[1568] "Emotion data" is data generated by analyzing customer emotions using an emotion analysis engine.
[1569] An "emotion analysis engine" is software or algorithms for analyzing a customer's emotional state from acquired video data.
[1570] "Storage" is a storage device that stores data for a long period of time and allows it to be accessed later.
[1571] In this invention, a system is constructed by integrating multiple pieces of hardware and software to improve the monitoring and security of unmanned stores.
[1572] Hardware and software used
[1573] 1. Hardware
[1574] Surveillance camera: Used to capture video data within the store.
[1575] Facial recognition camera: Used to capture facial images of customers when an alert is issued.
[1576] Communication terminal: A device that allows store managers to receive alert notifications.
[1577] Server: A central device that analyzes data and manages the generated behavioral and emotional data.
[1578] 2. Software
[1579] Generative AI models (e.g., OpenAI's GPT series): Analyze captured video data and convert customer behavior into text data.
[1580] Emotion engine (e.g. Microsoft's Emotion API): Analyzes customer emotion data.
[1581] Notification system (e.g., LINE API, Twilio): Sends warnings to the store manager's communication device based on the prohibited behavior detected.
[1582] System Operation Overview
[1583] Video data acquisition and analysis
[1584] Surveillance cameras capture real-time video data of customer behavior in the store, and the server passes this video data to a generative AI model to generate behavioral data.
[1585] Examples:
[1586] Customers are seen on surveillance cameras picking up items from the shelves.
[1587] The server inputs the video data into a generative AI model and generates behavioral data such as "a customer picking up a product from a shelf."
[1588] The server then uses an emotion engine to generate "emotional data" of the customer from the video data, and obtains the result that "the customer is interested."
[1589] Example prompt for a generative AI model:
[1590] Analyze the video data below and convert customer behavior into text data.
[1591] Video footage: A customer picks up an item from the shelf and heads straight for the exit.
[1592] Output of the generative AI model: The customer picks up the item and heads to the exit without going through the checkout.
[1593] Detecting and notifying prohibited behavior
[1594] The server analyzes the text and emotion data generated by the generative AI model and emotion engine to detect whether prohibited behavior is included. If prohibited behavior is detected, the server sends a warning to the store manager's communication device via the notification system.
[1595] Examples:
[1596] Behavioral data records that "the customer picked up the product and headed for the exit without going through the cash register."
[1597] The emotional data includes information that "the customer is nervous."
[1598] The server uses this information to detect prohibited behavior and sends a warning notification saying, "A customer may have stolen a product."
[1599] Activating the face recognition camera and acquiring video
[1600] Immediately after issuing the warning, the server automatically activates a facial recognition camera installed near the entrance to capture the customer's facial image, which is then saved by the server.
[1601] Examples:
[1602] The server will send a warning notification and simultaneously activate the facial recognition camera.
[1603] The camera captures the customer's face and sends the image to a server.
[1604] The server stores the acquired facial image as evidence.
[1605] Example of a prompt for the emotion engine:
[1606] Analyze the video data below and convert customer sentiment into text data.
[1607] Footage: Customers are seen picking up items and heading for the exit.
[1608] Emotion Engine Output: Customer is nervous.
[1609] Information storage and access
[1610] The server stores the captured video data, generated text data, and log data when prohibited behavior is detected. These data can be accessed by store managers and later downloaded and provided to external organizations as needed.
[1611] Examples:
[1612] The server stores the video data, text data, and emotional data when prohibited behavior is detected in dedicated storage.
[1613] Store managers can access the storage and download this data as needed, providing it to external agencies such as the police.
[1614] This system will enable the prevention and rapid response of theft in unmanned stores, significantly improving the safety of store operations. In addition, by using an emotion engine, it is possible to analyze not only customer behavior but also emotions, enabling more accurate detection of prohibited behavior.
[1615] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1616] Step 1:
[1617] The server acquires video data in real time from surveillance cameras installed in the store. Specifically, the cameras capture customer behavior in the store and send the video data to the server. The input data is the video data from the surveillance cameras, and the output data is the video data itself.
[1618] Step 2:
[1619] The server passes the acquired video data to a generative AI model and converts the customer's behavior into text data. Specifically, the server inputs the video data into a generative AI model (e.g., OpenAI's GPT series) and generates behavioral data such as "the customer picking up a product." The input data is video data, and the output data is behavioral data in text format.
[1620] Step 3:
[1621] The server then uses an emotion engine to analyze the customer's emotions from the video data. Specifically, the server inputs the video data into an emotion engine (e.g., Microsoft's Emotion API) and generates emotion data indicating that the customer is interested. The input data is video data, and the output data is emotion data in text format.
[1622] Step 4:
[1623] The server analyzes the generated behavioral and emotional data to detect whether it contains specific prohibited behavior. Specifically, it uses an analysis algorithm to identify prohibited behavior, such as "a customer picks up a product and heads toward the exit without going through the cash register," and also analyzes emotional data, such as "the customer is nervous." The input data are behavioral and emotional data, and the output data are the results of the detection of prohibited behavior.
[1624] Step 5:
[1625] If prohibited behavior is detected, the server sends a warning to the store manager's communication device via the notification system. Specifically, the server uses the LINE API or Twilio to send a warning message to the store manager's smartphone, such as "A customer may have stolen an item." The input data is the detection result, and the output data is the warning notification.
[1626] Step 6:
[1627] Immediately after issuing the warning, the server automatically activates the facial recognition camera installed near the entrance. Specifically, the server sends a start command to the facial recognition camera to start capturing facial images. The input data is the start command, and the output data is facial image data.
[1628] Step 7:
[1629] The facial recognition camera captures the customer's facial image and sends it to a server, which then stores the data. Specifically, the server receives the facial image data and stores it in dedicated storage. This storage can be accessed later. The input data is the customer's facial image, and the output data is the data stored in the storage.
[1630] Step 8:
[1631] The server saves the acquired video data, the generated behavioral and emotional data, and the notification log in storage. Specifically, the server stores each data in a dedicated storage device so that it can be preserved for a long period of time. The input data is the video data, behavioral data, emotional data, and notification log, and the output data is the data saved in storage.
[1632] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1633] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1634] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1635] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1636] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1637] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1638] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1639] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1640] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1641] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1642] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1643] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1644] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1645] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1646] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1647] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1648] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1649] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1650] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1651] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1652] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1653] The following is further disclosed regarding the above embodiment.
[1654] (Claim 1)
[1655] A means for acquiring video data from a surveillance camera installed in the store;
[1656] A means for analyzing the acquired video data and converting customer behavior into text data;
[1657] means for issuing a warning when a specific prohibited behavior is detected based on the text data;
[1658] means for notifying the warning to a communication terminal of a store manager;
[1659] a means for activating a face recognition camera installed near the entrance at the same time as issuing the warning and capturing an image of the customer's face;
[1660] A system including:
[1661] (Claim 2)
[1662] 10. The system of claim 1, further comprising means for storing video data obtained from the surveillance cameras and facial recognition cameras.
[1663] (Claim 3)
[1664] 10. The system of claim 1, further comprising means for storing the text data and video data in a storage accessible to a store manager.
[1665] "Example 1"
[1666] (Claim 1)
[1667] A means for acquiring video data from a surveillance camera installed in the store;
[1668] A means for analyzing the acquired video data and converting customer behavior into text data;
[1669] means for issuing a warning when a specific prohibited behavior is detected based on the text data;
[1670] means for notifying the warning to a communication terminal of a store manager;
[1671] a means for activating a face recognition camera installed near the entrance at the same time as issuing the warning and capturing an image of the customer's face;
[1672] A means for inputting the acquired video data and text data into a generative AI model;
[1673] A means for detecting whether a specific prohibited behavior is included based on the analysis result of the generative AI model;
[1674] A system including:
[1675] (Claim 2)
[1676] 10. The system of claim 1, further comprising means for storing video data obtained from the surveillance cameras and facial recognition cameras.
[1677] (Claim 3)
[1678] 10. The system of claim 1, further comprising means for storing the text data and video data in a storage accessible to a store manager.
[1679] "Application Example 1"
[1680] (Claim 1)
[1681] A means for acquiring video data from a surveillance camera installed in the store;
[1682] A means for analyzing the acquired video data and converting customer behavior into text data;
[1683] means for issuing a warning when a specific prohibited behavior is detected based on the text data;
[1684] means for notifying the warning to a communication terminal of a store manager;
[1685] a means for activating a face recognition camera installed near the entrance at the same time as issuing the warning and capturing an image of the customer's face;
[1686] means for using a generative AI model to analyze the text data;
[1687] means for using telecommunications technology to transmit said alert to a communications terminal;
[1688] means for using a server to transmit the video data and analyze the text data in real time;
[1689] A system including:
[1690] (Claim 2)
[1691] 10. The system of claim 1, further comprising means for storing video data obtained from the surveillance cameras and facial recognition cameras.
[1692] (Claim 3)
[1693] 10. The system of claim 1, further comprising means for storing the text data and video data in a storage accessible to a store manager.
[1694] "Example 2: Combining Emotion Engines"
[1695] (Claim 1)
[1696] A means for acquiring video data from a surveillance camera installed in the store;
[1697] A means for analyzing the acquired video data and converting customer behavior into text data;
[1698] means for issuing a warning when a specific prohibited behavior is detected based on the text data and emotion data;
[1699] means for notifying the warning to a communication terminal of a store manager;
[1700] a means for activating a face recognition camera installed near the entrance at the same time as issuing the warning and capturing an image of the customer's face;
[1701] A means incorporating an emotion recognition engine that analyzes customer behavior and emotions and converts them into text data and emotion data;
[1702] A system including:
[1703] (Claim 2)
[1704] 10. The system of claim 1, further comprising means for storing video data obtained from the surveillance cameras and facial recognition cameras.
[1705] (Claim 3)
[1706] 10. The system of claim 1, further comprising means for storing the text data and video data in a storage accessible to a store manager.
[1707] "Application example 2 when combining emotion engines"
[1708] (Claim 1)
[1709] A means for acquiring video data from a surveillance camera installed in the store;
[1710] A means for analyzing the acquired video data and converting customer behavior into text data;
[1711] means for issuing a warning when a specific prohibited behavior is detected based on the text data;
[1712] means for notifying the warning to a communication terminal of a store manager;
[1713] a means for activating a face recognition camera installed near the entrance at the same time as issuing the warning and capturing an image of the customer's face;
[1714] A means for analyzing customer emotion data using a sentiment analysis engine to improve the accuracy of detecting prohibited behavior;
[1715] A system including:
[1716] (Claim 2)
[1717] 10. The system of claim 1, further comprising means for storing video data acquired from the surveillance camera and the facial recognition camera, and for storing text data including emotion data.
[1718] (Claim 3)
[1719] The system of claim 1 , further comprising means for storing the text data, emotion data, and video data in a storage accessible to a store manager. [Explanation of symbols]
[1720] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring video data from a surveillance camera installed in the store; A means for analyzing the acquired video data and converting customer behavior into text data; means for issuing a warning when a specific prohibited behavior is detected based on the text data; means for notifying the warning to a communication terminal of a store manager; a means for activating a face recognition camera installed near the entrance at the same time as issuing the warning and capturing an image of the customer's face; A system including:
2. The system of claim 1 , further comprising means for storing video data acquired from the surveillance cameras and facial recognition cameras.
3. The system of claim 1 , further comprising means for storing the text data and the video data in a storage accessible by a store manager.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A