Information processing device, information processing system, and program

The information processing device uses facial and behavioral analysis to identify potential threats and generate alerts, effectively addressing the challenge of detecting harmful acts in facilities.

JP2025119296APending Publication Date: 2025-08-14TOSHIBA TEC KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024014108
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing systems fail to accurately detect individuals likely to commit harmful acts such as shoplifting or dangerous behavior within facilities and provide timely alerts to employees.

Method used

An information processing device equipped with a feature recognition model to analyze facial expressions, appearances, and behaviors from captured images, a suspicious feature list to identify potential threats, and a warning generation model to generate easy-to-understand alerts for staff.

Benefits of technology

Accurately detects individuals engaging in harmful activities and notifies staff in a clear manner, reducing merchandise loss by alerting them to potential shoplifters or dangerous behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025119296000001_ABST
    Figure 2025119296000001_ABST
Patent Text Reader

Abstract

To provide an information processing device capable of accurately detecting a person likely to commit harmful acts and giving a notice of the detected person in an easy-to-understand form.SOLUTION: An information processing device according to an embodiment comprises an acquisition unit, a recognition unit, a detection unit, a generation unit, and a notification unit. The acquisition unit acquires an image including a person from a camera installed in a facility. The recognition unit derives second text information by inputting an image captured by the camera into a learned model that has been learned to derive the second text information describing characteristics of a person included in the image from the acquired image on the basis of a teacher data set including the image including a person and first text information describing at least one characteristic of a person out of a facial expression, appearance, and behavior of the person included in the image. The detection unit detects a security-risk person related to harmful acts on the basis of the second text information. The generation unit generates third text information including characteristics of the detected security-risk person on the basis of detection results. The notification unit notifies an employee of the facility of the third text information.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to an information processing device, an information processing system, and a program. [Background technology]

[0002] Recently, there has been an increase in sales data processing devices and systems for self-service point-of-sale (POS) terminals, smartphones, cart POS systems, etc., which allow consumers to purchase products themselves without the assistance of an employee. It is known that the increase in self-service POS systems is likely to be accompanied by an increase in so-called shoplifting. Shoplifting leads to the loss of merchandise and has a particularly significant impact on small stores.

[0003] Incidentally, in recent years, advances in technologies such as artificial intelligence (AI) have made it possible to identify people's facial expressions, gestures, actions, etc. There is a demand for a technology that applies such technology to accurately detect people who are likely to commit harmful acts such as shoplifting or dangerous behavior from images of people taken inside a facility, and to alert employees and others in an easy-to-understand manner. Summary of the Invention [Problem to be solved by the invention]

[0004] The problem to be solved by the present invention is to provide an information processing device, an information processing system, and a program that are capable of detecting with high accuracy a person who is likely to commit a harmful act and notifying the person in an easy-to-understand manner. [Means for solving the problem]

[0005] An information processing device according to an embodiment includes an acquisition unit, a recognition unit, a detection unit, a generation unit, and an alarm unit. The acquisition unit acquires an image including a person from a camera installed in a facility. The recognition unit derives the second text information by inputting an image captured by the camera into a trained model trained to derive second text information describing characteristics of the person included in the image from the image based on a training dataset including an image including the person and first text information describing at least one characteristic of the person included in the image, such as facial expression, appearance, and behavior. The detection unit detects a person requiring attention who is involved in harmful behavior based on the second text information. The generation unit generates third text information including characteristics of the detected person requiring attention based on the detection result. The alarm unit notifies employees of the facility of the third text information. [Brief explanation of the drawings]

[0006] [Figure 1] FIG. 1 is a system diagram showing an example of the connection relationships of the devices in the system according to the embodiment. [Figure 2] FIG. 2 is a block diagram showing an example of the hardware configuration of the SC according to the embodiment. [Figure 3] FIG. 3 is a block diagram showing an example of the functional configuration of the SC according to the embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of a learning process of the SC according to the embodiment. [Figure 5] FIG. 5 is a flowchart showing an example of processing executed by the SC according to the embodiment. [Figure 6] FIG. 6 is a flowchart showing an example of processing executed by the SC according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0007] An information processing device, an information processing system, and a program according to an embodiment will be described in detail below with reference to Figures 1 to 6. In the embodiment described below, a store computer (SC) installed in a store (an example of a facility) such as a department store or supermarket will be described as an example of a learning device and an information processing device, but the present invention is not limited to the embodiment.

[0008] 1 is a system diagram showing an example of the connection relationships of devices in an information processing system S according to an embodiment. In FIG. 1, the system includes an SC (Store Computer) 1, a POS (Point of Sales) terminal 2, multiple cameras 3, and a mobile terminal 5.

[0009] The SC1, the POS terminal 2, and the multiple cameras 3 are connected to one another via a communication line 6 such as a LAN (Local Area Network). The mobile terminal 5 is connected to the SC1 via the communication line 6 and an access point 4, which is a wireless communication repeater.

[0010] Note that the number of each device shown in Fig. 1 is an example, and the number of each device included in the system is not limited to the number shown in Fig. 1. For example, the system may include a plurality of access points 4 and mobile terminals 5.

[0011] The POS terminal 2 is a sales data processing device that executes sales registration processing for products purchased at a store. For example, the POS terminal 2 is a dedicated self-service POS terminal, a smartphone, a cart POS, etc. Note that the POS terminal 2 is not limited to a self-service device. For example, the POS terminal 2 may be a device used by a store clerk to execute sales registration processing.

[0012] The POS terminal 2 generates product sales registration information and transmits it to the SC1 via the communication line 6. The POS terminal 2 also receives a warning message, which will be described later, from the SC1.

[0013] SC1 is a server device that collects and manages product sales registration information received from the POS terminal 2. SC1 may be configured as a single server device or as a plurality of server devices. SC1 may also be a cloud server. SC1 may also be provided on a network outside the store.

[0014] SC1 stores product information such as the prices and product names of products sold in the store. SC1 also inputs images captured by camera 3 into feature recognition model 142 (see FIG. 2), and based on text describing the characteristics of people output from feature recognition model 142, detects people who are predicted to be involved in harmful acts such as shoplifting, pickpocketing, and violent acts as suspicious persons.

[0015] Furthermore, when a suspicious person is detected, the SC1 generates a warning message to notify a store clerk (an example of an employee) that a suspicious person has been detected. The SC1 transmits the generated warning message to the POS terminal 2 and the mobile terminal 5.

[0016] The camera 3 captures images including people inside the store. For example, a plurality of cameras 3 are installed at regular intervals along the aisles, such as on the ceiling near the aisles, so that they can capture images of people passing through the store.

[0017] As an example, the multiple installed cameras 3 capture images of the path a person takes from the time they enter the store until they leave. In the example of Fig. 1, n cameras 3 are installed along the aisles in the store, and these cameras 3 capture images of person PA and person PB. The images captured by the cameras 3 include information indicating the time of capture and the capture position (for example, the installation position of the camera 3 that captured the image).

[0018] The mobile terminal 5 is carried by a store clerk and exchanges various information with the SC1. The mobile terminal 5 is, for example, a smartphone or a tablet terminal. For example, the mobile terminal 5 receives a warning message from the SC1.

[0019] Next, the hardware configuration of the SC 1 will be described with reference to Fig. 2, which is a block diagram showing an example of the hardware configuration of the SC 1.

[0020] As shown in Figure 2, SC1 includes a CPU (Central Processing Unit) 11 that serves as the control body, a ROM (Read Only Memory) 12 that stores various programs, a RAM (Random Access Memory) 13 that expands various data, and a memory unit 14 that stores various programs.

[0021] The CPU 11, ROM 12, RAM 13, and memory unit 14 are connected to one another via a data bus 15. The CPU 11, ROM 12, and RAM 13 constitute a control unit 100. That is, the control unit 100 executes various processes by the CPU 11 operating in accordance with a control program 141 stored in the ROM 12 or memory unit 14 and loaded into the RAM 13. The various processes will be described later.

[0022] The RAM 13 develops various programs including the control program 141, and also temporarily stores images captured by the camera 3 until they are stored in the memory unit .

[0023] The memory unit 14 is a non-volatile memory such as a hard disk drive (HDD) or flash memory that retains stored information even when the power is turned off, and stores programs including a control program 141. The memory unit 14 also has a feature recognition model 142, a suspicious feature list 143, and a warning generation model 144.

[0024] The feature recognition model 142 is a trained model trained using input training data including an image including a person and text describing the person's behavior (an example of first text information). The training process of the feature recognition model 142 will be described later.

[0025] As an example, the feature recognition model 142 is a model based on a machine learning model such as a neural network whose parameters are determined by deep learning. As the model, for example, a convolutional neural network (CNN) can be used, but other networks may also be used.

[0026] In this embodiment, the feature recognition model 142 is configured through learning to output feature text data (an example of second text information) describing the features of a person (e.g., facial expression, appearance, behavior, etc.) in response to an input of an image including the person captured by the camera 3.

[0027] The suspicious characteristic list 143 is a list of registered keywords that indicate the characteristics of suspicious individuals involved in harmful acts. Suspect individuals include individuals who are predicted to commit harmful acts in the future and individuals who have already committed harmful acts. For example, keywords that indicate the characteristics of suspicious individuals are character strings related to facial expressions, gestures, behavioral patterns, etc. that may suggest harmful acts.

[0028] If a keyword registered in the suspicious feature list 143 exists in the feature text data output by the feature recognition model 142, the SC1 detects a person included in the input image as a suspicious person.

[0029] For example, actions such as hiding merchandise, unnatural movements, sudden changes of direction, etc. are known to be actions that suggest shoplifting. Therefore, by registering character strings related to such actions as keywords in the suspicious characteristics list 143, SC1 can detect not only people who have actually shoplifted, but also people who may shoplift in the future as suspicious individuals.

[0030] The warning generation model 144 is a generation AI that generates sentences, such as a large language model (LLM). The warning generation model 144 generates warning text data (an example of third text information) including characteristics of a person detected as a suspicious person.

[0031] For example, the warning generation model 144 generates warning text data that notifies a store clerk of the presence of a person detected as a suspicious person in response to input of text data including characteristic text data of the person, which is constructed using known deep learning technology.

[0032] An operation unit 17 and a display unit 18 are also connected to the data bus 15 via a controller 16 .

[0033] The operation unit 17 receives various inputs from an operator such as a store clerk, etc. For example, the operation unit 17 includes a numeric keypad for entering numbers, various function keys, and the like.

[0034] The display unit 18 displays various types of information. For example, the display unit 18 displays a generated warning message. The display unit 18 may display an image of the inside of the store input from the camera 3. The display unit 18 may also display an image captured by a specific camera 3. The display unit 18 may also divide the screen and display images captured by multiple cameras 3 on the same screen at the same time.

[0035] The data bus 15 is also connected to a communication I / F 19 such as a LAN I / F (Interface). The communication I / F 19 is connected to a communication line 6.

[0036] The communication I / F 19 transmits and receives various types of information. For example, the communication I / F 19 receives images captured by the camera 3 in real time.

[0037] Next, the functional configuration of SC1 will be described. Fig. 3 is a functional block diagram showing an example of the functional configuration of SC1. By following various programs including a control program 141 stored in ROM 12 and memory unit 14, control unit 100 functions as a learning unit 101, an acquisition unit 102, a recognition unit 103, a detection unit 104, a generation unit 105, and a notification unit 106.

[0038] The learning unit 101 trains the feature recognition model 142 to learn features of people included in images, such as facial expressions, appearances, and behaviors of people. For example, the learning unit 101 collects images including people and text describing the facial expressions and behavioral features of the people included in the images. Then, the learning unit 101 generates input training data including the collected images and text.

[0039] Furthermore, the learning unit 101 generates a set of the generated input training data and feature text data (output training data) describing features of people included in the images of the input training data as a training data set. The learning unit 101 uses the generated training data set to train the feature recognition model 142 on the features of people included in the images. An example of the learning process will be described below with reference to FIG. 4.

[0040] 4 is a diagram illustrating an example of the learning process. First, the learning unit 101 collects images IA, IB, IC, etc. that include people. The learning unit 101 also collects texts TA, TB, TC, etc. that correspond to the images IA, IB, IC, etc. and describe the actions of the people included in the images IA, IB, IC, etc.

[0041] In the example of Fig. 4, image IA is an image including a person PU who is shoplifting. Furthermore, text TA corresponding to image IA is text data describing the act of shoplifting product M by person PU. The learning unit 101 generates data in which image IA and text TA are associated with each other as the first input training data.

[0042] The learning unit 101 also sets feature text data describing the facial expression, appearance, behavior, etc. of person PU included in image IA (for example, "slightly nervous facial expression, medium build male in his 50s, black shirt, black pants, shoplifting") as the first output training data (correct answer data). The learning unit 101 generates a pair of the output training data and the generated input training data as the first training data set.

[0043] Similarly, image IB is an image containing a pickpocket person PW. Text TB corresponding to image IB is text data describing the act of the person PW stealing (pickpocketing) a wallet W from another person PV. The learning unit 101 generates data in which image IB and text TB are associated with each other as second input training data.

[0044] The learning unit 101 also sets feature text data describing the facial expression, appearance, behavior, etc. of person PW included in image IB (for example, "laughing, slim, medium-height male in his 40s, wearing a blue jacket and black pants, shoplifting") as second output training data. The learning unit 101 generates a pair of the output training data and the generated input training data as a second training data set.

[0045] Similarly, image IC is an image including persons PX and PZ who are committing threatening acts. Furthermore, text TC corresponding to image IC is text data describing threatening acts by persons PX and PZ against another person PY. The learning unit 101 generates data in which image IC and text TC are associated with each other as the third input training data.

[0046] In addition, the learning unit 101 sets the third output training data as characteristic text data describing the facial expression, appearance, behavior, etc. of person PZ included in image IC (for example, "a slightly skinny, medium-sized teenage male with a mocking expression, wearing a white polo shirt and dark blue pants, threatening behavior"), and characteristic text data describing the facial expression, appearance, behavior, etc. of person PZ (for example, "a slightly overweight, tall male in his twenties with a mocking expression, wearing a red T-shirt and dark red shorts, threatening behavior")

[0047] The learning unit 101 generates a set of the output training data and the generated input training data as a third training data set. Then, the learning unit 101 causes the feature recognition model 142 to learn features of people included in images using the multiple training data sets generated as described above.

[0048] The processing performed by the learning unit 101 may be performed by an external device other than the SC1, such as an external server. In this case, the SC1 stores the feature recognition model 142 learned by the external server or the like in the memory unit 14.

[0049] The acquisition unit 102 acquires an image including a person.

[0050] For example, the acquisition unit 102 receives images from the camera 3 in real time via the communication I / F 19 and the communication line 6. Furthermore, the acquisition unit 102 acquires, as images including a person, images that are recognized as including a person using a known image recognition technology. As an example, in FIG. 1, the acquisition unit 102 acquires an image including a person PA and an image including a person PB, both of which are captured by the camera 3. Furthermore, the acquisition unit 102 sends the acquired images including the person to the recognition unit 103.

[0051] The recognition unit 103 recognizes the features of the person included in the image acquired by the acquisition unit 102.

[0052] For example, the recognition unit 103 inputs an image sent from the acquisition unit 102 to the feature recognition model 142. Then, in response to the input image, the recognition unit 103 sends feature text data output from the feature recognition model 142 to the detection unit 104 described below as a recognition result of the features of a person included in the image. At this time, the recognition unit 103 also sends information indicating the capture location of the image related to the recognition result to the detection unit 104.

[0053] For example, assume that an image of person PA in Fig. 1, a tall, medium-sized man in his twenties wearing a navy blue suit with a stern expression, making a sudden turn in front of a product shelf, is acquired by the acquisition unit 102. In this case, in response to the input of the image, the feature recognition model 142 outputs text data (feature text data) such as "a tall, medium-sized man in his twenties with a stern expression, wearing a navy blue suit, making a sudden turn in front of a product shelf."

[0054] The recognition unit 103 sends the recognition result (characteristic text data) that the characteristics of person PA are "a tall man in his twenties with a medium build, with a stern expression, wearing a navy blue suit, making a sudden turn in front of a store shelf" to the detection unit 104, together with information indicating the location where the image related to the recognition result was captured.

[0055] The detection unit 104 detects a suspicious person based on the recognition result by the recognition unit 103.

[0056] For example, the detection unit 104 refers to the suspicious feature list 143 and determines whether or not a keyword registered in the suspicious feature list 143 is present in the feature text data sent as the recognition result from the recognition unit 103. If the keyword is present, the detection unit 104 detects the person included in the image related to the recognition result as a suspicious person.

[0057] As an example, suppose that the characteristics of the recognized person PA are “a tall, medium-sized man in his twenties with a stern expression, wearing a navy blue suit, making a sudden turn in front of a product shelf” and that “sudden turn in front of a product shelf” is included in the suspicious characteristics list 143. In this case, the detection unit 104 detects the person PA as a suspicious person.

[0058] The detection unit 104 sends to the generation unit 105 information indicating that a person to be suspicious has been detected, characteristic text data of the person detected as a person to be suspicious, and information indicating the location where the image related to the recognition result sent from the recognition unit 103 was captured.

[0059] The detection unit 104 may use a known natural language processing technique or the like to determine whether the feature text data includes a character string related to a keyword registered in the suspicious feature list 143. If the character string related to the keyword is included, the detection unit 104 detects the person included in the image related to the recognition result as a suspicious person.

[0060] The generating unit 105 (an example of a generating unit and an identifying unit) generates a warning message to notify a store clerk that a suspicious person has been detected.

[0061] For example, when a suspicious person is detected by the detection unit 104, the generation unit 105 inputs the characteristic text data of the person sent from the detection unit 104 and text data (an example of fourth text information) obtained by converting information indicating the capture location of the image related to the recognition result into text, to the warning generation model 144. The generation unit 105 generates the warning text data output from the warning generation model 144 as a warning message. The generation unit 105 sends the generated warning message to the notification unit 106.

[0062] As an example, suppose that person PA is detected as a suspicious person, and the characteristics of person PA are "a tall man in his twenties with a medium build, with a stern expression, wearing a navy blue suit, making a sudden turn in front of a product shelf," and the location where the image related to the recognition result of person PA's characteristics was taken is in front of product shelf number 1.

[0063] In this case, the warning generation model 144 outputs warning text data such as, "A tall man in his twenties with a medium build and wearing a navy blue suit with a stern expression is standing near the first product shelf, engaging in suspicious behavior by making a sudden turn in front of the product shelf. Please be careful as this person may be a shoplifter." The generation unit 105 generates the warning text data as a warning message.

[0064] In this way, by generating a warning message including information about the location where the image of the suspicious person was captured, it becomes easier for the store clerk to grasp the location of the suspicious person.

[0065] The generator 105 may input text data and instruction information (prompt) to the warning generation model 144, thereby controlling the warning message output from the model according to the content specified by the instruction information.

[0066] For example, the instruction information may specify the output order of elements included in the text data (for example, to explain the location and then the characteristics of the person), or may instruct to summarize and output the contents of the text data. In this case, the warning generation model 144 generates a warning message according to the contents of the instruction information.

[0067] Furthermore, templates for generating instruction information are stored in advance in the memory unit 14 or the like. Note that a plurality of templates may be stored in the memory unit 14 or the like. In this case, the generation unit 105 may use a plurality of templates depending on the situation. For example, the generation unit 105 may switch templates depending on keywords for blacklist individuals.

[0068] Here, the process of detecting a suspicious person by the detection unit 104 is executed after the process of recognizing the characteristics of the person included in the image acquired by the acquisition unit 102 by the recognition unit 103, and before the process of generating a warning message by the generation unit 105.

[0069] The notification unit 106 notifies a store clerk that a suspicious person has been detected.

[0070] For example, the notification unit 106 controls the display unit 18 to display the warning message sent from the generation unit 105. Also, for example, the notification unit 106 transmits the warning message sent from the generation unit 105 to the POS terminal 2 via the communication I / F 19 and the communication line 6. Also, the notification unit 106 transmits the warning message to the mobile terminal 5 via the communication I / F 19, the communication line 6, and the access point 4. Also, the notification unit 106, the POS terminal 2, or the mobile terminal 5 may convert the warning message into sound and output it as sound, for example.

[0071] For example, the POS terminal 2 displays the received warning message on a display for store staff provided on the POS terminal 2, thereby alerting store staff working near the POS terminal 2 that a suspicious person has been detected.

[0072] Furthermore, for example, the mobile terminal 5 displays the received warning message on a display provided on the mobile terminal 5, thereby alerting store staff working in a location other than near the POS terminal 2, such as the back room, that a suspicious person has been detected.

[0073] Next, the processing executed by SC1 will be explained using Fig. 5 and Fig. 6. Fig. 5 and Fig. 6 are flowcharts showing an example of the processing executed by SC1. First, the learning processing will be explained. Fig. 5 shows an example of the flow of the learning processing executed by SC1.

[0074] First, the learning unit 101 collects images including people (step ST1). The collected images may be images taken by multiple cameras 3 inside the store, or may be images taken by a camera installed outside the store.

[0075] Next, the learning unit 101 collects text data describing the actions being performed by people in the acquired images (step ST2). For example, the learning unit 101 receives input of text describing the actions being performed by people in the images collected in step ST1 from the operator via the operation unit 17. The learning unit 101 collects the received text as text data.

[0076] In addition, when an image including a person and text explaining the behavior of the person in the image are published on a website, free paper, etc., the learning unit 101 may collect the image and the text as a set.

[0077] Next, the learning unit 101 generates input training data (step ST3). For example, the learning unit 101 generates data in which the images collected in step ST1 and the text data collected in step ST2 are associated with each other as the input training data.

[0078] Next, the learning unit 101 collects characteristic text data (output training data) that describes the characteristics of people in the images corresponding to the generated input training data (step ST4). For example, the learning unit 101 receives input of text that describes the facial expressions, appearances, behaviors, etc. of people in the images collected in step ST1 from the operator via the operation unit 17. The learning unit 101 collects the received text as characteristic text data.

[0079] Next, the learning unit 101 generates a learning data set (step ST5). For example, the learning unit 101 generates the learning data set as a set of the input training data generated in step ST3 and the output training data collected in step ST5.

[0080] Next, the learning unit 101 trains the feature recognition model 142 (step ST6). For example, the learning unit 101 uses the training data set generated in step ST5 to cause the feature recognition model 142 to learn features of people included in the image.

[0081] Next, the learning unit 101 determines whether the learning termination condition is satisfied (step ST7). For example, the learning unit 101 determines that the learning termination condition is satisfied when the number of learning data sets for which the learning process has been executed exceeds a threshold. If the termination condition is not satisfied (step ST7: No), the process returns to step ST1. On the other hand, if the termination condition is satisfied (step ST7: Yes), the process ends.

[0082] Next, the detection and notification process will be described. In this embodiment, the detection and notification process is a series of processes related to the detection of a suspicious person and notification of that fact. Fig. 6 shows an example of the flow of the detection and notification process executed by SC1.

[0083] First, the acquisition unit 102 acquires an image including a person (step ST21). For example, the acquisition unit 102 receives images captured by the multiple cameras 3 in real time from the multiple cameras 3 via the communication I / F 19 and the communication line 6. The acquisition unit 102 acquires an image that is recognized as including a person using a known image recognition technology as an image including a person. The acquisition unit 102 sends the acquired image to the recognition unit 103.

[0084] Next, the recognition unit 103 recognizes the features of the person in the image based on the acquired image (step ST22).

[0085] For example, the recognition unit 103 inputs the image sent from the acquisition unit 102 to the feature recognition model 142. The recognition unit 103 sends the feature text data output from the feature recognition model 142 to the detection unit 104 as the recognition result. At this time, the recognition unit 103 also sends information indicating the imaging location of the image acquired in step ST21 to the detection unit 104 together with the recognition result.

[0086] Next, the detection unit 104 determines whether a suspicious person has been detected (step ST23). For example, the detection unit 104 refers to the suspicious feature list 143, and determines that a suspicious person has been detected if a keyword registered in the suspicious feature list 143 exists in the feature text data sent as the recognition result from the recognition unit 103. If a suspicious person has not been detected (step ST23: No), the process returns to step ST21.

[0087] On the other hand, if a suspicious person is detected (step ST23: No), the detection unit 104 sends the characteristic text data and information indicating the image capture location sent to the detection unit 104 to the generation unit 105 after step ST22.

[0088] Next, the generating unit 105 generates a warning message (step S24).

[0089] For example, after step ST23, the generation unit 105 inputs the characteristic text data sent from the detection unit 104 and text data obtained by converting information indicating the image capture location into text, to the warning generation model 144. The generation unit 105 generates a warning message from the warning text data output from the warning generation model 144. The generation unit 105 sends the generated warning message to the notification unit 106.

[0090] Next, the notification unit 106 issues a warning to the store staff that a suspicious person has been detected (step ST25).

[0091] For example, after step S24, the notification unit 106 controls the display unit 18 to display the warning message sent from the generation unit 105. Also, for example, the notification unit 106 transmits the warning message to the POS terminal 2 via the communication I / F 19 and the communication line 6. Also, the notification unit 106 transmits the warning message to the mobile terminal 5 via the communication I / F 19, the communication line 6, and the access point 4.

[0092] Thereafter, the process returns to step ST21 and the same process is repeated.

[0093] As described above, the SC (learning device, information processing device) 1 according to this embodiment trains the feature recognition model 142 to derive feature text data describing features of a person from an image based on a teacher dataset including an image containing a person and text data describing the features of the person. Furthermore, the SC1 according to this embodiment detects a person who may be about to commit a harmful act such as shoplifting or a suspect person who has already committed a harmful act based on the feature text data derived by inputting images captured by a camera 3 installed in a store into the feature recognition model 142. Furthermore, when a suspect person is detected, the SC1 according to this embodiment trains the warning generation model 144 to derive warning text data from the feature text data, the warning text data describing the detection of a suspect person. Then, the SC1 according to this embodiment generates a warning message based on the warning text data derived by inputting the feature text data of the suspect person, and notifies a store clerk.

[0094] The SC1 according to this embodiment can accurately derive the characteristics of people present in a store by using a feature recognition model 142 trained with a training dataset including images containing people and text data describing the characteristics of the people. Furthermore, if the derived characteristics of the people are related to harmful activities such as shoplifting, the SC1 according to this embodiment can detect suspicious people who are involved in harmful activities. Therefore, the SC1 according to this embodiment can easily and accurately detect suspicious people by simply installing a camera 3 in a store and acquiring images captured by the camera. Furthermore, the SC1 according to this embodiment uses the warning generation model 124, which is an LLM, to generate a warning message from the feature text data describing the characteristics of the detected suspicious person, warning the store clerk that a suspicious person has been detected. Therefore, the SC1 according to this embodiment can notify the store clerk that a suspicious person has been detected in an easy-to-understand manner.

[0095] The above-described embodiment can be modified as needed by changing some of the configurations or functions of SC1. Therefore, below, several modifications of the above-described embodiment will be described as other embodiments. Below, differences from the above-described embodiment will be mainly described, and detailed descriptions of commonalities with the contents already described will be omitted. The modifications described below may be implemented individually or in appropriate combination.

[0096] (Variation 1) In the above-described embodiment, the detection unit 104 detects a suspicious person based on the feature text data output from the feature recognition model 142 and the suspicious feature list 143. In this modified example, a description will be given of a mode in which a suspicious person is detected without using the suspicious feature list 143.

[0097] In this modification, the feature recognition model 142 is configured to output only feature text data describing the features of people involved in harmful behavior. In other words, if a person in an image is engaging in behavior unrelated to harmful behavior, the feature recognition model 142 will not output feature text data.

[0098] Therefore, in this modification, when feature text data is output from the feature recognition model 142, the detection unit 104 detects a person in an image corresponding to the output feature text data as a suspicious person.

[0099] In this modification, the detection unit 104 does not need to compare the keywords registered in the suspicious feature list 143 with the character strings included in the output feature text data. In other words, this modification can reduce the processing load on the control unit 100.

[0100] (Variation 2) In the above-described first modification, the detection unit 104 detects a person to be suspicious based on the output of the feature recognition model 142. In this modification, a description will be given of a form in which a person to be suspicious is detected using a determination model that determines whether a person in an image is a person to be suspicious.

[0101] In this modification, a judgment model is stored in the memory unit 14. For example, the judgment model is configured by CNN or the like. The judgment model is a trained model that uses characteristic text data as input training data and is trained using information indicating whether or not a person related to the characteristic text data is a suspicious person as output training data (correct answer data).

[0102] For example, the determination model is configured to output information indicating whether or not a person included in an acquired image is a suspicious person in response to input of the feature text data output by the feature recognition model 142.

[0103] In this modification, when the determination model outputs information indicating that a person included in an acquired image is a person to be suspicious, the detection unit 104 detects the person in the acquired image as a person to be suspicious.

[0104] According to this modification, when it is predicted that the suspicious feature list 143 will result in an enormous number of registered keywords, it is expected that the processing time will be reduced.

[0105] Furthermore, in Modification 1, whether a person in an image is a person to be suspicious is determined based on whether or not feature text data is output, so if a process involving feature recognition model 142 is delayed for some reason, the determination may become difficult or an erroneous determination may be made. On the other hand, in this Modification, the determination model outputs information indicating whether or not the person is a person to be suspicious, so an erroneous determination caused by a processing delay does not occur.

[0106] (Variation 3) In the above-described embodiment, the recognition unit 103 recognizes features of a person in an image, such as facial expression, appearance, and behavior, using one feature recognition model 142. In this modified example, a description will be given of a configuration in which a plurality of models are used to recognize features of a person in an image, such as facial expression, appearance, and behavior.

[0107] In this modification, three models, namely, a facial expression recognition model, an appearance recognition model, and a behavior recognition model, have the same functions as the feature recognition model 142.

[0108] The facial expression recognition model is configured using CNN or the like. The facial expression recognition model is a trained model trained with a training dataset that includes images containing people and text data describing the facial expressions of the people. The facial expression recognition model is configured to output text data describing the facial expressions of the people in response to the input of an image containing a person.

[0109] The appearance recognition model is configured using CNN or the like. The appearance recognition model is a trained model trained with a training dataset that includes images containing people and text data describing the appearance characteristics of the people. The appearance recognition model is configured to output text data describing the appearance characteristics of a person in response to the input of an image containing a person.

[0110] The behavior recognition model is configured using CNN or the like. The behavior recognition model is a trained model trained with a training dataset that includes images containing people and text data describing the actions of the people. The behavior recognition model is functionally configured to output text data describing the actions of a person in response to the input of an image containing a person.

[0111] In this modification, the recognition unit 103 inputs the images acquired by the acquisition unit 102 into a facial expression recognition model, an appearance recognition model, and a behavior recognition model. At this time, the facial expression recognition model, the appearance recognition model, and the behavior recognition model each execute processing in parallel. Furthermore, the recognition unit 103 integrates the output results of the facial expression recognition model, the appearance recognition model, and the behavior recognition model, and recognizes features such as facial expressions, appearances, and behaviors of people in the images.

[0112] In the above example, the features of a person in an image are recognized using three models: a facial expression recognition model, an appearance recognition model, and a behavior recognition model. However, models other than these may be added to recognize the features of a person in an image. For example, a voice recognition model that recognizes the voice of a person in a video image captured by the camera 3, trained using known machine learning or deep learning technology, may be used to recognize the features of a person.

[0113] According to this modification, human features can be recognized using trained models trained for each feature type, such as facial features, appearance features, and behavioral features. This is expected to improve the recognition accuracy for each feature type. Furthermore, according to this modification, learning can be performed in parallel using multiple learning devices in the learning process. In this case, learning efficiency can be expected to be improved compared to when multiple features are learned simultaneously using a single learning device.

[0114] The program executed by SC1 in the above-described embodiment is provided as a file in an installable or executable format recorded on a non-transitory computer-readable recording medium such as a CD-ROM, a flexible disk (FD), a CD-R, or a DVD (Digital Versatile Disk).

[0115] The program executed by the SC1 of the above-described embodiment may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. The program executed by the SC1 of the embodiment may be provided or distributed via a network such as the Internet.

[0116] Furthermore, the program executed by the SC1 of the embodiment may be provided by being pre-installed in the ROM 12 or the like.

[0117] 3, the various functional units such as the learning unit 101, the acquisition unit 102, the recognition unit 103, the detection unit 104, the generation unit 105, and the notification unit 106 may be implemented by one or more processing circuits such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). In such a configuration, the processing circuit functioning as the detection unit 104 that detects a suspicious person based on second text information describing the person's characteristics is provided between the processing circuit functioning as the recognition unit 103 that derives the second text information and the processing circuit functioning as the generation unit 105 that generates third text information including the characteristics of the suspicious person. With such a configuration, it is possible to detect a person likely to commit a harmful act with high accuracy and to notify the person in an easy-to-understand manner.

[0118] Although the embodiments of the present invention have been described above, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, modifications, and combinations can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the inventions and their equivalents set forth in the claims. [Explanation of symbols]

[0119] 1 SC 2. POS terminals 3 Camera 4. Access Points 5. Mobile devices 6. Communication lines 11 CPU 12 ROM 13 RAM 14 Memory section 100 control section 101 Learning Department 102 Acquisition Department 103 Recognition part 104 Detector 105 Generation part 106 Information Department 142 Feature Recognition Model 143 List of features to watch out for 144 Warning Generation Model [Prior art documents] [Patent documents]

[0120] [Patent Document 1] Japanese Patent Publication No. 2023-098483

Claims

1. an acquisition unit that acquires an image including a person from a camera installed in the facility; a recognition unit that derives second text information by inputting an image captured by the camera into a trained model that has been trained to derive second text information describing characteristics of a person included in an image from the image, based on a teacher dataset that includes an image including a person and first text information describing at least one of facial expression, appearance, and behavior of the person included in the image; and a detection unit that detects a person to be suspicious who is involved in harmful acts based on the second text information; a generation unit that generates third text information including characteristics of the detected suspicious person based on the detection result; a notification unit that notifies an employee of the facility of the third text information; Equipped with Information processing device.

2. When the trained model derives the second text information including the feature suggesting the harmful act, or when the trained model derives the second text information including the feature indicating the harmful act itself, the detection unit detects a person corresponding to the feature as the suspicious person. The information processing device according to claim 1 .

3. the detection unit is provided between the recognition unit that derives the second text information describing characteristics of a person included in the image and the generation unit that generates the third text information including characteristics of the suspicious person; The information processing device according to claim 1 .

4. when the suspicious person is detected, the generation unit generates the third text information based on a character string obtained by inputting the second text information into a large-scale language model. The information processing device according to claim 1 .

5. The image acquired from the camera includes information indicating an imaging position, The image capturing device further includes an identifying unit that identifies a position where the suspicious person is captured, the generation unit generates the third text information including the location where the image of the suspect person was captured, based on a character string obtained by inputting the second text information and fourth text information describing the location where the image of the suspect person was captured into the large-scale language model. The information processing device according to claim 4 .

6. An information processing system including a camera and an information processing device, The camera is It is installed in the facility and captures images including people. The information processing device includes: an acquisition unit that acquires the image; a recognition unit that derives second text information by inputting an image captured by the camera into a trained model that has been trained to derive second text information describing characteristics of a person included in an image from the image, based on a teacher dataset that includes an image including a person and first text information describing at least one of facial expression, appearance, and behavior of the person included in the image; and a detection unit that detects a person to be suspicious who is involved in harmful acts based on the second text information; a generation unit that generates third text information including characteristics of the detected suspicious person based on the detection result; a notification unit that notifies an employee of the facility of the third text information; Equipped with Information processing system.

7. Computer, an acquisition unit that acquires an image including a person from a camera installed in the facility; a recognition unit that derives second text information by inputting an image captured by the camera into a trained model that has been trained to derive second text information describing characteristics of a person included in an image from the image, based on a teacher dataset that includes an image including a person and first text information describing at least one of facial expression, appearance, and behavior of the person included in the image; and a detection unit that detects a person to be suspicious who is involved in harmful acts based on the second text information; a generation unit that generates third text information including characteristics of the detected suspicious person based on the detection result; a notification unit that notifies an employee of the facility of the third text information; A program that functions as a

Citation Information

Patent Citations

  • Information processing system, information processing method, and program

    JP2020091812A

  • Digital autofile security system, method and program

    JP6773389B1

  • Action recognition system

    WO2025053080A1

  • Information processing program, method for processing information, and information processor

    JP2023098483A