MOTION DISCRETION PROGRAM, MOTION DISCRETION METHOD, AND MOTION DISCRETION DEVICE

The integration of object and skeletal information in a motion discrimination program and device allows for precise determination of normal and abnormal human motions in relation to objects, addressing the limitations of existing recognition technologies.

JP7680671B2Active Publication Date: 2025-05-21FUJITSU LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021095110
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-07
Publication Date
2025-05-21
Estimated Expiration
2041-06-07

AI Technical Summary

Technical Problem

Existing image recognition technologies struggle to accurately discriminate between normal and abnormal human motions in relation to objects, as they either focus solely on object recognition or human pose estimation, failing to integrate both aspects effectively.

Method used

A motion discrimination program and device that combines object and skeletal information from images to determine the normalcy of human motions, using technologies like HOID for object interaction detection and skeletal information extraction to analyze human-object interactions.

Benefits of technology

Enables accurate discrimination between normal and abnormal human motions by recognizing object positions and skeletal information, effectively identifying subtle differences in human-object interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007680671000001
    Figure 0007680671000001
  • Figure 0007680671000002
    Figure 0007680671000002
  • Figure 0007680671000003
    Figure 0007680671000003
Patent Text Reader

Abstract

To accurately determine whether an operation to an object from a person is normal or not.SOLUTION: An operation determination device 1 acquires a photographed image 2 obtained by photographing an operation of a person 3a. Then, the operation determination device 1 detects a position of an object 3b related to the person 3a from the acquired photographed image 2. Along with this operation, the operation determination device 1 detects skeleton information of the person 3a from the acquired photographed image 2. The operation determination device 1 determines whether the operation performed by the person 3a to the object 3b is normal or not on the basis of the detected position of the object 3b and the detected skeleton information. Consequently, whether the operation is normal or not can be accurately determined.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a motion discrimination program, a motion discrimination method, and a motion discrimination device. [Background technology]

[0002] Image recognition technology that recognizes a specific object from an image is widely used. In this technology, for example, the area of ​​a specific object in an image is identified as a bounding box. There is also a technology that performs image recognition of an object using machine learning. It is considered that such image recognition technology can be applied to, for example, monitoring the purchasing behavior of customers in a store and managing the work of workers in a factory. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2020-53019 A [Patent Document 2] JP 2014-132501 A Summary of the Invention [Problem to be solved by the invention]

[0004] Incidentally, when judging from an image whether a person's action against an object is normal or not, by recognizing the position and movement of the person from the image along with the position of the object, it becomes possible to distinguish fine differences between normal and abnormal movements of the person. However, the above-mentioned image recognition technology that recognizes specific objects can only recognize objects from an image, and cannot recognize people along with the objects. In addition, there is a pose estimation technology that recognizes people from an image, but this technology cannot recognize objects.

[0005] According to one aspect, the present invention provides a motion discrimination program, a motion discrimination method, and a motion discrimination device that are capable of discriminating with high accuracy whether a motion made by a person with respect to an object is normal or not. [Means for solving the problem]

[0006] In one proposal, a motion discrimination program is provided that causes a computer to execute a process of acquiring an image capturing a person's motion, detecting the positions of objects related to the person and skeletal information of the person from the acquired image, and discriminating whether the motion the person makes with respect to the object is normal or not based on the object positions and skeletal information.

[0007] Also, one proposal provides a motion discrimination method in which a computer executes a process similar to the process based on the motion discrimination program described above. Furthermore, in one proposal, a motion discrimination device is provided that executes a process similar to the process based on the motion discrimination program described above. Effect of the Invention

[0008] On the one hand, it can accurately determine whether an action a person takes toward an object is normal or not. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating a motion discrimination device according to a first embodiment. [Diagram 2] FIG. 13 illustrates an example of a configuration of a customer monitoring system according to a second embodiment. [Diagram 3] FIG. 2 illustrates an example of a hardware configuration of a monitoring device. [Figure 4] FIG. 13 is a diagram showing a comparative example of person and object recognition by HOID. [Diagram 5] 1A and 1B are diagrams illustrating examples of beverage cans and packaged goods. [Figure 6] FIG. 13 is a diagram showing an example of an image when an incorrect purchasing action is performed. [Figure 7]1A to 1C are diagrams showing an example of recognition processing using first and second comparative examples and applications thereof; [Figure 8] FIG. 2 is a diagram illustrating an example of a configuration of processing functions included in a monitoring device. [Figure 9] 11A and 11B are diagrams illustrating an example of data generated by an image feature extraction unit and a skeletal information extraction unit. [Figure 10] 13 is a diagram illustrating an example of the internal configuration of a purchasing behavior extraction unit. [Figure 11] 11 is a diagram for explaining the process of a determination unit serving as a predictor. FIG. [Figure 12] 13 is a flowchart illustrating an example of a learning process procedure of a predictor (determination unit). [Figure 13] 13 is an example of a flowchart illustrating a determination process procedure using a predictor. [Figure 14] 1 is a flowchart illustrating an example of a learning process procedure of a classifier (determination unit). [Figure 15] 13 is an example of a flowchart illustrating a determination process procedure using a classifier. [Figure 16] FIG. 13 illustrates an example of a configuration of processing functions included in a monitoring device according to a fourth embodiment. [Figure 17] 11 is a diagram for explaining a determination rule for determining whether a normal purchasing action has been performed. FIG. [Figure 18] 11A and 11B are diagrams illustrating an example of a correction process performed by an environmental difference correction unit. [Figure 19] 4 is a diagram illustrating an example of the internal configuration of a person difference correction unit. FIG. [Figure 20] 11A and 11B are diagrams for explaining a correction process by an individual difference correction unit. [Figure 21] 1 is a first example of a flowchart illustrating a procedure of a determination process based on a determination rule. [Figure 22] 13 is a second example of a flowchart illustrating a procedure of a determination process based on a determination rule. [Figure 23] 11 is an example of a flowchart showing a procedure for counting the number of scan points. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. First Embodiment Fig. 1 is a diagram showing a movement discrimination device according to a first embodiment. The movement discrimination device 1 shown in Fig. 1 is a device that discriminates whether a movement made by a person on an object is normal or not. The movement discrimination device 1 is realized as a computer such as a personal computer or a server device.

[0011] The action discrimination device 1 acquires a captured image 2 capturing an action of a person (step S1). The captured image 2 shows a person 3a performing an action on an object 3b. For example, the captured image 2 shown in Fig. 1 shows an action of the person 3a holding the object 3b in his / her hand and bringing it close to a reading device 3c, causing the reading device 3c to read the identification information added to the surface of the object 3b.

[0012] The action discrimination device 1 detects the position of an object 3b related to a person 3a from the acquired photographed image 2 (step S2a). In this detection, for example, the person 3a is recognized from the photographed image 2, and an object 3b that has an interaction with the recognized person 3a is recognized, and the positions of the person 3a and the object 3b are detected. In FIG. 1, as an example, an image area 4a containing the person 3a and an image area 4b containing the object 3b are detected. Such detection can be realized, for example, by using HOID (Human Object Interaction Detection).

[0013] HOID is a method of detecting a person and an object from a given image, and detecting the type of interaction between the person and the object. Therefore, when HOID is applied, for example, when a person interacts with an object in the captured image 2, the person, the object, and the type of interaction are detected. Then, based on the person, the object, and the type of interaction, an object (corresponding to object 3b) related to the person (corresponding to person 3a) is identified. For example, when a user holds a book, it is detected that the user, the book, and the book are being held. The interaction includes all interactions that can be recognized from an image, regardless of whether the interaction is directly caused by the user's consciousness, unconsciousness, contact, or non-contact.

[0014] At the same time, the motion discrimination device 1 detects skeletal information of the person 3a from the acquired photographed image 2 (step S2b). In this detection, for example, the positions of a plurality of predetermined joints of the person 3a are detected. In FIG. 1, as an example, the positions of the left and right shoulders, elbows, and wrists are detected as the joints of the person 3a. A skeletal line 5a indicates a line connecting the position of the right shoulder and the position of the right elbow, and a skeletal line 5b indicates a line connecting the position of the right elbow and the position of the right wrist. In addition, a skeletal line 6a indicates a line connecting the position of the left shoulder and the position of the left elbow, and a skeletal line 6b indicates a line connecting the position of the left elbow and the position of the left wrist.

[0015] Based on the position and skeletal information of the object 3b detected in this way, the action discrimination device 1 discriminates whether the action of the person 3a with respect to the object 3b is normal or not (step S3). In the example of Fig. 1, it is discriminated whether the action of the person 3a to cause the reading device 3c to read the identification information added to the object 3b is normal or not.

[0016] In the above process, it is possible to accurately determine whether the action of the person 3a with respect to the object 3b is normal or not. For example, in this embodiment, the position of an object is not simply recognized from the captured image 2, but the position of an object related to a person is recognized. This makes it possible to reliably recognize an object held by a person as in the example of FIG. 1, and even if an object unrelated to the person is photographed, it is not recognized. For example, an image recognition technology that recognizes a specific object from an image cannot extract and recognize only objects related to a person in this way.

[0017] Furthermore, in this embodiment, by using the skeletal information of the human body together with the position of the object, it becomes possible to distinguish fine differences in the movement of the person in normal and abnormal situations. For example, when the operation of making the reading device 3c read the identification information added to the object 3b is performed normally, the transition of the position of the object 3b and the transition of the positions of the joints of the person 3a are different from when the operation is not performed normally. For example, when this operation is not performed normally, the angle of the object 3b when the object 3b is brought close to the reading device 3c is different from that in normal situations. In this case, the state of the joints of the hand holding the object 3b (for example, the relative position between the joints) is also considered to be different from that in normal situations. In this embodiment, by using the skeletal information, it becomes possible to distinguish fine differences in such a person's movement.

[0018] Second Embodiment Next, a case where the processing of the movement determining device 1 is applied to a customer monitoring system in a store will be described as a second embodiment.

[0019] Fig. 2 is a diagram showing an example of the configuration of a customer monitoring system according to a second embodiment. The customer monitoring system shown in Fig. 2 is a system for monitoring the purchasing behavior of customers in a store where products are sold, and includes a monitoring device 100 and a camera 101 connected to the monitoring device 100. The monitoring device 100 is an example of the behavior discrimination device 1 shown in Fig. 1.

[0020] The camera 101 is installed in a store in which a cash register 50 is installed. The cash register 50 is a POS terminal included in a POS (Point Of Sale) system. The cash register 50 is a self-service type cash register in which the customer himself performs the checkout operation, and is sometimes called a "self-checkout."

[0021] Cash register 50 comprises a barcode scanner 51, a display 52, and a deposit / withdrawal unit 53. Barcode scanner 51 reads barcodes attached to products that indicate product codes. Display 52 displays the price of the product whose barcode has been read, the total price of the products to be purchased, the amount of change, etc. Deposit / withdrawal unit 53 accepts deposits from customers and dispenses change.

[0022] For example, a customer approaches cash register 50 with a store basket containing the items they wish to purchase, and performs a "scanning operation" by bringing each item in the basket close to barcode scanner 51 in order to have its barcode read. When the customer has finished scanning all the items, they perform a "settlement operation" to request payment. For example, if display 52 is a touch panel, the customer can perform the settlement operation by pressing a settlement button on the touch panel. After performing the settlement operation, the customer deposits the purchase amount into cash deposit / withdrawal unit 53 according to the information displayed on display 52, and receives any change from cash deposit / withdrawal unit 53.

[0023] The camera 101 captures an image of the front of the cash register 50 (particularly, the area around the barcode scanner 51) so as to capture the purchasing behavior of the customer using the cash register 50. The monitoring device 100 determines whether the customer has performed a correct purchasing behavior from the image captured by the camera 101, and can issue a warning if it is determined that the customer has performed an abnormal purchasing behavior.

[0024] Fig. 3 is a diagram showing an example of the hardware configuration of a monitoring device. The monitoring device 100 is realized, for example, as a computer as shown in Fig. 3. The monitoring device 100 shown in Fig. 3 has a processor 111, a random access memory (RAM) 112, a hard disk drive (HDD) 113, a graphics processing unit (GPU) 114, an input interface (I / F) 115, a reading device 116, a network interface (I / F) 117, and a communication interface (I / F) 118.

[0025] The processor 111 performs overall control of the entire monitoring device 100. The processor 111 is, for example, a central processing unit (CPU), a micro processing unit (MPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a programmable logic device (PLD). The processor 111 may also be a combination of two or more elements of a CPU, an MPU, a DSP, an ASIC, or a PLD.

[0026] The RAM 112 is used as a main storage device of the monitoring device 100. The RAM 112 temporarily stores at least a part of an OS (Operating System) program and application programs to be executed by the processor 111. The RAM 112 also stores various data necessary for processing by the processor 111.

[0027] The HDD 113 is used as an auxiliary storage device of the monitoring device 100. An OS program, an application program, and various data are stored in the HDD 113. Note that other types of non-volatile storage devices such as SSDs (Solid State Drives) can also be used as the auxiliary storage device.

[0028] A display device 114a is connected to the GPU 114. The GPU 114 displays an image on the display device 114a in accordance with an instruction from the processor 111. The display device may be a liquid crystal display or an organic EL (ElectroLuminescence) display.

[0029] An input device 115a is connected to the input interface 115. The input interface 115 transmits a signal output from the input device 115a to the processor 111. The input device 115a may be a keyboard or a pointing device. The pointing device may be a mouse, a touch panel, a tablet, a touch pad, a trackball, or the like.

[0030] Portable recording medium 116a is detachably attached to reader 116. Reader 116 reads data recorded on portable recording medium 116a and transmits it to processor 111. Portable recording medium 116a may be an optical disk, a magneto-optical disk, a semiconductor memory, or the like.

[0031] A network interface 117 transmits and receives data to and from other devices via a network 117a. The communication interface 118 transmits and receives data to and from the camera 101 .

[0032] The processing functions of the monitoring device 100 can be realized by the above hardware configuration. Meanwhile, self-service cash registers are becoming increasingly popular to address labor shortages caused by population decline, to relieve congestion, and to prevent virus infection. However, there is a problem with self-service cash registers, where customers may make incorrect purchase actions, resulting in unpaid bills. Such incorrect purchase actions may be accidental or intentional, and there are various types of actions, as follows:

[0033] Examples of negligent purchasing errors include "scan failure" where a customer forgets to scan an item and moves the item directly from the in-store basket to the customer's bag, or "basket failure" where a customer forgets to scan an item in the lower basket when in-store baskets can be placed above and below the shopping cart.

[0034] On the other hand, deliberately incorrect purchasing behaviors include "barcode hiding," where a customer pretends to scan the barcode while covering it with their finger. There is also "barcode scanning error," where, when multiple identical products are in a package, a customer scans the barcode of one product exposed from the package instead of the barcode on the package.

[0035] Since the appearance of the product and the body movements of these erroneous purchasing actions differ depending on the type, it can be said that it is technically difficult to automatically recognize all types of actions as erroneous purchasing actions. Therefore, for example, a method of recognizing purchasing actions using multiple types of sensors or multiple sensors of the same type can be considered. However, since the introduction cost of the equipment for this method is high, it is desirable to be able to recognize using only a single camera.

[0036] Here, technology for recognizing a specific product from an image captured by a camera is widely used. With this technology, for example, the area of ​​the product in the image (for example, a bounding box) is identified. There is also technology that performs such image recognition of products using machine learning. However, with these technologies, templates and learning data must be prepared in advance for each product. For this reason, it is not realistic to use these technologies in a store where there are many products and new products are frequently introduced. Furthermore, this technology only identifies the image area of ​​a specific product, and is not able to recognize the relationship with a person.

[0037] On the other hand, image data of people is easy to collect, and people tend to be large in size in images. For this reason, people can be more easily recognized than non-human objects, for example, by machine learning using a neural network (NN). Therefore, the difficulty of recognizing objects related to people can be reduced by first recognizing people from an image, and then recognizing objects by focusing on the positions of the person's body parts (for example, the position of the hands). HOID is a technology that uses this approach to recognize people and objects related to them.

[0038] Figure 4 shows a comparative example of person and object recognition by HOID. In HOID, objects are recognized using information about people in an image. Of the objects in the image, only those that have interactions with people are recognized, and objects that do not have interactions are ignored.

[0039] 4 shows an image 200 of a person 201 holding some kind of object 202. When this image 200 is input into a learning model trained by HOID, for example, the position information of the person, the position information of the object that is interacting with the person, the class name of the object, the class name of the interaction, and the confidence score for the class name are output.

[0040] Position information of a person and an object is output, for example, as a bounding box indicating a rectangular area circumscribing those areas. In the image 200a illustrated in FIG. 4, a bounding box 203 indicating the position of a person 201 and a bounding box 204 indicating the position of an object 202 are detected from the image 200. In this example, a class name indicating "has" is output as the class name of the interaction. The confidence score may actually represent a confidence score (probability value) that when there is a person area and an object area, the object belongs to a certain object class name and there is a relationship of a certain interaction class name between the object and the person.

[0041] By applying this HOID, it becomes possible to recognize products related to people from images without pre-registering the product image, as long as the product does not blend in with the background or hands. In other words, any object in an image can be recognized as "Something" (some object other than a person, which may or may not be a pre-registered object) and a bounding box indicating the area of ​​that object can be estimated. Hereinafter, this type of person and object recognition processing will be referred to as "First Comparative Example."

[0042] On the other hand, posture estimation technology is a technology for recognizing a person from an image. The posture estimation technology is a technology for estimating the posture of a human body by detecting skeletal information of the human body. As the skeletal information, for example, the positions of a number of predetermined joints contained in the human body in an image are detected. Hereinafter, the process of recognizing a person by detecting the skeleton is referred to as a "second comparative example."

[0043] Here, an example of an erroneous purchasing behavior using a self-service cash register 50 is shown, and a process for recognizing this purchasing behavior is shown using first and second comparative examples. Here, the above-mentioned "barcode scan error" is shown as an example of an erroneous purchasing behavior. Also, a beverage can and its packaged product are shown as examples of products to be purchased.

[0044] Fig. 5 is a diagram showing examples of a beverage can and a packaged product. Beverage can 211 shown in Fig. 5 is a can containing a beverage such as beer. A barcode 212 indicating a product code is attached to the side of beverage can 211. Packaged product 213 shown in Fig. 5 is a product in which multiple beverage cans 211 are sold as a set, and multiple beverage cans 211 are stored inside an exterior packaging. A barcode 214 indicating the product code of packaged product 213 is attached to the exterior packaging.

[0045] Many of these packaged goods 213 have openings on the exterior, such as at both ends, and the beverage can 211 inside is partially exposed through the openings. For this reason, as shown in Fig. 5, a customer may intentionally cause a barcode 215 attached to the beverage can 211 exposed from the exterior of the packaged goods 213 to be scanned by a barcode scanner 51 of a cash register 50. In this case, multiple beverage cans are fraudulently purchased for the price of one can.

[0046] For example, one method for detecting fraudulent purchasing behavior is to compare the number of scanned items with the number of items taken out. However, when a purchasing behavior resulting from such a "barcode scan error" is performed, the number of scanned items will match the number of items taken out, making it difficult to detect fraudulent behavior.

[0047] Fig. 6 is a diagram showing an example of an image when an incorrect purchasing action is performed. The images 221 and 222 shown in Fig. 6 show a situation when a purchasing action of "wrong barcode scan" is performed for the package product 213 as described above.

[0048] Image 221 shows a state where a customer takes out a packaged product 213 from an in-store basket 216 in his / her hand. Image 222 shows a state where the customer then brings the packaged product 213 held in his / her hand close to the barcode scanner 51 to perform a scanning operation. However, in image 222, the customer is trying to have the barcode scanner 51 scan the barcode attached to the beverage can inside the packaged product 213, not the barcode attached to the packaged product 213. When performing such a scanning operation, the customer rotates the packaged product 213 held in his / her hand and moves it to a position where the barcode of the beverage can can be scanned, which is different from a normal purchasing operation. For example, in order to rotate the packaged product 213, the arm, fingers (thumb), etc. make a unique movement that is different from a normal one.

[0049] Therefore, if such actions by customers can be recognized from images as at least actions that are different from normal, it will be possible to issue a warning when such actions are performed. However, as shown in the following Figure 7, when using the above first comparative example using HOID or the above second comparative example using skeletal detection, it is difficult to recognize the above actions as incorrect purchasing actions.

[0050] FIG. 7 is a diagram showing an example of recognition processing using the first and second comparative examples and its application. Images 221a and 222a shown at the top of Fig. 7 show the case where the position of a product is recognized by applying the first comparative example based on images 221 and 222 in Fig. 6. In this case, a bounding box (not shown) indicating the area of ​​a person is first detected by HOID, and then a bounding box 217 indicating the area of ​​a product that has a correlation with this person is detected. In reality, the product is recognized as some kind of object (Something) held by the person in his / her hand, but since the person is holding it in front of the cash register 50, it is possible to identify the object as a product.

[0051] In this way, when the first comparative example using HOID is applied, the position of the product can be detected from the captured image. Therefore, it is possible to recognize that the customer has brought the product closer to the barcode scanner 51. However, since the customer's body movements cannot be detected, it is not possible to recognize the specific body movements (e.g., finger movements) made when rotating the packaged product 213. Therefore, it is not possible to distinguish the above-mentioned customer purchasing behavior from normal purchasing behavior.

[0052] On the other hand, images 221b and 222b shown at the bottom of Fig. 7 show a case where the movement of a person is recognized by applying the second comparative example based on images 221 and 222 of Fig. 6. In this case, the positions of the person's joints are detected by skeleton detection. In images 221b and 222b, the lines connecting the wrist and thumb joints are represented by thick lines 218a to 218c.

[0053] In this way, when the second comparative example using skeletal detection is applied, it is possible to detect from the captured image the specific body movements (e.g., finger movements) made when rotating a product for an unauthorized scanning operation. However, since the position of the product cannot be detected, the detected body movements cannot be associated with the product and recognized as "movements while holding the product."

[0054] Therefore, the monitoring device 100 of this embodiment is able to detect both the position of the product and the specific body movement by combining the technology of the first comparative example using HOID and the technology of the second comparative example using skeletal detection. This allows the monitoring device 100 to recognize the product held by the customer in front of the cash register 50 without having to register images of many products in advance. The monitoring device 100 can accurately recognize the movement trajectory of the product and how it appears in the image, as well as the body movement of the customer holding the product and performing a scan operation. As a result, the monitoring device 100 can distinguish various erroneous purchasing actions, including the above-mentioned "barcode scan error," from normal purchasing actions.

[0055] 8 is a diagram showing an example of the configuration of the processing functions of the monitoring device. As shown in FIG. 8, the monitoring device 100 includes an image acquisition unit 121, an image feature extraction unit 122, a skeletal information extraction unit 123, a purchasing behavior extraction unit 124, a learning unit 125, a determination unit 126, an image storage unit 131, and a learning model storage unit 132. In the monitoring device 100, the processes of the image acquisition unit 121, the image feature extraction unit 122, the skeletal information extraction unit 123, the purchasing behavior extraction unit 124, the learning unit 125, and the determination unit 126 are realized, for example, by the processor 111 included in the monitoring device 100 executing a predetermined program. The image storage unit 131 and the learning model storage unit 132 are realized by the storage area of ​​a storage device included in the monitoring device 100, such as the RAM 112 and the HDD 113.

[0056] Image acquisition unit 121 acquires data of a moving image captured by camera 101. During learning, the acquired data of the moving image is stored in image storage unit 131 as learning data, and then read from image storage unit 131 and input to image feature extraction unit 122 and skeletal information extraction unit 123. On the other hand, during judgment of a purchasing behavior, the acquired data of the moving image is input to image feature extraction unit 122 and skeletal information extraction unit 123 in sequence.

[0057] The image feature extraction unit 122 inputs the input video into a learning model (here, NN) of HOID to extract information on the person and information on the product that interacts with the person. Specifically, the appearance information of the product and the position information of the person and the product are extracted.

[0058] The skeleton information extraction unit 123 detects skeletons from the input video image and extracts human skeleton information. The purchasing behavior extraction unit 124 generates a feature quantity indicating a purchasing behavior based on the information extracted from the video by the image feature extraction unit 122 and the skeletal information extraction unit 123.

[0059] The learning unit 125 uses the feature amounts extracted by the purchasing behavior extraction unit 124 from a large number of video images as learning data to learn a classifier for discriminating between normal purchasing behavior and abnormal purchasing behavior. In this embodiment, deep learning using a NN is executed. The learning unit 125 stores data indicating a learning model (NN) obtained by learning in the learning model storage unit 132.

[0060] The determination unit 126 operates as a classifier based on the data of the learning model stored in the learning model storage unit 132. The determination unit 126 distinguishes between normal and abnormal purchasing behaviors using the feature amount extracted by the purchasing behavior extraction unit 124 from the video captured by the camera 101, and issues a warning when an abnormal purchasing behavior is detected.

[0061] In this embodiment, video images obtained by recording actual purchasing behavior are used as learning data without being labeled. Among such videos, the majority are videos of normal purchasing behavior, and a very small number are videos of abnormal purchasing behavior. However, ideally, only videos of normal purchasing behavior may be used. The learning unit 125 is an example of a classifier that learns a predictor that predicts the feature amount at the next time from the feature amount at a certain time based on such videos. Therefore, the determination unit 126 operates as such a predictor.

[0062] 9 is a diagram showing an example of data generated by the image feature extraction unit and the skeletal information extraction unit. The image feature extraction unit 122 and the skeletal information extraction unit 123 generate the following data every time frame data is input.

[0063] The image feature extraction unit 122 outputs HOID information by inputting the frame into a learning model (NN) of HOID. The HOID information includes information indicating an interaction between a person and an object recognized from the frame. The information indicating the interaction includes, for example, a person ID that identifies a person, an object ID that identifies an object that interacts with the person, and an action ID that indicates the type of action of the person with respect to the object.

[0064] In this embodiment, the image feature extraction unit 122 outputs HOID information including an action ID indicating "holding an object" as an action type. For example, when multiple combinations of people and objects that interact with each other are extracted from a captured image, only the HOID information for the combination whose action ID indicates "holding an object" is output.

[0065] The HOID information further includes information about a person and information about an object. The information about a person includes a person ID and position information of a person area. The information about an object includes an object ID, position information of an object area, and an object type ID indicating a type of the object. Note that, for example, information indicating the positions of bounding boxes indicating the person area and the object area, respectively (for example, coordinates of the four corners) is used as the position information of the person area and the object area.

[0066] The image feature extraction unit 122 generates appearance information and position information about the object from the above HOID information, and outputs the information to the purchasing behavior extraction unit 124. The appearance information about the object includes an image obtained by cutting out an object area from a frame. This appearance information may include, for example, information indicating the color, shape, and size of the object. The position information about the object includes position information about the object area (bounding box). This position information includes, for example, the coordinates of the four corners of the object area. Furthermore, the appearance information may further include information indicating the relative positional relationship between the object area and the person area.

[0067] In practice, it is desirable to perform a supplementary process so that the HOID accurately recognizes the commodity as an object. For example, a ROI (Region Of Interest) is set at a predetermined position in the captured image, and the commodity is accurately recognized as an object based on the positional relationship between the recognized object and the ROI. As an example, an area in front of the cash register 50 where the commodity is located when it is taken out of the in-store basket, or an area where the commodity is temporarily placed, is set as the ROI. When the position of an object newly recognized by the HOID is within the ROI, the image feature extraction unit 122 recognizes the object as a commodity, whereas when the position of the newly recognized object is within a bounding box of a person, the image feature extraction unit 122 recognizes the object as not a commodity (for example, a customer's personal property such as a wallet). The image feature extraction unit 122 may also distinguish the commodity as an object based on an image change (such as background difference) within the ROI before and after the object enters the ROI.

[0068] Meanwhile, the skeleton information extraction unit 123 generates skeleton information of the person from the frames and outputs it to the purchase behavior extraction unit 124. The skeleton information includes, for each detected joint point, a combination of a joint point ID that identifies the joint point (the center point of the joint) and position information (coordinates) of the joint point. In addition, the skeleton information may include a reliability score of the joint point as information for each joint point.

[0069] Note that the skeletal information of a person is associated with a person or object detected by the image feature extraction unit 122, for example, based on a comparison result between the position of each joint point and the position of the person area or object area detected by the image feature extraction unit 122. When joint points (e.g., a predetermined number or more of joint points) detected by the skeletal information extraction unit 123 are included within the bounding box of a person detected by the image feature extraction unit 122, the person corresponding to the bounding box and the person corresponding to the joint points are determined to be the same person. In this case, the person ID of the person corresponding to the bounding box is added to the data of the joint points.

[0070] 10 is a diagram showing an example of the internal configuration of the purchasing behavior extraction unit 124. As shown in FIG.

[0071] Environmental difference correction unit 141 corrects the appearance information and position information of an object input from image feature extraction unit 122 and the skeletal information of a person input from skeletal information extraction unit 123 in accordance with the shooting environment of camera 101. For example, based on the relative distance between camera 101 and cash register 50 and the number of pixels in the captured image, the appearance information and position information of the object and the skeletal information of the person are corrected so that the distance at a reference position in the shooting space (for example, a predetermined point on the surface of cash register 50) and the number of pixels on the image are always constant.

[0072] The feature amount calculation unit 142 calculates feature amount vectors from each of the object appearance information, object position information, and person skeletal information, and integrates these feature amount vectors into one feature amount vector. This feature amount calculation unit 142 includes an appearance feature extraction unit 151, an object position feature extraction unit 152, a skeletal feature extraction unit 153, and a feature amount vector integration unit 154.

[0073] The appearance feature extraction unit 151 converts the appearance information of the corrected object into a feature vector and calculates the appearance feature vector. For example, the feature vector is generated from the appearance information (a partial image cut out from the captured image) by the ROI Align method.

[0074] The object position feature extraction unit 152 converts the corrected object position information into a feature vector to calculate an object position feature vector. For example, at least one of the following is vectorized: the center coordinates, width, height, size, aspect ratio, ratio of the size of the object region to the human body region, overlap rate (Intersection over Union: IoU) between the human body region and the object region, distance between the center of the human body region and the center of the object region, and relative coordinates between the object region and a reference position in the shooting space (for example, a stand on which a commodity is temporarily placed next to the cash register 50). Then, the generated vector is linearly transformed to generate an object position feature vector.

[0075] The skeleton feature extraction unit 153 converts the corrected skeleton information of the person into a feature vector and calculates a skeleton position feature vector. For example, at least one of the coordinates of each joint point, the reliability score of each joint point, the relative coordinate between a certain joint point and another joint point, the relative coordinate between each joint point and a reference position in the shooting space, and the relative coordinate between each joint point and the center of the object area is vectorized. Then, the generated vector is linearly transformed to generate a skeleton position feature vector.

[0076] The feature vector integration unit 154 integrates the calculated appearance feature vector, object position feature vector, and skeleton position feature vector to calculate an integrated feature vector. In this vector integration, each feature vector is integrated for each joint point. That is, the skeleton position feature vector, object position feature vector, and appearance feature vector for one joint point are integrated to calculate an integrated feature vector for that joint point, and such an integrated feature vector is calculated for each joint point. This allows the feature vector to be generated as learning data that can accurately learn the movement of each joint point.

[0077] The feature vector integration unit 154 integrates the feature vectors of multiple dimensions into a feature vector of the total number of dimensions by, for example, concatenation, and calculates an integrated feature vector by performing linear transformation or nonlinear transformation. For example, the appearance feature vector is expressed as x 1 , the object position feature vector is x 2 , the skeleton position feature vector of joint point i is y i Then, the integrated feature vector z i is expressed by the following equation (1). z i =W 2 *σ(W 1 *[x 1 ,x 2 ,y i ]) ···(1) In addition, in formula (1), W 1 ,W 2 represents a predetermined weighting coefficient, [] represents concatenation, σ represents a nonlinear transformation, and * represents an inner product.

[0078] The integrated feature vectors for each joint point generated by the above procedure may be input directly to the learning unit 125 and the determination unit 126. On the other hand, in this embodiment, these integrated feature vectors are processed by the time-series information processing unit 143 and then input to the learning unit 125 and the determination unit 126. The time-series information processing unit 143 processes the generated integrated feature vectors for each joint point according to temporal continuity to generate a feature vector representing the entire purchasing behavior. This process is executed using, for example, ST-GCN (Spatial Temporal-Graph Convolutional Networks). In ST-GCN, both spatial patterns and temporal patterns based on the positions of the joint points are learned from the integrated feature vectors for each joint point.

[0079] Next, the learning unit 125 and the determination unit 126 will be described. As described above, in this embodiment, the determination unit 126 operates as a predictor that predicts a feature amount at the next time from a feature amount at a certain time. The learning unit 125 learns such a predictor.

[0080] Fig. 11 is a diagram for explaining the processing of the determination unit as a predictor. The feature space 231 shown in Fig. 11 is a simplified representation of the coordinate space having each dimension of the integrated feature vector output from the purchasing behavior extraction unit 124, as a two-dimensional coordinate space. Here, the frame period is taken as the unit time, and the change in feature between a certain time T and a time (T+1) is considered. When the frame at time T is the current frame, the frame at time (T+1) is the next frame.

[0081] The learning unit 125 learns a predictor that predicts the feature amount from time T to time (T+1) based on an integrated feature amount vector generated from a video image obtained by recording an actual purchasing behavior. Such learning is performed, for example, by using an RNN (Recurrent Neural Network) as a learning model.

[0082] Here, a label indicating normality or abnormality does not need to be added to the integrated feature vector as learning data input to the learning unit 125. However, the majority of such videos are videos when normal purchasing behavior is performed, and videos when abnormal purchasing behavior is performed are very few. For this reason, the learning unit 125 learns a predictor that predicts features when normal purchasing behavior is performed.

[0083] The predictor trained in this way predicts the feature values ​​at time (T+1) when a normal purchasing behavior is performed when the feature values ​​at time T are input. For example, if a product or finger joints are in a certain position at time T, the predictor can predict to which position the product or finger joints will move at time (T+1) if the purchasing behavior is normal. Conversely, when the feature values ​​at time T are input, the predictor cannot predict the feature values ​​at time (T+1) when an abnormal purchasing behavior is performed.

[0084] For example, in the feature space 231 shown in FIG. 11, the feature F T When the input is input to the predictor, the predictor selects the feature F' at time (T+1) from the set of normal behavior features. T+1 However, the input feature F T If is the feature value when abnormal purchasing behavior occurs, the feature value F' at the predicted time (T+1) is T+1 and the actual feature value F at time (T+1) T+1 Therefore, the distance D between the predicted feature quantity F' in the feature quantity space 231 increases. T+1 and the actual feature F T+1 When the spatial distance D between the shopping cart and the shopping cart is greater than a predetermined threshold, it can be determined that an abnormal purchasing behavior has occurred.

[0085] This type of processing makes it possible to determine that an abnormal purchasing behavior has occurred when a customer performs an action that differs from the normal movement of a product or joints, such as the aforementioned "barcode scan error."

[0086] Next, FIG. 12 is an example of a flowchart showing a learning process procedure of the predictor (determining unit). [Step S11] The image acquisition unit 121 collects video images of purchasing behavior captured by the camera 101. These videos include a relatively large number of videos of normal purchasing behavior, and a small number of videos of abnormal purchasing behavior. However, only videos of normal purchasing behavior may be collected. The image acquisition unit 121 stores data of each collected video in the image storage unit 131.

[0087] [Step S12] For each video sequence stored in the image storage unit 131, the feature generation loop process up to step S18 is executed. [Step S13] A frame processing loop up to step S17 is executed for each frame included in the video.

[0088] [Step S14] The image feature extraction unit 122 inputs the video data into the HOID learning model (NN) to calculate HOID information. Based on the HOID information, the image feature extraction unit 122 extracts object appearance information and position information as feature quantities. In addition, the skeletal information extraction unit 123 detects skeletons from the video and extracts human skeletal information as feature quantities.

[0089] [Step S15] The environmental difference correction unit 141 of the purchasing behavior extraction unit 124 corrects the extracted feature amounts (appearance information, position information, and skeletal information) in accordance with the image capture environment. [Step S16] A feature vector is calculated based on the corrected feature.

[0090] Specifically, the appearance feature extraction unit 151 calculates an appearance feature vector based on the appearance information of an object. Furthermore, the object position feature extraction unit 152 calculates an object position feature vector based on the object position information. Furthermore, the skeleton feature extraction unit 153 calculates a skeleton position feature vector based on the person's skeleton information. Then, the feature vector integration unit 154 integrates the appearance feature vector, the object position feature vector, and the skeleton position feature vector to calculate an integrated feature vector for each joint point. The time series information processing unit 143 processes the integrated feature vector for each joint point according to the temporal continuity to generate a feature vector representing the entire purchasing behavior.

[0091] [Step S17] When the processes of steps S14 to S16 have been performed on all frames included in the video, the frame processing loop ends and the process proceeds to step S18.

[0092] [Step S18] When the processes of steps S13 to S17 have been executed for all the moving images stored in the image storage unit 131, the feature generation loop ends and the process proceeds to step S19.

[0093] [Step S19] The learning unit 125 learns a predictor that predicts the feature quantity at time (T+1) from the feature quantity at time T based on the feature quantity vector that represents the entire purchasing behavior. Data of the learning model (NN) generated by this learning is stored in the learning model storage unit 132.

[0094] FIG. 13 is an example of a flowchart illustrating a determination process procedure using a predictor. [Step S21] The image acquisition unit 121 acquires frames of a moving image captured by the camera 101.

[0095] [Step S22] The image feature extraction unit 122 inputs the frame data into the HOID learning model (NN) to calculate HOID information. This frame is the frame acquired in the previous step S21 or step S29. The image feature extraction unit 122 extracts appearance information and position information of an object as features based on the HOID information. In addition, the skeletal information extraction unit 123 detects skeletons from the video and extracts human skeletal information as features.

[0096] [Step S23] The environmental difference correction unit 141 of the purchasing behavior extraction unit 124 corrects the extracted feature amounts (appearance information, position information, and skeletal information) in accordance with the image capture environment. [Step S24] A feature vector is calculated based on the corrected feature amounts.

[0097] Specifically, the appearance feature extraction unit 151 calculates an appearance feature vector based on the appearance information of an object. Furthermore, the object position feature extraction unit 152 calculates an object position feature vector based on the object position information. Furthermore, the skeleton feature extraction unit 153 calculates a skeleton position feature vector based on the person's skeleton information. Then, the feature vector integration unit 154 integrates the appearance feature vector, the object position feature vector, and the skeleton position feature vector to calculate an integrated feature vector for each joint point. The time series information processing unit 143 processes the integrated feature vector for each joint point according to the temporal continuity to generate a feature vector representing the entire purchasing behavior.

[0098] [Step S25] The determination unit 126 inputs the generated feature vector to a predictor based on the learning model data stored in the learning model storage unit 132 to predict the feature vector of the next frame. The determination unit 126 temporarily stores the prediction result of the feature vector of the next frame in the RAM 112.

[0099] [Step S26] If the prediction result of the feature vector of the current frame, which is predicted based on the feature vector of the previous frame, is stored in RAM 112, the processes of steps S26 to S28 are executed. This prediction result is the one stored in RAM 112 by the process of step S25 for the previous frame. The determination unit 126 acquires this prediction result from RAM 112, and calculates the distance between the feature vector indicated by this prediction result and the feature vector calculated from the current frame in step S24.

[0100] [Step S27] The determination unit 126 compares the calculated distance with a predetermined threshold value. If the distance exceeds the threshold value, the process proceeds to step S28. If the distance is equal to or smaller than the threshold value, the process proceeds to step S29.

[0101] [Step S28] If the distance exceeds the threshold, it is determined that an abnormal purchasing behavior has occurred. The determining unit 126 executes a process to issue a warning that an abnormal purchasing behavior has occurred. For example, the determination unit 126 may cause the display device 114a to display image information indicating that an abnormal purchasing behavior has occurred. When the monitoring device 100 is capable of communicating with the cash register 50, the determination unit 126 may cause the display 52 of the cash register 50 to display such image information.

[0102] Also, a warning may be given by voice. For example, the determination unit 126 causes a speaker connected to the monitoring device 100 to output a warning voice that warns that an abnormal purchasing behavior has occurred. Also, when a store clerk is wearing earphones that can hear voice via wireless communication, the determination unit 126 may transmit voice information that warns that an abnormal purchasing behavior has occurred and cause the earphones to output the voice to notify the store clerk of the occurrence of the abnormality.

[0103] <Step S29> The image acquisition unit 121 acquires the next frame captured by the camera 101. [Step S30] The image feature extraction unit 122 or the skeletal information extraction unit 123 determines whether the same person as in the previous frame is detected. If the same person is detected, the process proceeds to step S22. On the other hand, if the same person is not detected (if the person has moved outside the shooting area), the determination process in FIG. 13 ends.

[0104] In the second embodiment described above, the determination unit 126 can be realized to determine with high accuracy whether or not a purchasing action has been performed correctly by learning based on the detection results of products by HOID and the detection results of human skeleton information. In particular, it becomes possible to determine that an illegal purchasing action, such as the above-mentioned barcode scan error where the scanning operation of the barcode of a product is actually performed, is not a correct purchasing action.

[0105] Third embodiment In the third embodiment, a part of the processing of the monitoring device 100 in the second embodiment is modified. In the second embodiment, the determination unit 126 operates as a predictor that predicts a feature amount at time (T+1) from a feature amount at time T, and determines whether an abnormal action has occurred from the difference between the predicted value of the feature amount at time (T+1) and the actual feature amount. In contrast, in the third embodiment, video images to which a label indicating whether the action is normal or abnormal is added are used as learning data, and a classifier that explicitly distinguishes between normal and abnormal actions is trained. Then, the processing of the determination unit 126 is modified so that it operates as such a classifier.

[0106] Fig. 14 is an example of a flowchart showing the learning process procedure of a classifier (determination unit). In Fig. 14, the same step numbers are assigned to the processing steps with the same processing contents as in Fig. 12. In the learning process of the classifier shown in Fig. 14, step S41 is executed between step S11 and step S12 in Fig. 12, and step S42 is executed instead of step S19 in Fig. 12.

[0107] [Step S41] Annotation is performed on each video collected in step S11 and stored in the image storage unit 131. For example, when the classifier distinguishes between normal and abnormal actions, either a normal label indicating normal action or an abnormal label indicating abnormal action is added to the data of the video. Preferably, the abnormal label is added only to frames in the video during a period in which an abnormal purchasing action is being performed, and the normal label is added to frames in other periods. For example, in a video showing the action of "barcode scan error" for the packaged product 213 described in FIG. 6 and FIG. 7, the abnormal label may be added to frames in a period from when the customer starts to rotate the packaged product 213 in the process of holding the packaged product 213 and bringing it close to the barcode scanner 51 until the barcode of the beverage can is scanned.

[0108] In addition, if the classifier further identifies the type of abnormal behavior, a label for each type of abnormal behavior is added to the video as an abnormal label. For example, to identify the four types of abnormal behavior mentioned above, "missed scan," "missing basket," "hiding the barcode," and "bad barcode scan," four types of abnormal labels are used. In this case, one of five labels, namely a normal label and one of the four types of abnormal labels, is added to the frames of the video.

[0109] In step S11, unlike the case of Fig. 12, it is desirable to collect a certain number of videos when abnormal purchasing behavior is performed. In particular, when making it possible to identify the type of abnormal behavior as described above, it is desirable to collect a certain number of videos when the corresponding type of abnormal behavior is performed for each type of abnormal behavior.

[0110] [Step S42] The feature vectors of each video calculated by the purchasing behavior extraction unit 124 with the above-mentioned labels added are input to the learning unit 125. The learning unit 125 trains a classifier that identifies purchasing behaviors based on the input feature vectors and labels. Data of the learning model (NN) generated by this training is stored in the learning model storage unit 132.

[0111] Fig. 15 is an example of a flowchart showing a procedure of a determination process using a classifier. In Fig. 15, the same step numbers are assigned to the processing steps having the same processing contents as those in Fig. 13. In the determination process shown in Fig. 15, steps S51 to S53 are executed instead of steps S25 to S28 in Fig. 13.

[0112] [Step S51] The determination unit 126 inputs the generated feature vector to a classifier based on the learning model data stored in the learning model storage unit 132, and determines whether the operation is normal or abnormal.

[0113] [Step S52] If an abnormal operation has occurred, the process proceeds to step S53; if a normal operation has occurred, the process proceeds to step S29. [Step S53] The judgment unit 126 executes a process of issuing a warning that an abnormal purchasing behavior has occurred. The warning may be issued in the same manner as in step S28 of FIG.

[0114] Here, it is assumed that the type of abnormal operation is also identified by the identifier. In this case, a warning can be issued according to the type of abnormal operation. For example, when a warning is issued by display information or voice as in step S28 of FIG. 13, the type of abnormal operation is notified by the display information or voice. The method of warning can also be changed depending on the type of abnormal operation. For example, if an abnormal operation is performed due to customer negligence, such as "missed scanning" or "missing the basket," a warning is issued to the customer by the display 52 of the cash register 50 or voice output, and the customer is prompted to try the scan operation again. On the other hand, if an abnormal operation is performed intentionally by the customer, such as "hiding the barcode" or "scanning the barcode incorrectly," a warning is issued to the store clerk by display information or voice.

[0115] In addition, in order to prevent trouble with customers, when an intentional abnormal operation is performed, information such as the current time, the identification number of the cash register 50, and the type of abnormal operation may be stored in a storage device without issuing a warning that would be noticeable to the customer, or captured video data may be stored in a storage device as evidence.

[0116] In the above third embodiment, by learning based on the product detection results by HOID and the detection results of the person's skeletal information, a classifier can be generated that can identify abnormal purchasing actions in which the product or hand joints move in a specific way. In particular, when the learning data is labeled for each type of abnormal action, a classifier can be generated that can explicitly identify multiple types of abnormal actions. By using the product detection results by HOID and the detection results of the person's skeletal information, it becomes possible to distinguish subtle differences in purchasing actions, and a classifier can be generated that can identify multiple types of abnormal actions with high accuracy.

[0117] [Fourth embodiment] In the fourth embodiment, a part of the processing of the monitoring device 100 in the second embodiment is modified. Specifically, whether or not a normal purchasing action has been performed is determined according to a predetermined determination rule based on the position on the image of the object (product) detected by the image feature extraction unit 122. In addition, based on the skeletal information detected by the skeletal information extraction unit 123, the content of the determination rule is corrected according to the physique and standing position of the customer.

[0118] Fig. 16 is a diagram showing an example of the configuration of processing functions included in a monitoring device according to a fourth embodiment. In Fig. 16, the same reference numerals are used to denote processing functions that execute the same processes as those in Fig. 8. The monitoring device 100a shown in Fig. 16 includes a determination rule storage unit 161, an environmental difference correction unit 162, an individual difference correction unit 163, a customer operation recognition unit 164, and a purchasing behavior determination unit 124a in addition to the image acquisition unit 121, the image feature extraction unit 122, and the skeletal information extraction unit 123 shown in Fig. 8.

[0119] The judgment rule storage unit 161 is realized as a storage area of ​​a storage device included in the monitoring device 100a. The judgment rule storage unit 161 stores information (judgment rule information) indicating a judgment rule for judging whether or not a normal purchasing behavior has been performed. For example, information indicating a plurality of judgment areas set on an image for recognizing that a customer has performed a specific behavior from the position of a product is stored as the judgment rule information.

[0120] The processes of the environmental difference correction unit 162, the personal difference correction unit 163, the customer operation recognition unit 164, and the purchasing behavior determination unit 124a are realized, for example, by a processor included in the monitoring device 100a executing a predetermined program.

[0121] The environmental difference correction unit 162 corrects the image acquired by the image acquisition unit 121 in accordance with the shooting environment so that the determination rule can be applied correctly. The individual difference correction unit 163 corrects the judgment rule information stored in the judgment rule storage unit 161 based on the skeletal information detected by the skeletal information extraction unit 123. Through this correction, the judgment rule information is corrected according to the physique and standing position of the customer.

[0122] The customer operation recognition unit 164 recognizes operations by a customer using the cash register 50. For example, a scanning operation of an item and a settlement start operation performed after the scanning operation are completed are recognized as customer operations.

[0123] The customer operation recognition unit 164 recognizes these operations, for example, based on the image captured by the camera 101. For example, a lamp mounted on the cash register 50 may light up when a scanning operation or a settlement start operation is performed. A different lamp may light up for each operation, or the lamp may light up in a different color for each operation. In such cases, the customer operation recognition unit 164 can recognize that the above operations have been performed by detecting the lighting of the lamp or the color when it is lit from the captured image.

[0124] Furthermore, the customer operation recognition unit 164 may recognize each operation by recognizing from the captured image that the display content of the display 52 of the cash register 50 changes in response to each operation.

[0125] Furthermore, for example, if the cash register 50 generates different notification sounds for each operation, the customer operation recognition unit 164 can detect the generation of the notification sounds via a microphone and recognize that the above operations have been performed. Also, if the monitoring device 100a and the cash register 50 are capable of communicating with each other, the customer operation recognition unit 164 may receive a notification from the cash register 50 indicating that the above operations have been performed.

[0126] The purchasing behavior determination unit 124a determines whether a normal purchasing behavior has been performed based on the position information of the object (product) output from the image feature extraction unit 122 and in accordance with the determination rule indicated by the determination rule information corrected by the individual difference correction unit 163. The purchasing behavior determination unit 124a also counts the number of scan operations performed based on the determination result, and counts the number of scan operations recognized by the customer operation recognition unit 164. The purchasing behavior determination unit 124a compares the two count values ​​when the customer performs a checkout start operation, and if they do not match, executes processing to warn that a normal purchasing behavior has not been performed.

[0127] FIG. 17 is a diagram for explaining a determination rule for determining whether or not a normal purchasing action has been performed. 17 is an example of a captured image acquired by the image acquisition unit 121. This image 241 is an image captured from above the cash register 50, near the front surface of the cash register 50 (the surface on which the barcode scanner 51 is mounted). The image 241 also illustrates an example of a product area (bounding box) 242 for the product 213, which is indicated by the HOID information output from the image feature extraction unit 122. Furthermore, a center position 243 of the product area 242 is also illustrated.

[0128] In the judgment rule, for example, a plurality of judgment regions are used to recognize that each customer has performed a specific action. In this case, the judgment rule information stored in the judgment rule storage unit 161 includes information indicating the position of each of these judgment regions. As an example, it is assumed here that two judgment regions, a take-out region R1 and a take-out region R2, are set. An image 241a shown in FIG. 17 is an image 241 with the take-out region R1 and the take-out region R2 superimposed on it.

[0129] The removal area R1 is an area for detecting the action of a customer picking up an item to perform a scanning operation. In many cases, a customer approaches the cash register 50 with the items to be purchased placed in the in-store basket 216, and removes the items one by one from the basket 216 to perform a scanning operation. In this case, the removal area R1 is an area for detecting that the customer has removed the items from the basket 216 to perform a scanning operation. Therefore, in the following description, the above-mentioned action to be detected based on the removal area R1 is referred to as a "removal action." Here, it is assumed that the removal action is started when the center position 243 of the item area 242 newly enters the removal area R1, and that the removal action is ended when the center position 243 moves outside the removal area R1.

[0130] The removal area R1 is set at a position a certain distance away from the front of the cash register 50. In addition, in the example of Fig. 17, the removal area R1 is also set in the area close to the front of the cash register 50, between the area where the barcode scanner 51 is located and the area where the basket 216 is placed.

[0131] However, removal area R1 is set so as not to include the area where basket 216 is placed. For example, multiple products are often close to each other in basket 216. For this reason, according to the HOID process, when a customer is carrying basket 216, multiple products in basket 216 may be detected as being held by the customer. Also, when placing basket 216 to remove products, if a product that was underneath another product is removed, it may be erroneously detected that the customer is holding the upper product that is not the target for removal. By setting removal area R1 so as not to include the area where basket 216 is placed, such erroneous detection can be prevented.

[0132] On the other hand, take-out area R2 is an area for detecting the action of scanning an item picked up by a customer, and is set in a position close to barcode scanner 51. Take-out area R2 can also be said to be an area for detecting whether a customer has performed a normal action to take the item picked up out of the store. Take-out area R2 is set so that at least one of its boundaries abuts the boundary of removal area R1. Specifically, of the boundaries of take-out area R2, the boundary that separates the customer from the cash register 50 (the boundary on the cash register 50 side, parallel to the front of the cash register 50) is set so as to abut the boundary of removal area R1.

[0133] From the relationship between the take-out area R1 and carry-out area R2 and the center position 243 of the product area 242, it is determined whether or not a normal purchasing action has been performed, for example, according to the following determination rule. When the center position 243 moves into the take-out area R1, the purchasing action determination unit 124a recognizes that one product has been taken out (a take-out action has been performed). When the center position 243 moves from the take-out area R1 to the take-out area R2 from that state, the purchasing action determination unit 124a recognizes that a scan operation has been performed (a correct take-out action has been performed). When such a series of actions is detected, it is determined that a normal purchasing action has been performed.

[0134] On the other hand, if the center position 243 does not move into the take-out area R2, it is determined that a normal purchasing behavior has not been performed. For example, if the center position 243 moves from within the take-out area R1 to outside the take-out area R1 without moving into the take-out area R2, it is assumed that the customer put the product in his / her pocket or bag without scanning it. In such a case, the purchasing behavior determination unit 124a recognizes that an abnormal purchasing behavior has been performed and can execute a process of issuing a warning.

[0135] During the above determination, a method for recognizing that the same product is being held in a hand between consecutive frames can be, for example, a method for recognizing based on the trajectory of the product's position, a method for recognizing based on image information of the product area, etc. As a method using image information, for example, a method for recognizing by comparing the luminance or color histogram within a bounding box can be used.

[0136] In the above example, the purchasing behavior is determined from the relationship between the set area on the image and the position of the product, but the purchasing behavior may also be determined by using, for example, image changes (background difference) within the set area before and after the product enters the set area. For example, by detecting image changes in the removal area R1, it becomes possible to accurately determine that the recognized object is a product and trace the movement trajectory of the product.

[0137] Fig. 18 is a diagram showing examples of correction processing by the environmental difference correction unit, Fig. 18(A) shows a first correction processing example, and Fig. 18(B) shows a second correction processing example. The environmental difference correction unit 162 corrects the image acquired by the image acquisition unit 121 in accordance with the shooting environment so that the judgment rule can be correctly applied to execute the judgment process.

[0138] For example, the environmental difference correction unit 162 corrects the captured image based on the imaging environment information in which information indicating the imaging environment of the camera 101 is recorded. The imaging environment information records, for example, the number of pixels of the image captured by the camera 101, the imaging direction of the camera 101, the distance between the camera 101 and a reference position in real space (for example, a predetermined position in front of the cash register 50), and the like. The determination rule storage unit 161 registers reference values ​​related to the information recorded in the imaging environment information, and the environmental difference correction unit 162 compares the imaging environment information with the reference values ​​to correct the captured image so that the number of pixels of the captured image, the imaging direction, and the distance from the reference position match the reference values.

[0139] Moreover, the environmental difference correction unit 162 may perform such correction using a plurality of markers placed in real space instead of the shooting environment information. In this case, the determination rule storage unit 161 registers a reference value for the positional relationship between the markers on the captured image (the distance between the markers and the relative position). The environmental difference correction unit 162 detects the position of each marker in the captured image, and corrects the captured image so that the positional relationship between the markers on the captured image matches the reference value registered in the determination rule storage unit 161.

[0140] 18A shows an example in which the distance between markers on a captured image 251 is smaller than a reference value. In this example, markers M1 to M4 are placed in real space, and these markers M1 to M4 appear in the captured image 251.

[0141] Here, it is assumed that in the determination rule information, the boundary line L1 that separates the customer from the cash register 50 among the boundaries of the removal area R1 is set to a position 300 pixels from the top of the image. If the distance between the markers on the captured image 251 captured by the camera 101 is smaller than a reference value, if the position of the boundary line L1 set in the determination rule information is used as is, it is not possible to correctly determine whether or not a product is located in the removal area R1. For this reason, the environmental difference correction unit 162 performs a correction to enlarge the captured image 251 so that the distance between the markers matches the reference value. The determination process is performed using the corrected (enlarged) captured image 251a, so that the determination process can be performed accurately.

[0142] Furthermore, when markers are used, a correction can be performed by rotating the captured image as shown in Fig. 18(B). In the captured image 252 shown in Fig. 18(B), a comparison between the positional relationship between the markers M1 to M4 and the reference value registered in the judgment rule storage unit 161 reveals that the camera 101 has rotated with respect to the lens optical axis. In such a case, the environmental difference correction unit 162 rotates the captured image 252 so that the angular relationship between the markers matches the reference value, and then cuts out an image region 253 from the rotated captured image 252 such that the distance between the markers matches the reference value. By such a correction, a corrected captured image 252a is obtained.

[0143] 19 is a diagram showing an example of the internal configuration of the individual difference correction unit 163. As shown in FIG. 19, the individual difference correction unit 163 includes an individual difference correction unit 171 and a person position correction unit 172. The individual difference correction unit 171 estimates a person's physique from the skeletal information detected by the skeletal information extraction unit 123, and corrects the judgment rule information stored in the judgment rule memory unit 161 so that accurate judgment can be performed regardless of physique differences between people.

[0144] The person position correction unit 172 recognizes the position where the person is standing (standing position) based on the position information of the person detected by the image feature extraction unit 122 or the skeletal information detected by the skeletal information extraction unit 123. The person position correction unit 172 corrects the determination rule information stored in the determination rule storage unit 161 based on the recognized standing position so that accurate determination can be performed regardless of the person's standing position.

[0145] FIG. 20 is a diagram for explaining the correction process by the person difference correction unit. As described above, the purchase action determination unit 124a detects the take-out action when the position of the product enters the take-out area R1, and then detects the execution of a scan operation when the position of the product enters the take-out area R2. For example, the position where a normal take-out action is performed and the position where a normal scan operation is performed are roughly determined by the length between the joint points of the person (particularly, the length of the arm from the elbow to the wrist) and the standing position of the person. Therefore, by optimizing the boundary between the take-out area R1 and the take-out area R2, which is actually used to distinguish between the take-out action and the scan operation, according to the length between the predetermined joint points of the person in the captured image and the standing position of the person, it becomes possible to perform accurate action determination.

[0146] 20 illustrates an example of a take-out area R1 and a carry-out area R2. A boundary line L2 between the take-out area R1 and the carry-out area R2 is set horizontally on the photographed image 261 so as to separate the customer from the cash register 50. The person difference correction unit 163 corrects the vertical position of the boundary line L2 according to the length between predetermined joint points of a person and the standing position of the person.

[0147] For example, the individual difference correction unit 171 calculates the length of the person's arm based on the skeletal information detected by the skeletal information extraction unit 123. If the person has a long arm, the vertical range when the product removal action is performed correctly may be wide in the downward direction (toward the cash register 50). Therefore, if the person's arm is longer than a predetermined threshold, the individual difference correction unit 171 corrects the boundary line L2 downward on the image, and if the person's arm is equal to or shorter than the threshold, the individual difference correction unit 171 corrects the boundary line L2 upward on the image. The amount of correction may be determined according to, for example, the difference between the arm length and the threshold.

[0148] Furthermore, the person position correction unit 172 detects the standing position of the person in the captured image 261. For example, the standing position of the person can be detected from the position information of the person detected by the image feature extraction unit 122. Alternatively, when the skeletal information extraction unit 123 detects the joint points of the person's feet (for example, the ankles, knees, etc.), the positions of the joint points can also be detected as the standing position of the person.

[0149] The further downward the person is positioned on the image, the wider the vertical range in the downward direction (toward cash register 50) will be when the product removal action is performed correctly. Therefore, person position correction unit 172 corrects boundary line L2 downward on the image when the y coordinate (vertical coordinate) indicating the person's position is greater than a predetermined reference vertical coordinate (when located downward), and corrects boundary line L2 upward on the image when the y coordinate indicating the person's position is equal to or smaller than the reference vertical coordinate. The amount of correction may be determined, for example, according to the difference between the y coordinate indicating the person's position and the reference vertical coordinate.

[0150] Next, the process of the monitoring device 100a according to the fourth embodiment will be described with reference to a flowchart. 21 and 22 are examples of flowcharts showing the procedure of a judgment process based on judgment rules. In these Figs. 21 and 22, it is judged whether a correct purchasing action has been performed. At the same time, the number of times that the barcode scanning operation is recognized to have been performed by the judgment is compared with the actual number of times that the scanning operation has been performed in the cash register 50, and a notification or warning process is executed according to the comparison result.

[0151] [Step S61] First, as a pre-processing for the judgment process, a process for setting correction parameters used by the environmental difference correction unit 162 is executed. At this time, a plurality of markers are placed at predetermined positions within the shooting range, and the shooting range including each marker is photographed by the camera 101. When the image acquisition unit 121 acquires the photographed image, the environmental difference correction unit 162 determines correction parameters for correcting the photographed image by comparing the positional relationship of each marker on the photographed image with a reference value of the positional relationship of each marker registered in the judgment rule storage unit 161. As the correction parameters, for example, the enlargement / reduction ratio of the image, the rotation angle of the image, etc. are determined. The environmental difference correction unit 162 stores the determined correction parameters in a storage device such as the HDD 113.

[0152] In addition, when shooting environment information indicating the shooting environment of the camera 101 is provided, the correction parameters can be determined by comparing the shooting environment information with the reference values ​​of each piece of information contained in the shooting environment information, without taking a photograph using the camera 101.

[0153] Once the correction parameters have been set in the above manner, the camera 101 starts capturing images for executing the actual determination process. At this time, the markers are removed from the capturing range. [Step S62] The image captured by the camera 101 is corrected by the environmental difference correction unit 162 based on the correction parameters set in step S61, and input to the image feature extraction unit 122 and the skeletal information extraction unit 123. After this, every time a captured image is obtained, correction is performed by the environmental difference correction unit 162, and the corrected captured image is input to the image feature extraction unit 122 and the skeletal information extraction unit 123.

[0154] [Step S63] The purchasing behavior determination unit 124a determines whether a person and an object (product) that interacts with the person have been recognized by the image feature extraction unit 122. If a person has been recognized but not an object, or if a person has not been recognized, a wait state is entered and the process of step S63 is executed again for the next frame. On the other hand, if a person and an object have been recognized, the process proceeds to step S64.

[0155] The processes in the following steps S64 to S69 are executed for each frame. [Step S64] The individual difference correction unit 163 corrects the determination rule information stored in the determination rule storage unit 161, based on the skeletal information detected by the skeletal information extraction unit 123. For example, as described in Fig. 20, the ranges of the take-out region R1 and the take-out region R2 are adjusted according to the length between the predetermined joint points of the person and the standing position of the person.

[0156] [Step S65] The purchasing behavior determination unit 124a executes a process of counting the number of items using the determination rule information corrected in step S64. In this process, the number of times that the purchasing behavior determination unit 124a determines that the scanning operation has been performed correctly is counted as the number of items. The process content of step S65 will be described in detail later with reference to FIG. 23.

[0157] [Step S66] The purchasing behavior determination unit 124a counts, as a scan point, the number of times that the barcode of a product is scanned at the cash register 50. In this process, when the customer operation recognition unit 164 recognizes that a barcode has been scanned, the scan point is incremented.

[0158] [Step S67] The purchasing behavior determination unit 124a checks whether the number of products is greater than the number of scanned products. many Determine whether the number of items is greater than the number of scans. many If so, the process proceeds to step S68, and if the product quantity and the scanned quantity match, the process proceeds to step S69.

[0159] [Step S68] The purchasing behavior determination unit 124a notifies the customer to make a correction. For example, a voice is output indicating that an error has occurred. Alternatively, display information indicating that an error has occurred is displayed on the display 52 of the cash register 50. The customer then makes a correction. For example, the customer inputs information to the cash register 50 to redo the scan operation. The customer also picks up the product that he or she was holding in their hand immediately before. When the customer operation recognition unit 164 detects an input operation to redo the scan operation, for example, the purchasing behavior determination unit 124a counts down (increments) the number of scans. Then, the process proceeds to step S64.

[0160] [Step S69] The purchasing behavior determination unit 124a determines whether the customer will perform the settlement process. For example, when the customer operation recognition unit 164 detects that the customer has performed an operation on the cash register 50 to request the settlement process, it is determined that the customer will perform the settlement process. If it is not determined that the customer will perform the settlement process, the process proceeds to step S64. On the other hand, if it is determined that the customer will perform the settlement process, the process proceeds to step S71 in FIG. 22.

[0161] The explanation will be continued below with reference to FIG. [Step S71] The payment process is carried out by the cash register 50. For example, the display 52 of the cash register 50 shows the total price of the scanned items.

[0162] [Step S72] The purchasing behavior determination unit 124a checks whether the number of products is greater than the number of scanned products. many Determine whether the number of items is greater than the number of scans. many If so, the process proceeds to step S73. On the other hand, if the number of items matches the number of scanned items, the payment process continues, and when the payment process is completed (i.e., payment is finished), the process proceeds to step S75.

[0163] [Step S73] The purchasing behavior determination unit 124a executes a process of issuing a warning that the number of scanned products is insufficient. For example, warning information is displayed on the display 52 of the cash register 50. Alternatively, the warning information is output as audio through earphones worn by the store clerk.

[0164] [Step S74] The customer makes a correction. For example, the customer inputs a command to redo the scan operation in the cash register 50, and the payment is requested again. When the payment is completed, the process proceeds to step S75.

[0165] [Step S75] It is determined whether to end the determination process. If the determination process is to be continued, the process proceeds to step S63 in Fig. 21, where new person and object recognition processes are performed. On the other hand, if the determination process is to be ended, the process ends.

[0166] Fig. 23 is an example of a flowchart showing the procedure of the process of counting the number of scan points. The process in Fig. 23 corresponds to the process of step S65 in Fig. 21. In the process in Fig. 23, information indicating a newly recognized object (object information) is sequentially recorded in RAM 112. Flag information indicating whether or not a scan has been performed is added to this object information.

[0167] [Step S81] The purchasing behavior determination unit 124a determines whether the center position of the object region (bounding box) is within the take-out region R1. If the center position is within the take-out region R1, the process proceeds to step S82. If the center position is not within the take-out region R1, the process proceeds to step S86.

[0168] [Step S82] The purchasing behavior determination unit 124a determines whether the object determined to be in the take-out area R1 in step S81 is an object that has already been recorded. A recorded object means that object information about the object has been recorded in the RAM 112 as the most recent object information. If the object is a recorded object, the state in which the object is located in the take-out area R1 continues from the previous frame, and the process of FIG. 23 ends. On the other hand, if the object is not a recorded object, the process proceeds to step S83.

[0169] [Step S83] The purchasing behavior determination unit 124a newly registers, in the RAM 112, object information of the object determined to be within the take-out area R1 in step S81. At this time, flag information indicating that the object has not been scanned is added to the object information.

[0170] [Step S84] The purchasing behavior determination unit 124a refers to the object information registered immediately before the object information newly registered in step S83 (i.e., the object information of the object immediately before that determined to be located in the removal area R1). Based on the flag information added to this object information, the purchasing behavior determination unit 124a determines whether the object corresponding to this object information has not been scanned. If this object has not been scanned, the process proceeds to step S85, and if this object has been scanned, the process of FIG. 23 ends.

[0171] [Step S85] A case where the answer to the determination in step S84 is Yes is a case where the previous object has moved outside the range of the removal area R1 without moving from the removal area R1 to the take-out area R2. In this case, the product may have been taken out without being scanned. Therefore, the purchasing behavior determination unit 124a executes a process to warn the customer that the scan was not performed correctly, for example. For example, a sound is output indicating that the scan was not performed correctly. Or, display information indicating that the scan was not performed correctly is displayed on the display 52 of the cash register 50. After this, the process in FIG. 23 ends.

[0172] In step S85, for example, instead of issuing a warning, information indicating that an abnormal purchasing behavior has been performed may be recorded in RAM 112. In this case, if it is determined that the number of products and the number of scans are the same based on the recorded information in step S67 of Fig. 21 or step S72 of Fig. 22, a process of issuing a warning that an abnormality has occurred may be executed.

[0173] [Step S86] The purchasing behavior determination unit 124a determines whether the center position of the object region (bounding box) is within the take-out region R2. In this process, if the center position is within the take-out region R2 and the corresponding object is an object already recorded in the RAM 112, the process proceeds to step S87. This means that the object has moved from the take-out region R1 to the take-out region R2. On the other hand, if the center position is outside the take-out region R2, or if the center position is within the take-out region R2 but the object information of the corresponding object is not recorded in the RAM 112, the process of FIG. 23 ends.

[0174] [Step S87] The purchasing behavior determination unit 124a counts up the number of products. [Step S88] The purchasing behavior determination unit 124a records the corresponding object as an object that has been scanned. That is, the purchasing behavior determination unit 124a updates the flag information added to the object information of the corresponding object to indicate that the object has been scanned.

[0175] According to the fourth embodiment described above, products related to a person are detected, and whether or not the purchase behavior is correct is determined based on the position of the product. At the same time, the determination rule information used to determine the purchase behavior is optimized based on the detection result of the skeleton information. This makes it possible to accurately determine whether or not a person has performed a correct purchase behavior without using templates or learning data for each product, and regardless of the person's physique or the person's position in the shooting range.

[0176] The processing functions of the devices (e.g., the motion discrimination device 1, the monitoring device 100) shown in each of the above embodiments can be realized by a computer. In this case, a program describing the processing contents of the functions that each device should have is provided, and the above processing functions are realized on the computer by executing the program on a computer. The program describing the processing contents can be recorded on a computer-readable recording medium. Examples of computer-readable recording media include magnetic storage devices, optical disks, and semiconductor memories. Examples of magnetic storage devices include hard disk drives (HDDs) and magnetic tapes. Examples of optical disks include CDs (Compact Discs), DVDs (Digital Versatile Discs), and Blu-ray Discs (BD, registered trademark).

[0177] When distributing a program, for example, the program is recorded on a portable recording medium such as a DVD or CD and then sold. The program can also be stored in a storage device of a server computer and transferred from the server computer to other computers via a network.

[0178] A computer that executes a program stores, for example, a program recorded on a portable recording medium or a program transferred from a server computer in its own storage device. The computer then reads the program from its own storage device and executes processing according to the program. The computer can also read the program directly from a portable recording medium and execute processing according to the program. The computer can also execute processing according to the received program each time a program is transferred from a server computer connected via a network. [Explanation of symbols]

[0179] 1 Operation determination device 2. Captured images 3a person 3b Object 3c Reading device 4a,4b Image area 5a, 5b, 6a, 6b Skeleton line

Claims

1. On the computer, Acquire a photographed image of a person's movements; Detecting a position of an object related to the person and skeletal information of the person from the acquired photographed image; Calculating a first feature amount based on the position of the object detected from a first frame in the captured image and the skeleton information; inputting the first feature amount to a predictor configured to predict a feature amount in a second frame subsequent to the first frame, thereby calculating a predicted value of the feature amount in the second frame; Calculating a second feature amount based on the position of the object detected from the second frame and the skeleton information; determining whether or not the action of the person with respect to the object is normal based on a distance between the second feature amount and the predicted value; An action determination program that executes processing.

2. the predictor is generated by learning using the first feature amount detected from each of a plurality of past images captured in the past as learning data. The motion discrimination program according to claim 1.

3. Each of the plurality of past images is an image captured in a state in which the person normally performs an action on the object. The motion discrimination program according to claim 2.

4. By inputting the first feature amount into a classifier, it is possible to classify whether an action performed by the person with respect to the object is a normal action or one of multiple types of abnormal actions.

2. The motion discrimination program according to claim 1, which causes the computer to execute processing.

5. In the determination, calculating a first feature vector based on a position of the object and a second feature vector for each joint of the person based on the skeletal information; calculating the first feature amount by integrating the first feature amount vector and the second feature amount vector for each joint; 5. The motion discrimination program according to claim 1, 2, 3 or 4.

6. On the computer, Acquire a photographed image of a person's movements; Detecting a position of an object related to the person and skeletal information of the person from the acquired photographed image; Based on the position of the object and the skeletal information, it is determined whether or not the person has normally held the object and caused a scanner mounted in a cash register to read the product information attached to the object. An action determination program that executes processing.

7. On the computer, Acquire a photographed image of a person's movements; Detecting a position of an object related to the person and skeletal information of the person from the acquired photographed image; referring to area information indicating a plurality of image areas set in the photographed image, the area information being stored in a storage unit of the computer, and correcting the setting positions of the plurality of image areas in the photographed image based on the skeleton information; determining whether or not a motion of the person with respect to the object is normal based on a time-series relationship between the plurality of image regions in which the setting positions have been corrected and the position of the object; An action determination program that executes processing.

8. In the determination, it is determined whether a scanning operation in which the person holds the object in his / her hand and causes a scanner mounted on a cash register to read product information attached to the object has been performed normally. The motion discrimination program according to claim 7.

9. The computer includes: counting a first number of times that the determination indicates that the scanning operation has been performed normally and a second number of times that the cash register indicates that the scanning operation has been performed; outputting a predetermined warning information when the first number of times does not match the second number of times; Further processing may be performed, The motion discrimination program according to claim 8.

10. The computer Acquire a photographed image of a person's movements; Detecting a position of an object related to the person and skeletal information of the person from the acquired photographed image; Calculating a first feature amount based on the position of the object detected from a first frame in the captured image and the skeleton information; inputting the first feature amount to a predictor configured to predict a feature amount in a second frame subsequent to the first frame, thereby calculating a predicted value of the feature amount in the second frame; Calculating a second feature amount based on the position of the object detected from the second frame and the skeleton information; determining whether or not the action of the person with respect to the object is normal based on a distance between the second feature amount and the predicted value; Operation determination method.

11. The computer Acquire a photographed image of a person's movements; Detecting a position of an object related to the person and skeletal information of the person from the acquired photographed image; Based on the position of the object and the skeletal information, it is determined whether or not the person has normally held the object and caused a scanner mounted in a cash register to read the product information attached to the object. Operation determination method.

12. The computer Acquire a photographed image of a person's movements; Detecting a position of an object related to the person and skeletal information of the person from the acquired photographed image; referring to area information indicating a plurality of image areas set in the photographed image, the area information being stored in a storage unit of the computer, and correcting the setting positions of the plurality of image areas in the photographed image based on the skeleton information; determining whether or not a motion of the person with respect to the object is normal based on a time-series relationship between the plurality of image regions in which the setting positions have been corrected and the position of the object; Operation determination method.

13. Acquire a photographed image of a person's movements; Detecting a position of an object related to the person and skeletal information of the person from the acquired photographed image; Calculating a first feature amount based on the position of the object detected from a first frame in the captured image and the skeleton information; inputting the first feature amount to a predictor configured to predict a feature amount in a second frame subsequent to the first frame, thereby calculating a predicted value of the feature amount in the second frame; Calculating a second feature amount based on the position of the object detected from the second frame and the skeleton information; a processing unit that determines whether or not a motion of the person with respect to the object is normal based on a distance between the second feature amount and the predicted value; A motion discrimination device having the above configuration.

14. Acquire a photographed image of a person's movements; Detecting a position of an object related to the person and skeletal information of the person from the acquired photographed image; a processing unit that determines whether or not an operation of the person holding the object in his / her hand and having a scanner mounted in a cash register read product information attached to the object has been performed normally, based on the position of the object and the skeletal information; A motion discrimination device having the above configuration.

15. a storage unit that stores area information indicating a plurality of image areas that are set in a captured image capturing the motion of a person; Acquire the captured image, Detecting a position of an object related to the person and skeletal information of the person from the acquired photographed image; correcting setting positions of the plurality of image regions in the photographed image based on the skeletal information; a processing unit that determines whether or not a motion of the person with respect to the object is normal based on a time-series relationship between the plurality of image regions in which the setting positions have been corrected and a position of the object; A motion discrimination device having the above configuration.

Citation Information

Patent Citations

  • Examination cheat detection method and apparatus thereof

    CN107491717A

  • Crime prevention apparatus, crime prevention system, and crime prevention program

    JP2007148894A

  • Self-POS device and operation method therefor

    JP2014132501A

  • Monitoring method, monitoring device, and monitoring program

    JP2015176227A

  • Autonomous store tracking system

    JP2020053019A