Motion determination device, motion determination method, and program
The motion determination device efficiently identifies and corrects user actions in self-service facilities by analyzing characteristic movement patterns, enhancing accuracy and enabling real-time alerts or controls.
Patent Information
- Application Number
- JP2024504083
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-02
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-03-02
AI Technical Summary
Existing motion detection systems are not fast and efficient enough for monitoring user actions in self-service facilities.
A motion determination device and method that identifies a person's position and analyzes movements based on characteristic patterns to determine if actions are in a predetermined order, using a non-transitory computer-readable medium for faster and more accurate motion detection.
Enables quick and efficient determination of user actions, reducing erroneous detections and improving accuracy by identifying movements according to registered patterns, and allowing for real-time alerts or control actions.
Smart Images

Figure 0007790548000001 
Figure 0007790548000002 
Figure 0007790548000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a motion determination device, a motion determination method, and a non-transitory computer-readable medium. [Background technology]
[0002] At facilities such as self-service gas stations and pay parking lots, users are required to follow specific procedures for safety and efficiency reasons. There are technologies that monitor whether users are following the correct procedures.
[0003] For example, Patent Document 1 discloses a self-service refueling system that includes a means for capturing an image of a user refueling, a means for determining whether the user's behavior is abnormal, normal, or unknown based on the captured image, a means for stopping or prohibiting refueling if the user's behavior for each refueling process is determined to be abnormal or unknown, and a management terminal that displays the image that serves as the basis for the determination. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] International Publication No. 2021 / 117346 Summary of the Invention [Problem to be solved by the invention]
[0005] There is a demand for faster and more efficient motion detection.
[0006] In view of the above-mentioned problems, an object of the present disclosure is to provide a motion determination device, a motion determination method, and a non-transitory computer-readable medium that can determine a person's motion faster and more efficiently. [Means for solving the problem]
[0007] A motion determination device according to one aspect of the present disclosure includes: a position specifying means for specifying a position of a person from the acquired image data; a movement identification means for analyzing a movement of the person in the image data to identify a first characteristic movement according to a first characteristic movement pattern associated with a first position of the identified person, and for analyzing a movement of the person in the image data to identify a second characteristic movement according to a second characteristic movement pattern associated with a second position of the identified person; and a determining means for determining whether the identified characteristic actions are in a predetermined order.
[0008] A motion determination method according to one aspect of the present disclosure includes: The position of the person is identified from the acquired image data, Identifying a first characteristic motion by analyzing the motion of the person in the image data according to a first characteristic motion pattern associated with a first position of the identified person, and identifying a second characteristic motion by analyzing the motion of the person in the image data according to a second characteristic motion pattern associated with a second position of the identified person; It is determined whether the identified characteristic actions are in a predetermined order.
[0009] According to one aspect of the present disclosure, there is provided a non-transitory computer-readable medium, comprising: A process of identifying the position of a person from the acquired image data; a process of analyzing a movement of the person in the image data according to a first characteristic movement pattern associated with a first position of the identified person to identify a first characteristic movement, and analyzing a movement of the person in the image data according to a second characteristic movement pattern associated with a second position of the identified person to identify a second characteristic movement; and a process of determining whether the identified characteristic actions are in a predetermined order. [Effects of the Invention]
[0010] The present disclosure provides a motion determination device, a motion determination method, and a non-transitory computer-readable medium that can determine a person's motion faster and more efficiently. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a block diagram showing a configuration of a movement determining device according to a first embodiment. [Figure 2] 1 is a flowchart showing the flow of a motion determining method according to the first embodiment. [Figure 3] FIG. 10 is a diagram showing the overall configuration of a motion determining system according to a second embodiment. [Figure 4] FIG. 10 is a block diagram showing the configuration of a server and a terminal device according to a second embodiment. [Figure 5] FIG. 10 is a diagram showing skeletal information extracted from a frame image included in video data according to the second embodiment. [Figure 6] 10 is a flowchart showing the flow of a method for transmitting video data by a terminal device according to a second embodiment. [Figure 7] 10 is a flowchart showing the flow of a method for registering a registration operation ID and a registration operation sequence by a server according to a second embodiment. [Figure 8] FIG. 10 is a diagram for explaining a registration operation according to the second embodiment. [Figure 9] FIG. 10 is a diagram for explaining a normal operation sequence according to the second embodiment. [Figure 10] FIG. 10 is a diagram for explaining a caution action sequence according to the second embodiment. [Figure 11] 10 is a flowchart showing the flow of an operation determination method by a server according to a second embodiment. [Figure 12] 1 is a block diagram showing an example of the hardware configuration of an action determining device 100 and the like. DETAILED DESCRIPTION OF THE INVENTION
[0012] The present disclosure will be described below through embodiments, but the disclosure according to the claims is not limited to the following embodiments. Furthermore, not all of the configurations described in the embodiments are necessarily essential as means for solving the problems. In each drawing, the same elements are denoted by the same reference numerals, and redundant explanations are omitted as necessary.
[0013] <Embodiment 1> 1 is a block diagram showing the configuration of a motion determination device 100a according to a first embodiment. The motion determination device 100a is a computer system that identifies multiple motions performed by a user and determines whether the multiple motions have been performed in a predetermined order. The motion determination device 100a includes a position identification unit 106a, a motion identification unit 108a, and a determination unit 111a.
[0014] The movement determination device 100a includes a position determination unit 106a that determines the position of a person from acquired image data, a movement determination unit 108a that analyzes the movement of the person in the image in accordance with a first characteristic movement pattern associated with a first position of the identified person to determine a first characteristic movement, and that analyzes the movement of the person in the image in accordance with a second characteristic movement pattern associated with a second position of the identified person to determine a second characteristic movement, and a determination unit 111a that determines whether the identified characteristic movements are in a predetermined order.
[0015] The position identification unit 106a is also called a position identification means. The position identification unit 106a can identify the position of a person for each acquired image frame. The position of a person here may include a three-dimensional position in a world coordinate system (e.g., within a predetermined distance from a device, or the latitude and longitude of the person's feet) and the position of the person within the image. The position of a person within the image may be defined as a position related to the person's movement, for example, the position of a hand or a leg within a few pixels of a certain object. The position within the image may also define a three-dimensional position in real space, including the depth as seen from the camera.
[0016] The motion identification unit 108a is also referred to as motion identification means. The motion identification unit 108a can identify at least two different characteristic motions of a person according to the stored characteristic motion patterns associated with specific positions. The "first characteristic motion pattern associated with a first position" can be, for example, a pattern including a posture or motion of a person touching an anti-static pad near a fuel pump at a gas station. The "second characteristic motion pattern associated with a second position" can be, for example, a pattern including a motion of a person removing a fuel filler cap from a vehicle or a posture or motion of a person grabbing the fuel nozzle of a fuel pump at a gas station. These motion patterns can be stored as normal motion patterns associated with positions.
[0017] The determination unit 111a is also called a determination means. The determination unit 111a determines whether the characteristic movements are performed in the correct order. For example, if the correct order is a first characteristic movement followed by a second characteristic movement, and the movement identification unit 108a identifies the second characteristic movement after the first characteristic movement, the determination unit 111a determines that the order is correct. On the other hand, if the movement identification unit 108a identifies the first characteristic movement after the second characteristic movement, or if the movement identification unit 108a identifies the second characteristic movement without identifying the first characteristic movement, the determination unit 111a determines that the order is incorrect.
[0018] In some embodiments, the movement determining device further includes an output unit that outputs an alert. The output unit can output the alert when the determining unit 111a determines that the order is incorrect. The alert may be displayed on the display unit or output as sound from a speaker. In other embodiments, the movement determining device includes a processing control unit that controls the fuel dispenser to stop refueling.
[0019] In some embodiments, the movement determining device includes a storage unit that stores the characteristic movement patterns associated with the above-described predetermined sequences and positions.
[0020] 2 is a flowchart showing the flow of the movement determination method according to embodiment 1. First, the position specifying unit 106a specifies the position of a person from the acquired image data (S11).
[0021] The movement identification unit 108a analyzes the movement of the person in the image according to a first characteristic movement pattern associated with a first position of the identified person to identify a first characteristic movement, and analyzes the movement of the person in the image according to a second characteristic movement pattern associated with a second position of the identified person to identify a second characteristic movement (S12).
[0022] The determination unit 111a determines whether the identified characteristic actions are in a predetermined order (S13).
[0023] As described above, according to the first embodiment, the movement determining device 100a can quickly and efficiently determine whether the order of the user's movements is correct by identifying the user's movements according to the characteristic movement patterns registered in advance. Furthermore, by identifying the person's movements after identifying the position, it is possible to prevent erroneous detection of similar movements performed at different positions, and improve the accuracy of movement identification and movement order determination.
[0024] <Embodiment 2> 3 is a diagram showing the overall configuration of the movement determination system 1 according to embodiment 2. The movement determination system 1 is a computer system that monitors the movements of a user U who visits a fuel dispenser 50 at a gas station, determines whether the movements are performed in a predetermined order, and executes predetermined processing according to the determination result.
[0025] As an example, the normal flow when a user U refuels a vehicle 60 at a fuel dispenser 50 at a gas station is as follows. (1) First, the user U gets into the vehicle 60, stops the vehicle next to the fuel dispenser 50 at the gas station, and turns off the engine. (2) The user U operates a predetermined button inside the vehicle to open the fuel tank cap, opens the door and exits the vehicle. (3) User U operates the display panel of the fueling machine 50 to select the payment method (e.g., cash or credit card), the desired type of fuel (e.g., high-octane, regular, diesel), and the fueling conditions (e.g., full tank, fixed amount, fixed price). (4) User U touches the anti-static pad (also called an anti-static sheet) with his / her hand. (5) The user U manually turns and opens the cap inside the fuel filler opening 61 of the vehicle that was opened in (2) above. (6) The user U places the cap in a predetermined location (for example, behind the fuel filler opening 61 or in a predetermined location on the fuel pump). (7) The user U grabs one of the fuel nozzles 51a to 51c that corresponds to the selected type of fuel (for example, high octane, regular, diesel, etc.) and inserts the fuel nozzle into the fuel filler opening. (8) User U fills up the fuel tank by squeezing the trigger of the fuel nozzle. The nozzle is equipped with a sensor, so the fuel tank stops filling automatically when the tank is full. (9) After refueling is completed, the user U removes the fuel nozzle from the fuel filler opening 61 and returns it to the designated mounting location on the fuel pump. (10) The user U takes the receipt issued from the fuel dispenser 50 or the payment machine (not shown). (11) The user U removes the cap that was previously placed, closes the hole of the fuel filler opening 61 with the cap, and closes the fuel filler opening 61.
[0026] 3, the movement determination system 1 includes a server 100, a terminal device 200 in a fuel tanker 50, and one or more cameras 300. The server 100 and the terminal device 200 are communicatively connected via a network N. The network N may be wired or wireless.
[0027] The camera 300 is a camera that photographs the user U standing in front of the fuel tanker 50 and monitors the user U. The camera 300 is disposed at a position and angle that allows it to photograph at least a portion of the body of the user U standing in front of the fuel tanker 50. The camera 300 may be a plurality of cameras, one of which is disposed at a position and angle that allows it to photograph a vehicle in front of the fuel tanker 50. Note that while FIG. 3 illustrates a vehicle, this embodiment may also be applied to various types of mobile vehicles, such as motorcycles, trucks, and buses.
[0028] The terminal device 200 is a computer having a memory, a processor, etc. that controls the refueling machine 50. The terminal device 200 acquires video data from the camera 300 and transmits the video data to the server 100 via the network N. The terminal device 200 also receives warning information indicating that the server 100 has identified a cautionary action of the user U, and outputs the warning information using the display unit 203 or the audio output unit 204 (the display panel 55 or speaker of the refueling machine 50). The display panel 55 of the refueling machine 50 may be installed in a position that is easily visible to the user U or the store staff. The speaker (not shown) of the refueling machine 50 may also be installed in a position that is easily heard by the user U or the store staff. The terminal device 200 accepts a selection input made by the user U through a touch operation on the display panel 55 of the refueling machine 50, and executes predetermined processing on each piece of hardware in the refueling machine.
[0029] The server 100 is a computer that identifies actions, action sequences, and cautionary actions performed by the user U related to the refueling machine 50 based on the video data received from the terminal device 200. The server 100 detects the user's action sequence, determines whether the sequence is correct, and transmits the determination result to the terminal device 200 via the network N. If the server 100 detects a cautionary action (for example, an incorrectly sequenced action or a dangerous unit action), it transmits warning information to the terminal device 200 via the network N.
[0030] 4 is a block diagram showing the configuration of the server 100 and the terminal device 200 according to the second embodiment. The server 100 is also called an operation determining device.
[0031] (Terminal device 200) The terminal device 200 includes a communication unit 201, a control unit 202, a display unit 203, and an audio output unit 204. The terminal device 200 controls the fuel tanker 50, acquires video data from the camera 300, and transmits the video data to the server 100 as appropriate.
[0032] The communication unit 201 is also called a communication means. The communication unit 201 is a communication interface with the network N. The communication unit 201 is also connected to the camera 300, and acquires video data from the camera 300 at predetermined time intervals.
[0033] The control unit 202 is also referred to as control means. The control unit 202 controls the hardware of the terminal device 200 and the fuel dispenser 50. The control unit 202 controls, for example, the display panel 55 (touch panel), the fuel dispenser nozzle 51, the camera 300, and the receipt issuing machine (not shown) of the fuel dispenser 50. For example, when the control unit 202 detects a start trigger, it starts transmitting video data acquired from the camera 300 to the server 100. The detection of the start trigger refers to the above-mentioned "detection that the user's vehicle has visited the fuel dispenser." Furthermore, for example, when the control unit 202 detects an end trigger, it stops transmitting video data acquired from the camera 300 to the server 100. The detection of the end trigger refers to the above-mentioned "detection that the user U's vehicle has left the fuel dispenser 50."
[0034] When the communication unit 201 receives warning information regarding a cautionary action or action sequence of the user U from the server 100, the control unit 202 causes the display unit 203 to display the warning information. The control unit 202 may also cause the audio output unit 204 to output the warning information.
[0035] The display unit 203 is a display panel. The audio output unit 204 is an audio output device including a speaker.
[0036] (Server 100) The server 100 includes a registration information acquisition unit 101, a registration unit 102, an action DB 103, an action sequence table 104, an image acquisition unit 105, a position identification unit 106, an extraction unit 107, an action identification unit 108, a generation unit 109, an object recognition unit 110, a judgment unit 111, and a processing control unit 112.
[0037] The registration information acquisition unit 101 is also called a registration information acquisition means. The registration information acquisition unit 101 acquires a plurality of pieces of registration video data in response to an action registration request from the terminal device 200 or an operation by an administrator of the server 100. In the second embodiment, each piece of registration video data is video data showing an individual action included in a person's normal action or a cautionary action (for example, an action of operating a fuel pump, an action of touching an anti-static pad, etc.). Note that in the second embodiment, the registration video data is a video including a plurality of frame images, but may also be a still image (one frame image).
[0038] Furthermore, the registration information acquisition unit 101 acquires information on a plurality of registered action IDs and the chronological order in which the actions are performed in a series of actions in response to a sequence registration request from the terminal device 200 or an operation by the administrator of the server 100.
[0039] The registration information acquisition unit 101 supplies the acquired information to the registration unit 102.
[0040] The registration unit 102 is also called a registration means. First, the registration unit 102 executes an action registration process in response to an action registration request. Specifically, the registration unit 102 supplies registration video data to the extraction unit 107 (described later), and acquires skeleton information extracted from the registration video data from the extraction unit 107 as registered skeleton information. The registration unit 102 then registers the acquired registered skeleton information in the action DB 103 in association with a registered action ID.
[0041] Next, the registration unit 102 executes a sequence registration process in response to the sequence registration request. Specifically, the registration unit 102 generates a registration action sequence by arranging the registration action IDs in chronological order based on the chronological order information. At this time, if the sequence registration request relates to a normal action, the registration unit 102 registers the generated registration action sequence in the action sequence table 104 as a normal action sequence NS. On the other hand, if the sequence registration request relates to a caution action, the registration unit 102 registers the generated registration action sequence in the action sequence table 104 as a caution action sequence IS. Examples of caution actions include, but are not limited to, the action of smoking a cigarette and the action of lighting a lighter.
[0042] The movement DB 103 is a storage device that stores registered skeleton information corresponding to each of the movements included in normal actions in association with a registered movement ID. The movement DB 103 may also store registered skeleton information corresponding to each of the movements included in caution movements in association with a registered movement ID.
[0043] The action sequence table 104 stores a normal action sequence NS and a caution action sequence IS. In the second embodiment, the action sequence table 104 stores a plurality of normal action sequences NS and a plurality of caution action sequences IS. The action DB and the action sequence table may also be simply referred to as a storage unit.
[0044] The image acquisition unit 105 is also called image acquisition means. The image acquisition unit 105 acquires video data captured by the camera 300 from the terminal device 200 during operation of the fuel tanker 50. That is, the image acquisition unit 105 acquires video data in response to detection of a start trigger. The image acquisition unit 105 supplies frame images included in the acquired video data to the position identification unit 106, the extraction unit 107, the target recognition unit 110, etc.
[0045] The position identification unit 106 is also called a position identification means. The position identification unit 106 identifies the position of a person from the image data acquired by the image acquisition unit 105. The position identification unit 106 identifies the position of a person in a store (e.g., the position where the person is located near a static neutralization pad on a fuel pump, a fuel nozzle, a vehicle's fuel filler, etc.). If the distance between the position of a person's hand, recognized by a known image recognition technology, and an object (e.g., a static neutralization pad on a fuel pump, a fuel nozzle, a vehicle's fuel filler, etc.) is within a predetermined distance (e.g., within a few pixels in the image), the person can be identified as being located at the object's position. In another example, for example, since the camera's angle of view is fixed to a store (e.g., a gas station), a correspondence relationship between the position of a person in a captured image and the position of the person in the store (e.g., a relatively distant positional relationship such as the position of a fuel pump and the position of a vehicle's fuel filler) can be defined in advance, and the position in the image can be converted to a position in the store based on the definition. More specifically, in the first step, the height, azimuth, and elevation angles of the camera that captures images inside the store, as well as the focal length of the camera (hereinafter referred to as camera parameters) are estimated from the captured images using existing technology. These may be measured or referenced. In the second step, existing technology is used to convert the position of a person's feet from two-dimensional coordinates on the image (hereinafter referred to as image coordinates) to three-dimensional coordinates in the real world (hereinafter referred to as world coordinates) based on the camera parameters. Note that the conversion from image coordinates to world coordinates is usually not uniquely determined, but it can be made unique by fixing the coordinate value of the feet in the height direction to zero, for example. In the third step, a three-dimensional map of the transportation facility is prepared in advance, and the world coordinates obtained in the second step are projected onto the map to identify the person's position inside the store.
[0046] The extraction unit 107 is also called extraction means. The extraction unit 107 detects an image area (body area) of a person's body from a frame image included in the video data and extracts (e.g., cuts out) it as a body image. Then, the extraction unit 107 uses a skeleton estimation technique using machine learning to extract skeleton information of at least a part of the person's body based on features such as the person's joints recognized in the body image. The skeleton information is information composed of "key points" that are characteristic points such as joints, and "bones (bone links)" that indicate links between the key points. The extraction unit 107 may use a skeleton estimation technique such as OpenPose. The extraction unit 107 supplies the extracted skeleton information to the motion identification unit 108.
[0047] The action identification unit 108 is also called action identification means. The action identification unit 108 converts skeletal information extracted from video data acquired during operation into an action ID using the action DB 103. In this way, the action identification unit 108 identifies the action of the person. Specifically, first, the action identification unit 108 identifies registered skeletal information from the registered skeletal information registered in the action DB 103, whose similarity to the skeletal information extracted by the extraction unit 107 is equal to or greater than a predetermined threshold. Then, the action identification unit 108 identifies the registered action ID associated with the identified registered skeletal information as the action ID corresponding to the person included in the acquired frame image.
[0048] The motion identification unit 108 analyzes the motion of the person in the image according to a motion pattern associated with a first position of the identified person to identify a first characteristic motion, and then analyzes the motion of the person in the image according to a motion pattern associated with a second position of the identified person to identify a second characteristic motion. For example, the first characteristic motion may be the motion of the person touching an anti-static pad, which is associated with the position of the static discharge pad. The second characteristic motion may be the motion of the person grabbing a fuel nozzle, which is associated with the position of the fuel nozzle. In another example, another characteristic motion may be the motion of the person inserting the fuel nozzle into the fuel nozzle, which is associated with the position of the fuel filler cap of the vehicle.
[0049] Here, the motion identification unit 108 may identify one motion ID based on skeletal information corresponding to one frame image, or may identify one motion ID based on time-series data of skeletal information corresponding to each of multiple frame images. When identifying one motion ID using multiple frame images, the motion identification unit 108 may extract only skeletal information with large movements and compare the extracted skeletal information with registered skeletal information in the motion DB 103. Extracting only skeletal information with large movements may mean extracting skeletal information with a difference of a predetermined amount or more between skeletal information of different frame images included within a predetermined period. This reduced comparison reduces the computational load and the amount of registered skeletal information. Furthermore, since the duration of motions varies depending on the person, only skeletal information with large movements is compared, thereby making motion identification robust.
[0050] In addition to the above-described method, various other methods are possible for identifying the action ID. For example, there is a method of estimating the action ID from the target video data using an action estimation model trained on video data that has been assigned correct answers with action IDs as training data. However, collecting this training data is difficult and expensive. In contrast, in the second embodiment, skeleton information is used to estimate the action ID, and the action DB 103 is used to compare it with pre-registered skeleton information. Therefore, in the second embodiment, the server 100 can more easily identify the action ID.
[0051] The generation unit 109 is also called a generation means. The generation unit 109 generates an action sequence based on the multiple action IDs identified by the action identification unit 108. The action sequence is configured to include the multiple action IDs in chronological order. The generation unit 109 supplies the generated action sequence to the determination unit 111.
[0052] The object recognition unit 110 is also called an object recognition means. The object recognition unit 110 can recognize objects, particularly moving objects (e.g., vehicles and people), from the acquired image data using known image recognition technology or the like. The object recognition unit 110 can recognize vehicles that enter the angle of view captured by the camera 300, and can also recognize open fuel filler doors of vehicles. Furthermore, the object recognition unit 110 can also identify the position of the vehicle and the position of the fuel filler door in cooperation with the above-mentioned position identification unit 106.
[0053] In another embodiment, by incorporating a marker or the like that can be recognized by an image into the tip of the fuel nozzle, the target recognition unit 110 can recognize that the tip of the fuel nozzle has been inserted into the vehicle's fuel filler opening.
[0054] In another embodiment, when there are multiple people in the image, the target recognition unit 110 can recognize each person and determine whether the person who touched the anti-static pad is refueling. If the person touching the static neutralization pad is filling up the gas tank, the determination unit (described later) can determine that this is a correct action, and conversely, if the person touching the static neutralization pad is different from the person holding the gas tank nozzle, the action can be determined to be incorrect. In other words, the object recognition unit 110 can acquire the action history for each person in cooperation with the extraction unit 107, action identification unit 108, and generation unit 109 described above.
[0055] The determination unit 111 is also called a determination means. The determination unit 111 determines whether the generated operation sequence matches (corresponds to) any of the normal operation sequences NS registered in the operation sequence table 104. For example, the determination unit 111 determines whether the above-mentioned first characteristic operation and second characteristic operation are in the correct order (i.e., whether the second characteristic operation is performed after the first characteristic operation).
[0056] In some embodiments, the object recognition unit 110 recognizes a vehicle present in a predetermined stopping area in the image, and the determination unit 111 can determine whether or not a vehicle is stopped at the vehicle stopping position. When the determination unit detects that a person has performed a characteristic action, it can determine whether or not a vehicle is stopped at the vehicle stopping position.
[0057] The process control unit 112 is an example of the process control unit 21 described above. If the process control unit 112 determines that the generated operation sequence does not correspond to any of the normal operation sequences NS, it outputs warning information to the terminal device 200. In this case, the process control unit 112 is also referred to as an output means. For example, in the example of a gas station shown in FIG. 3 , the correct order is for a user U to touch the static elimination pad 53 and then grab the fuel nozzle 51 to remove static electricity from the human body. Furthermore, as described above in relation to the normal flow when refueling a vehicle, the order of operations performed by the user U can be arbitrarily set as the normal operation sequence.
[0058] In some embodiments, the memory unit stores a time limit for each step in which a unit characteristic action (e.g., an action of touching an anti-static pad, an action of grabbing a fuel nozzle, etc.) is performed, the judgment unit 111 judges whether or not the time limit has been exceeded for each step, and the output means outputs judgment information at the point in time when the time limit has been exceeded.
[0059] In some embodiments, the memory unit may store the time limit for each process, the judgment unit may judge whether the time limit has been exceeded for each process, and the output means may output judgment information at the point when the time limit has been exceeded.
[0060] If the determination unit 111 determines that the operation sequence does not correspond to any of the normal operation sequences NS, it may determine whether the operation sequence corresponds to any of the caution operation sequences. In this case, the process control unit 112 may output predetermined information corresponding to the type of caution operation sequence to the terminal device 200. As an example, the display mode (font, color, thickness, or blinking of characters, etc.) when displaying warning information may be changed depending on the type of caution operation sequence, or the volume or sound volume when outputting the warning information as audio may be changed. This allows the user U or store staff to recognize the content of the caution operation and respond promptly and appropriately to the caution operation. Furthermore, the process control unit 112 may record the time, location, and video of the caution operation together with information on the type of caution operation sequence as history information. This allows store staff to recognize the content of the caution operation and take appropriate preventive measures against the caution operation.
[0061] In some embodiments, the processing controller 112 can prevent a refueling button (e.g., on a display panel) from responding if the person does not touch the static discharge pad. That is, the processing controller 112 can limit activation of the refueling button to start refueling if the person does not touch the static discharge pad.
[0062] In some embodiments, the processing controller 112 can prevent a trigger on the fuel nozzle for initiating refueling from responding if the person does not touch the static neutralization pad, i.e., the processing controller 112 can limit activation of the trigger on the fuel nozzle for initiating refueling if the person does not touch the static neutralization pad.
[0063] In another embodiment, if the action identification unit 108 identifies a caution action in which the person or another person in the vicinity is smoking, the processing control unit 112 may stop the fuel pump from dispensing fuel to the fuel nozzle, or may notify store staff of the caution action.
[0064] In yet another embodiment, when the action identification unit 108 identifies one characteristic action of a person, the processing control unit 112 may display the next action in the normal action sequence on the display panel 55 to communicate this to the person.
[0065] FIG. 5 is a diagram showing skeletal information extracted from a frame image IMG400 included in video data according to the second embodiment. The frame image 400 is an image captured from the side of a user U performing a touch operation on the display panel 55 of a fuel tanker 50. The image area includes an image of the entire body of the user U. The skeletal information shown in FIG. 5 also includes multiple key points and multiple bones detected from the entire body. As an example, in FIG. 5, the following key points are detected: a left ear A12, a left eye A22, a nose A3, a neck A4, a right shoulder A51, a left shoulder A52, a right elbow A61, a left elbow A62, a left hand A72, a right hip A81, a left hip A82, a right knee A91, a left knee A92, a right foot A101, and a left foot A102.
[0066] The server 100 compares such skeletal information with registered skeletal information corresponding to the entire body and determines whether they are similar to each other, thereby identifying each action. For example, when identifying a refueling action, it is important to determine whether the person's hand approaches a predetermined object (e.g., an anti-static pad, a fuel nozzle, a vehicle's fuel filler cap, etc.), and for the actions of "removing the fuel nozzle from the attachment" or "inserting the fuel nozzle into the vehicle's fuel filler cap," the positions of the right and left hands in the frame image 400 are important. Therefore, the server 100 may calculate the similarity by weighting the positions of the right hand A71 and the left hand A72. The server 100 may also calculate the similarity by weighting the right shoulder A51, the left shoulder A52, the right elbow A61, and the left elbow A62 in addition to the right hand A71 and the left hand A72.
[0067] Furthermore, when identifying each action, an object (for example, a fuel nozzle, a static neutralization pad, or a vehicle) may be recognized from the image. Furthermore, before identifying an action, the position of a person may be identified from the image.
[0068] 6 is a flowchart showing the flow of a method for transmitting video data by the terminal device 200 according to the second embodiment. First, the control unit 202 of the terminal device 200 determines whether or not a start trigger has been detected (S20). If the control unit 202 determines that a start trigger has been detected (Yes in S20), it starts transmitting the video data acquired from the camera 300 to the server 100 (S21). On the other hand, if the control unit 202 does not determine that a start trigger has been detected (No in S20), it repeats the process shown in S20.
[0069] Next, the control unit 202 of the terminal device 200 determines whether or not an end trigger has been detected (S22). If the control unit 202 determines that an end trigger has been detected (Yes in S22), it ends the transmission of the video data acquired from the camera 300 to the server 100 (S23). On the other hand, if the control unit 202 does not determine that an end trigger has been detected (No in S22), it repeats the process shown in S22 while transmitting the video data.
[0070] In this way, by limiting the transmission period of video data to the period between a predetermined start trigger and end trigger, the amount of communication data can be minimized. Also, outside of this period, the motion identification process in the server 100 can be omitted, thereby saving computational resources.
[0071] In another embodiment, the movement of a person may be tracked, but the tracking data may be discarded if the person in the image moves out of the frame, thereby reducing resource usage.
[0072] 7 is a flowchart showing a flow of a method for registering a registration motion ID and a registration motion sequence by the server 100 according to the second embodiment. First, the registration information acquisition unit 101 of the server 100 receives a motion registration request including registration video data and a registration motion ID from the terminal device 200 (S30). Next, the registration unit 102 supplies the registration video data to the extraction unit 107. Having acquired the registration video data, the extraction unit 107 extracts a body image from a frame image included in the registration video data (S31). Next, the extraction unit 107 extracts skeletal information from the body image (S32). Next, the registration unit 102 acquires the skeletal information from the extraction unit 107 and registers the acquired skeletal information as registration skeletal information in the motion DB 103 in association with the registration motion ID (S33). Note that the registration unit 102 may register all of the skeletal information extracted from the body image as the registration skeletal information, or may register only a portion of the skeletal information (for example, skeletal information of the shoulders, elbows, and hands) as the registration skeletal information.
[0073] FIG. 8 is a diagram for explaining the registration actions according to the second embodiment. As an example, the action DB 103 may store registration framework information for eight registration actions having registration action IDs "A" to "H." The registration action "A" is an action of operating the display panel of the fuel dispenser 50 (for example, determining the type of fuel or payment method). The registration action "B" is an action of touching the static elimination pad 53 with a hand. The registration action "C" is an action of removing the cap from the fuel filler opening 61 of the vehicle. The registration action "D" is an action of removing the fuel filler nozzle (any of 51a to 51c) from the attachment part 52 of the fuel dispenser 50. The registration action "E" is an action of inserting the fuel filler nozzle into the fuel filler opening 61 of the vehicle. The registration action "F" is an action of squeezing the trigger of the fuel filler nozzle (any of 51a to 51c) to fill fuel. The registered action "G" is the action of removing the fuel nozzle from the fuel filler opening 61 of the vehicle and returning it to the mounting portion 52 of the fuel pump. The registered action "H" is the action of closing the hole of the fuel filler opening 61 of the vehicle with a cap. These registered action patterns can be stored as normal action patterns associated with positions. Note that these registered actions are examples and are not limited to these. Note that in some embodiments, cautionary actions may also be registered. Examples of cautionary actions include the posture or action of smoking a cigarette inside a gas station. Therefore, cautionary actions may be registered in association with a specific position (e.g., near the fuel filler nozzle), or may be registered in association with a broader area (e.g., within the area of the gas station).
[0074] Returning to Fig. 7, the explanation will be continued. Next, the registration information acquisition unit 101 receives a sequence registration request including a plurality of registration action IDs and information on the chronological order of each action from the terminal device 200 (S34). Next, the registration unit 102 registers a registration action sequence (normal action sequence NS or caution action sequence IS) in which the registration action IDs are arranged based on the information on the chronological order in the action sequence table 104 (S35). Then, the server 100 ends the processing.
[0075] FIG. 9 is a diagram illustrating a normal operation sequence NS according to the second embodiment. As an example, the operation sequence table 104 may include at least four normal operation sequences NS having normal operation sequence IDs of "11" to "14." The normal operation sequence "11" is a sequence (B→C→D) including an action of removing the cap of the fuel filler opening 61 after touching the static neutralization pad 53. The normal operation sequence "12" is a sequence (A→B→D) including an action of gripping the fuel filler nozzle 51 after touching the static neutralization pad 53. The normal operation sequence "13" is a sequence (A→B→C→D→E) including an action of inserting the fuel filler nozzle 51 into the fuel filler opening 61 after touching the static neutralization pad 53. The normal operation sequence "14" is a sequence (A→B→C→D→E→F→G) including an action of gripping the trigger of the fuel filler nozzle 51 and starting refueling after touching the static neutralization pad 53.
[0076] FIG. 10 is a diagram illustrating a caution action sequence IS according to the second embodiment. The action sequence table 104 may include at least two caution action sequences IS having caution action sequence IDs of "21" to "24." The caution action sequence "21" is a sequence (A→C→D) including an action of removing the cap of a fuel filler neck of a vehicle without touching the static neutralization pad. The caution action sequence "22" is a sequence (A→D→B) including an action of grabbing the fuel filler nozzle without touching the static neutralization pad. The caution action sequence "23" is a sequence (A→C→D→E) including an action of inserting the fuel filler nozzle into the fuel filler neck of a vehicle without touching the static neutralization pad. The caution action sequence "24" is a sequence (A→C→D→E→F→G) including an action of gripping the trigger of the fuel filler nozzle and starting refueling without touching the static neutralization pad.
[0077] Different warning information may be set to be issued for each of the caution operation sequences shown in Fig. 10. For example, a sequence including an operation of squeezing the trigger of the fuel nozzle and starting refueling without touching the static neutralization pad may indicate a higher level of danger than a sequence including an operation of removing the cap of the fuel filler opening of a vehicle without touching the static neutralization pad. Therefore, different warning information may be set to be issued in stages depending on the danger level. Furthermore, the operation sequences shown in Figs. 9 and 10 are merely examples, and various modifications are possible.
[0078] FIG. 11 is a flowchart showing the flow of a motion determination method performed by the server 100 according to the second embodiment. First, the image acquisition unit 105 of the server 100 starts acquiring video data from the terminal device 200 (S400). The position identification unit 106 identifies the position of a person from frame images included in the video data (S401). Furthermore, the object recognition unit 110 recognizes people and vehicles from frame images included in the video data (S402) and stores the recognized people and vehicles for subsequent processing of the same people and vehicles. The extraction unit 107 extracts body images from frame images included in the video data (S403). Next, the extraction unit 107 extracts skeletal information from the body images (S404). The motion identification unit 108 calculates the similarity between at least a portion of the extracted skeletal information and each registered skeletal information registered in the motion DB 103, and identifies, as a motion ID, a registered motion ID associated with registered skeletal information whose similarity is equal to or exceeds a predetermined threshold (S405). Next, the generation unit 109 adds the action ID to the action sequence. Specifically, in the first cycle, the generation unit 109 sets the action ID identified in S405 as the action sequence, and in the next and subsequent cycles, the generation unit 109 adds the action ID identified in S405 to the action sequence that has already been generated.
[0079] The determination unit 111 determines whether the operation sequence corresponds to any of the normal operation sequences NS in the operation sequence table 104 (S407). The determination unit 111 determines, for each unit operation, whether it corresponds to a normal operation sequence NS. If the operation sequence corresponds to a normal operation sequence NS (Yes in S407), the determination unit 111 proceeds to S410, and if it does not correspond (No in S407), the determination unit 111 proceeds to S408.
[0080] The determination unit 111 determines the type of the caution action by determining which of the caution action sequences IS in the action sequence table 104 the action sequence corresponds to (S408). Then, the process control unit 112 transmits warning information according to the type of the caution action to the terminal device 200 (S409).
[0081] The server 100 determines whether or not the acquisition of the video data has been completed (S410). If the server 100 determines that the acquisition of the video data has been completed (Yes in S410), the server 100 terminates the process. On the other hand, if the server 100 does not determine that the acquisition of the video data has been completed (No in S410), the server 100 returns the process to S403 and repeats the additional process of the operation sequence. By returning the process to S403, it is possible to monitor the operations from the end of the refueling operation until the user U leaves the refueling machine 50.
[0082] Thus, according to the second embodiment, the server 100 determines whether the actions of the user U who visited the refueling machine 50 are normal by comparing the action sequence showing the flow of the actions of the user U with the normal action sequence NS. In this way, by registering in advance a plurality of normal action sequences NS that are in line with the flow of operations using the refueling machine 50, it is possible to realize detection of actions that require caution in line with the actual situation. Note that the second embodiment also achieves the same effects as the first embodiment.
[0083] A general machine learning method has been disclosed that acquires and learns from videos of normal and abnormal behavior. This method suffers from a problem of reduced accuracy due to factors unrelated to behavior, such as the background, clothing, belongings, and shooting orientation. As a result, the technology has low versatility. In contrast, the present invention solves this problem by using information about a person's posture and, moreover, by using features that are robust to the person's orientation.
[0084] <Other embodiments> In the above embodiment, the server 100 and the terminal device 200 realize some functions in a distributed manner, but all functions may be integrated. Also, the camera 300 may be an intelligent camera equipped with a processor, memory, etc., and may include some or all of the functions of the server 100 and the terminal device 200.
[0085] Although the above embodiment has been described using an example of a gas station, the present invention may be applicable to various other applications. For example, the present invention may be applicable to other situations where static electricity may occur. For example, in the case of a wheelchair taxi, the present invention may be applicable to the procedure in which the driver gets off the taxi and prepares a ramp for the wheelchair. In the case of paid parking, the present invention may be applicable to guidance on the operation procedure at the payment machine.
[0086] 12 is a block diagram showing an example of the hardware configuration of the movement determining apparatuses 100a, 100, and terminal apparatus 200 (hereinafter referred to as movement determining apparatus 100, etc.) described in the above-mentioned embodiments. Referring to FIG. 12, the movement determining apparatus 100, etc. includes a processor 1201 and a memory 1202.
[0087] The processor 1201 reads and executes software (computer programs) from the memory 1202 to perform the processing of the movement determining device 100 and the like described using flowcharts in the above-described embodiments. The processor 1201 may be, for example, a microprocessor, an MPU (Micro Processing Unit), or a CPU (Central Processing Unit). The processor 1201 may include multiple processors.
[0088] The memory 1202 is configured by a combination of volatile memory and non-volatile memory. The memory 1202 may include storage located remotely from the processor 1201. In this case, the processor 1201 may access the memory 1202 via an I / O interface (not shown).
[0089] 12, the memory 1202 is used to store a group of software modules. The processor 1201 reads and executes these software modules from the memory 1202, thereby performing the processing of the motion determination device 100 and the like described in the above-described embodiment.
[0090] As described using FIG. 2 or 11, each of the processors included in the motion determination device 100 or the like executes one or more programs including a group of instructions for causing a computer to execute the algorithm described using the drawings.
[0091] Although the above-described embodiments have been described as hardware configurations, the present disclosure is not limited to such configurations. Any processing in the present disclosure can also be realized by causing a processor to execute a computer program.
[0092] In the above examples, the program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The program may be stored on a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable medium or tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD), Blu-ray® disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of propagated signals.
[0093] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. (Appendix 1) a position specifying means for specifying a position of a person from the acquired image data; a movement identification means for analyzing a movement of the person in the image data to identify a first characteristic movement according to a first characteristic movement pattern associated with a first position of the identified person, and for analyzing a movement of the person in the image data to identify a second characteristic movement according to a second characteristic movement pattern associated with a second position of the identified person; and determining means for determining whether the identified characteristic movements are in a predetermined order. (Appendix 2) 2. The movement determination device according to claim 1, further comprising a storage means for storing the predetermined order, a plurality of positions of the person, and a characteristic movement pattern performed by the person associated with the positions. (Appendix 3) 3. The motion determination device of claim 2, wherein the plurality of positions are associated with a plurality of regions within an image. (Appendix 4) 4. The motion determining device according to any one of claims 1 to 3, wherein the motion specifying means sets feature points and a pseudo skeleton of the person's body based on the image data. (Appendix 5) 5. The movement determining device according to any one of appendices 1 to 4, wherein the movement identifying means recognizes a time-series body movement of the person based on a plurality of consecutive image frames. (Appendix 6) 3. The movement determination device according to claim 2, wherein the storage means stores the characteristic movement pattern based on a plurality of consecutive image frames. (Appendix 7) The movement determination device described in Appendix 3, wherein the storage means stores each of the areas defined as adjacent or separated ranges in the image, and stores each posture in which the person's hands are present in a position approximating each of the areas as the characteristic movement pattern. (Appendix 8) Further comprising an object recognition means for recognizing an object, 8. The motion determining device according to any one of claims 1 to 7, wherein the motion identifying means identifies a plurality of the characteristic motions for the same object recognized by the object recognizing means. (Appendix 9) further comprising an output means for outputting a determination result regarding the determination; The motion determination device according to any one of appendices 1 to 8, wherein the determination means sequentially determines the unit characteristic motion of each step in the predetermined sequence, and the output means sequentially outputs the determination result for each step. (Appendix 10) the storage means stores a time limit for each step in the predetermined order, and the determination means determines whether or not the time limit has been exceeded in each step; 3. The motion determination device according to claim 2, further comprising an output unit that outputs a determination result when the time limit is exceeded. (Appendix 11) the storage means further stores an attention action pattern, and the action identification means identifies an attention action of the person in accordance with the attention action pattern; The determination means determines that the state is an attention state when the attention action is identified regardless of the order, 3. The movement determining device according to claim 2, further comprising an output means for outputting caution information indicating that the device is in a caution state. (Appendix 12) further comprising an object recognition means for recognizing a vehicle present in a predetermined stopping area in the image data; The action determination device according to any one of appendices 1 to 11, wherein, when the target recognition means recognizes that a vehicle is stopped at a predetermined position, the action identification means identifies that the person has performed the characteristic action. (Appendix 13) further comprising an object recognition means for recognizing a vehicle present in a predetermined stopping area in the image data; 13. The movement determining device according to any one of appendices 1 to 12, wherein the determining means starts the determination when the target recognizing means recognizes that a vehicle is stopped at a stop position. (Appendix 14) 13. The operation determination device according to any one of claims 1 to 12, further comprising an object recognition means for recognizing a fuel nozzle, wherein the determination means determines whether or not the fuel nozzle is inserted into a fuel filler opening of a vehicle. (Appendix 15) The position of the person is identified from the acquired image data, Identifying a first characteristic motion by analyzing the motion of the person in the image data according to a first characteristic motion pattern associated with a first position of the identified person, and identifying a second characteristic motion by analyzing the motion of the person in the image data according to a second characteristic motion pattern associated with a second position of the identified person; A motion determination method that determines whether the identified characteristic motions are in a predetermined order. (Appendix 16) A process of identifying the position of a person from the acquired image data; a process of analyzing a movement of the person in the image data according to a first characteristic movement pattern associated with a first position of the identified person to identify a first characteristic movement, and analyzing a movement of the person in the image data according to a second characteristic movement pattern associated with a second position of the identified person to identify a second characteristic movement; and a process of determining whether the identified characteristic actions are in a predetermined order. [Explanation of symbols]
[0094] 1. Motion detection system 50 Refueling Machine 51 Fuel nozzle 52 Mounting part 53 Anti-static pad 55 Display panel 60 vehicles 61 Fuel filler 100 servers 100a Motion determination device 101 Registration Information Acquisition Department 102 Registration Department 103 Operation DB 104 Operation Sequence Table 105 Image acquisition unit 106, 106a Position identification part 107 Extraction part 108,108a Operation specific part 109 Generation part 110 Object Recognition Unit 111,111a Judgment part 112 Processing control section 200 Terminal Device 201 Communications Department 202 Control section 203 Display section 204 Audio output section 300 cameras IMG400 Frame image N Network
Claims
1. a storage means for storing a predetermined order, a plurality of positions in an image, and characteristic movement patterns associated with the positions, the storage means storing the posture of the person when the person's hand is present at the position as the characteristic movement pattern; an object recognition means for recognizing a vehicle and an object person from the acquired image data; a position specifying means for specifying the position of the recognized hand of the target person when it is recognized from the acquired image data that the vehicle is present in a predetermined stopping area; a movement identification means for, when the vehicle is recognized to be present in a predetermined stopping area, analyzing the movement of the target person in the image data to identify a first characteristic movement according to a first characteristic movement pattern associated with a first position of the identified hand of the target person, and for analyzing the movement of the target person in the image data to identify a second characteristic movement according to a second characteristic movement pattern associated with a second position of the identified hand of the target person; and determining means for determining whether the identified characteristic movements are in a predetermined order.
2. The motion determination device according to claim 1 , wherein the plurality of positions are associated with a plurality of regions within the image.
3. The motion determination device according to claim 1 , wherein the motion identification means sets physical feature points and a pseudo skeleton of the person based on the image data.
4. 4. The motion determination device according to claim 1, wherein the motion identification means recognizes a time-series body motion of the person based on a plurality of consecutive image frames.
5. 3. The movement determination device according to claim 2, wherein the storage means stores each of the areas defined as adjacent or separated ranges in the image, and stores each posture in which the person's hands are present in a position approximating each of the areas as the characteristic movement pattern.
6. further comprising an output means for outputting a determination result regarding the determination; 6. The motion determination device according to claim 1, wherein the determination means sequentially determines the unit characteristic motion of each step in the predetermined sequence, and the output means sequentially outputs a determination result for each step.
7. the storage means stores a time limit for each step in the predetermined order, and the determination means determines whether or not the time limit has been exceeded in each step; The motion determining device according to claim 2 , further comprising an output unit that outputs a determination result when the time limit is exceeded.
8. A computer-implemented motion determination method, comprising: storing a predetermined order, a plurality of positions in the image, and characteristic motion patterns associated with the positions, and storing postures of the person when the person's hands are present at the positions as the characteristic motion patterns; Recognizes vehicles and people from the acquired image data, When the vehicle is recognized to be in a predetermined stopping area from the acquired image data, the position of the recognized hand of the target person is identified; When the vehicle is recognized to be present in a predetermined stopping area, a first characteristic motion is identified by analyzing the motion of the target person in the image data according to a first characteristic motion pattern associated with a first position of the identified target person's hand, and a second characteristic motion is identified by analyzing the motion of the target person in the image data according to a second characteristic motion pattern associated with a second position of the identified target person's hand; A motion determination method that determines whether the identified characteristic motions are in a predetermined order.
9. a process of storing a predetermined order, a plurality of positions in the image, and characteristic movement patterns associated with the positions, and storing postures of the person when the person's hands are present at the positions as the characteristic movement patterns; Recognizes vehicles and people from the acquired image data, a process of identifying the position of the recognized hand of the target person when the vehicle is recognized to be present in a predetermined stopping area from the acquired image data; When the vehicle is recognized to be present in a predetermined stopping area, a process of analyzing the movement of the target person in the image data to identify a first characteristic movement according to a first characteristic movement pattern associated with a first position of the identified hand of the target person, and analyzing the movement of the target person in the image data to identify a second characteristic movement according to a second characteristic movement pattern associated with a second position of the identified hand of the target person; and a process of determining whether the identified characteristic actions are in a predetermined order.
Citation Information
Patent Citations
Oil supply work monitoring device, oil supply work monitoring system, oil supply work monitoring method, and oil supply work monitoring program
JP2014118153A
Work support device, and work support program
JP2018156279A
Information processor, information processing method, computer program, and storage medium
JP2019101919A
Image analysis system and image analysis method
JP2020140236A
Work analyzer and work analysis program
JP2021163293A