Action recognition program, action recognition method, and information processing apparatus

The action recognition program uses camera-based skeletal analysis and machine learning to accurately recognize human actions at a low cost, addressing the limitations of existing technologies by enhancing detection accuracy and reducing false positives.

JP7711441B2Active Publication Date: 2025-07-23FUJITSU LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021098038
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-11
Publication Date
2025-07-23
Estimated Expiration
2041-06-11

AI Technical Summary

Technical Problem

Existing action recognition technologies face challenges in accurately recognizing human actions at low cost, with sensor-based systems being complex and costly, and camera-based systems struggling with up-down, left-right, and depth direction detection.

Method used

An action recognition program that utilizes a camera to detect skeletal information, acquires region information, estimates actions based on skeletal and region information, and recognizes actions by analyzing the distance between a person and an object, employing machine learning models for improved accuracy.

Benefits of technology

Accurately recognizes human actions at a low cost using a camera-based system, enhancing applications such as work analysis and purchase analysis by improving detection accuracy and reducing false positives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711441000001
    Figure 0007711441000001
  • Figure 0007711441000002
    Figure 0007711441000002
  • Figure 0007711441000003
    Figure 0007711441000003
Patent Text Reader

Abstract

To recognize an action of a target person at a low cost with high accuracy.SOLUTION: An information processing apparatus detects skeleton information of a person from an image obtained by an imaging apparatus. The information processing apparatus acquires region information on a position where the person was imaged, and positions of a target region and an object. The information processing apparatus estimates an action performed by the person with respect to the object, based on the skeleton information of the person and the region information. The information processing apparatus recognizes an action of the person based on a distance between the person and the object, and the estimated action.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an action recognition program, an action recognition method, and an information processing apparatus.

Background Art

[0002] With the development of AI (Artificial Intelligence) technology, technologies have been developed and utilized to recognize people and objects from images, and automatically detect the actions, postures, states, and behaviors of the recognized people from their skeletal information. For example, technologies for automatically detecting elderly people or people with disabilities and determining whether they are in a dangerous state, and technologies for recognizing the postures and processes of workers to determine whether they are entering a dangerous area, working in an unreasonable posture, or following the procedures.

[0003] As described above, by utilizing AI technology, it is possible to automatically analyze human actions and states, so applications in various fields such as purchase analysis, work analysis, on-site monitoring, surveillance, and detection of suspicious persons are desired. For example, a technology that combines a camera and a sensor to recognize the actions of a target person, particularly actions such as reaching out a hand or performing a task, is known.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, with the above technology, it is difficult to accurately recognize the actions of the target person at low cost. For example, recognition technologies using sensors generally use special sensors such as lasers and wireless, which have a complex configuration and high cost. Although recognition technologies using only cameras are also known, they cannot accurately detect the up-down, left-right, and depth directions.

[0006] On one aspect, an object is to provide an action recognition program, an action recognition method, and an information processing apparatus capable of accurately recognizing the actions of a target person at low cost.

Means for Solving the Problems

[0007] In a first aspect, the action recognition program causes a computer to detect skeletal information of a person from an image acquired by an imaging device, acquire region information regarding the position where the person is imaged, a target region, and the position of a target object, estimate an action that the person performs on the target object based on the skeletal information of the person and the region information, and recognize the action of the person based on the distance between the person and the target object and the estimated action.

Effects of the Invention

[0008] According to one embodiment, the actions of a target person can be accurately recognized at low cost.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Mode for Carrying Out the Invention

[0010] Hereinafter, embodiments of the action recognition program, action recognition method, and information processing apparatus disclosed in the present application will be described in detail with reference to the drawings. Note that the present invention is not limited by this embodiment. Also, each embodiment can be appropriately combined within a non - contradictory range.

Embodiment

[0011] [Overall Configuration] FIG. 1 is a diagram for explaining an information processing apparatus 10 according to Embodiment 1. As shown in FIG. 1, this system is a system in which an imaging device such as a camera installed in a retail store that sells clothing, food products, etc. to consumers is connected to an information processing apparatus 10 by wire or wirelessly, and the information processing apparatus 10 grasps the purchase situation in the retail store. The information processing apparatus 10 is an example of a computer device that acquires and analyzes video data inside a retail store, and generates and outputs information that can be used for identifying a consumer group that purchases a certain product, identifying a consumer group that picks up a certain product, etc.

[0012] Specifically, the information processing apparatus 10 identifies the skeletal information of a person from image data acquired by the imaging device. Then, the information processing apparatus 10 acquires region information regarding the position where the person is imaged, the target region, and the position of the target object, and estimates the action that the person performs on the target object based on the skeletal information of the person and the region information. After that, the information processing apparatus 10 recognizes the action of the person based on the distance between the person and the target object and the estimated action.

[0013] For example, when the information processing apparatus 10 recognizes the skeletal information of a person from the image data acquired from the camera, and the skeletal information meets the condition "including actions such as looking at (seeing) the shelf, walking towards the shelf, standing or squatting, stretching out the hand, etc., and the distance between the shelf and the person is within the reach of the hand", the information processing apparatus 10 recognizes that the person present in the image is "performing an action of stretching out the hand to a product on the shelf". And when the information processing apparatus 10 recognizes that "performing an action of stretching out the hand to a product on the shelf", it identifies the position where the hand is stretched out from the skeletal information.

[0014] By doing so, the information processing apparatus 10 can accurately recognize the actions of the target person by means of a simple and low-cost method using the video of the camera. As a result, the information processing apparatus 10 can accurately perform action analysis such as the target person's reaching motion analysis and work analysis such as shelf picking operations.

[0015] [Functional Configuration] FIG. 2 is a functional block diagram showing the functional configuration of the information processing apparatus 10 according to the first embodiment. As shown in FIG. 2, the information processing apparatus 10 includes a communication unit 11, an output unit 12, a storage unit 13, and a control unit 20.

[0016] The communication unit 11 is a processing unit that controls communication between other devices, and is realized by, for example, a communication interface or the like. For example, the communication unit 11 receives data such as video, images, or moving images from an imaging device such as a camera.

[0017] The output unit 12 is a processing unit that displays various types of information, and is realized by, for example, a display or a touch panel. For example, the output unit 12 outputs the result of action recognition recognized by the control unit 20 described later.

[0018] The storage unit 13 is a processing unit that stores various types of data and programs executed by the control unit 20, and is realized by, for example, a memory or a hard disk. This storage unit 13 stores a machine learning model DB 14 and a target location information DB 15.

[0019] The machine learning model DB 14 is a database that stores trained machine learning models. For example, the machine learning model DB 14 stores a first machine learning model that identifies a person shown in the image data in response to the input of image data such as each frame of the video data, and outputs region information of the person shown. Further, the machine learning model DB 14 stores a second machine learning model that outputs the skeletal information of the person shown in response to the input of the person's region information and the image data.

[0020] The target location information database 15 is a database that stores target location information regarding the area imaged by the imaging device. Specifically, the target location information database 15 stores, for each installed imaging device, target location information indicating the areas and positions of product shelves and products within the imaging area of each imaging device.

[0021] Figure 3 is a diagram showing an example of the information stored in the target location information database 15. As shown in Figure 3, the target location information database 15 stores the shelf ROI 30, which is the ROI (Region of Interest) of the product shelf, and the aisle ROI 40, which is the ROI of the aisle. In addition to these, the target location information database 15 stores the coordinates of the product shelf, the coordinates of each product installed on the product shelf (shelf ROI 30), the coordinates of the aisle, and so on.

[0022] Note that the target location information may be information in which the positions of the left, right, upper, and lower ends are recorded in advance as the area of the shelf. Also, the target location information may use the shelf area that has been learned and recognized in advance (Semantic Segmentation). Further, the target location information may use a group of areas in which the products on the shelf have been learned in advance as an object model and the objects have been detected using video analysis technology.

[0023] The control unit 20 is a processing unit that controls the entire information processing device 10 and is realized by, for example, a processor or the like. This control unit 20 includes a video acquisition unit 21, a person detection unit 22, a skeleton detection unit 23, a motion estimation unit 24, a range estimation unit 25, an arm-reaching action recognition unit 26, and an arm-reaching position calculation unit 27. Note that the video acquisition unit 21, the person detection unit 22, the skeleton detection unit 23, the motion estimation unit 24, the range estimation unit 25, the arm-reaching action recognition unit 26, and the arm-reaching position calculation unit 27 are realized by electronic circuits included in the processor or processes executed by the processor.

[0024] The image acquisition unit 21 is a processing unit that acquires video data. For example, the image acquisition unit 21 acquires video data of a target area from a camera connected via USB (Universal Serial Bus), LAN (Local Area Network), wireless, etc., and outputs it to the person detection unit 22. Note that the image acquisition unit 21 can also acquire pre-shot moving image data in addition to real-time video data from the camera.

[0025] The person detection unit 22 is a processing unit that detects a person shown in the video data from the video data. Specifically, for each frame of the video data acquired by the image acquisition unit 21, the person detection unit 22 performs person detection and outputs information about the detected person to the skeleton detection unit 23. For example, the person detection unit 22 performs person detection on the acquired video using video analysis techniques such as deep learning. Note that parameters such as a threshold value can be used to determine the person-likeness, and it is also possible to move on to the next frame video acquisition process.

[0026] The skeleton detection unit 23 is a processing unit that detects the skeleton information of a person from the image acquired by the imaging device. For example, the skeleton detection unit 23 performs skeleton detection on the person area detected by the person detection unit 22 using the same video analysis technique. Note that parameters such as a threshold value can be used to determine the skeleton-likeness, and it is also possible to move on to the next frame video acquisition process.

[0027] Here, person detection and skeleton detection will be specifically described. FIG. 4 is a diagram for explaining person detection and skeleton detection. As shown in FIG. 4, the person detection unit 22 inputs one frame of video data into the first machine learning model and acquires position information (area information) including the area of the detected person. Here, the person detection unit 22 can also adopt the output position information of the person when the score (probability) of the output value of the first machine learning model is equal to or higher than the threshold value.

[0028] Subsequently, the skeleton detection unit 23 inputs the image data in which a person is detected and the position information of the person obtained from the first machine learning model into the second machine learning model, and acquires the skeleton information of the detected person. Here, the skeleton detection unit 23 can also adopt the output skeleton information of the person when the score (probability) of the output value of the second machine learning model is equal to or greater than the threshold value.

[0029] FIG. 5 is a diagram for explaining an example of skeleton information. FIG. 5 shows the definition information of the skeleton acquired as the skeleton information. As shown in FIG. 5, the skeleton information stores 18 pieces (from No. 0 to No. 17) of definition information in which each joint specified by a known skeleton model is numbered. For example, No. 7 is assigned to the right shoulder joint (SHOULDER_RIGHT), No. 5 is assigned to the left elbow joint (ELBOW_LEFT), No. 11 is assigned to the left knee joint (KNEE_LEFT), and No. 15 is assigned to the right hip joint (HIP_RIGHT).

[0030] Therefore, the skeleton detection unit 23 acquires the coordinate information of the 18 skeletons shown in FIG. 5 from the second machine learning model. For example, the skeleton detection unit 23 acquires "X coordinate = X7, Y coordinate = Y7, Z coordinate = Z7" as the position of the right shoulder joint of No. 7. For example, the Z-axis can be defined as the distance direction from the imaging device to the object, the Y-axis can be defined as the height direction perpendicular to the Z-axis, and the X-axis can be defined as the horizontal direction.

[0031] Returning to FIG. 2, the motion estimation unit 24 is a processing unit that acquires region information regarding the region of an object at the position where a person is imaged, and estimates the motion that the person performs on the object based on the region information and the skeleton information of the person. Specifically, the motion estimation unit 24 estimates whether the motion that the person performs on the object matches any one of the conditions of an action of stretching a hand toward a shelf, an action of looking at the shelf, or an action of facing the shelf (a directly facing action) by the person detection unit 22, and outputs the estimation result to the range estimation unit 25 and the hand stretching action recognition unit 26.

[0032] FIG. 6 is a diagram for explaining the estimation of a person's motion. In FIG. 6, a person is depicted by connecting each joint included in the skeletal information with a line. As shown in FIG. 6, the motion estimation unit 24 uses the position information of the shelf stored in the target location information DB 15 and the skeletal information detected by the skeletal detection unit 23 to determine whether the detected motion of the person matches a motion learned in advance.

[0033] For example, when the motion estimation unit 24 is located within a predetermined position from the shelf and the detected skeletal information is similar to the skeletal information indicating motion A prepared in advance, the motion A is set as the estimation result. Also, when the motion estimation unit 24 is in a state of being located within a predetermined position from the shelf and the transition of each skeletal information detected so far is similar to the transition of the skeletal information of motion B prepared in advance, the motion B is set as the estimation result. Note that a known method such as whether the total value of the differences in each coordinate is less than a threshold value can be adopted as the similarity.

[0034] Here, when the motion estimation unit 24 makes a determination such as looking at the shelf or facing the shelf, for example, it uses the fact that the vertical vectors such as the face, shoulders, and torso intersect the area of object information such as the shelf and products set in advance.

[0035] For example, as an operation of stretching the hand, the motion estimation unit 24 may use, as a determination condition, that the angle of the arm is the largest among a series of operations. Specifically, when the angle of the arm using the 5th or 8th skeletal information is equal to or greater than the threshold value, the motion estimation unit 24 can also estimate that the hand has been stretched. Also, in addition to the operation of stretching the hand and then stretching the hand in the left-right, up-down directions, when the hand moves in a one-directional trajectory, the motion estimation unit 24 can also estimate that the hand has been stretched.

[0036] In addition, as an action of looking at a location, the motion estimation unit 24 may use, as an action of looking at the hand detected, in addition to the above-mentioned previously specified areas and object areas such as shelves, etc., to determine as an assumed action of looking at a location. For example, the motion estimation unit 24 monitors the transition of the third (HEAD) of the skeleton information using several frames, and when it detects an action of looking at the hand after an action of stretching out the hand, it estimates it as an action of looking at a location (shelf). In addition, when the motion estimation unit 24 detects an action of moving the face multiple times and then detects an action of stretching out the hand, it estimates it as an action of looking at a location (shelf).

[0037] In addition, as an action of facing the shelf, in addition to the actions of standing, squatting, and sitting, the motion estimation unit 24 may use, as a determination condition, that the body is facing the front or obliquely with respect to the area or object. For example, when the motion estimation unit 24 detects any one of the actions of standing, squatting, and sitting in a state where it is located within a predetermined position from the shelf and the body is facing the object (shelf), it estimates it as an action of facing the shelf. Note that each of the above-mentioned actions can be specified by monitoring the skeleton information and the transition of the skeleton information. In addition, the motion estimation unit 24 may estimate a plurality of actions.

[0038] Returning to FIG. 2, the range estimation unit 25 is a processing unit that estimates the reach range of a person's hand from the person's skeleton information and the area information of the shelf as the distance between the person and the object (shelf). For example, the range estimation unit 25 uses the estimation result by the motion estimation unit 24 and the previously set position information (object area) of the shelf, and for the position where the person stands, specified by, for example, the center point of both feet, etc., it estimates a certain distance (for example, corresponding to the shoulder width) from the shortest perpendicular line of the shelf as the reach range of the hand. In addition, the range estimation unit 25 may estimate the reach range of the hand using the length of the hand estimated from the skeleton information such as the length of the foot, height of the back, shoulder width, etc., detected by the skeleton detection unit 23, and the position of the hand estimated from biological characteristics such as the angle of the arm not bending more than 180 degrees.

[0039] FIG. 7 is a diagram for explaining the estimation of the reach range. As shown in FIG. 7, the range estimation unit 25 first detects the center points of both feet specified by the skeleton information detected by the skeleton detection unit 23 as the standing position (S1). Subsequently, the range estimation unit 25 calculates, as the shortest distance from the center points of both feet, a perpendicular line to the lower line of the shelf ROI 30, for example, with respect to the shelf (position information) stored in the target location information DB 15 (S2). Then, the range estimation unit 25 calculates a perpendicular line in the height direction of the shelf (position information) with respect to the shortest distance (S3). After that, the range estimation unit 25 uses the height (S4) from the foot to the shoulder and the shoulder width information (S5) in the skeleton information detected by the skeleton detection unit 23 to estimate, for example, 90 to 110% of the height and ±100% of the shoulder width as the reach range (the reach range of the hand) (S6). Then, the range estimation unit 25 outputs the estimation result to the reach action recognition unit 26. Note that the numerical values shown here are merely examples and can be arbitrarily changed.

[0040] Note that the range estimation unit 25 can estimate the reach range in the sitting state by the above determination from S1 to S6 even when it is determined that the operation is in the sitting state, not limited to the standing state. Further, the range estimation unit 25 can also estimate the reach range of the hand from the skeleton information of the shoulder, the elbow, and the fingertip and the perpendicular line of the shelf using the length of the arm specified by the skeleton information of the shoulder.

[0041] Returning to FIG. 2, the reach action recognition unit 26 is a processing unit that recognizes the action of a person reaching for an object from the action estimated by the action estimation unit 24 and the reach range of the person's hand estimated by the range estimation unit 25. Then, the reach action recognition unit 26 outputs the recognition result to the reach position calculation unit 27.

[0042] Specifically, the reach action recognition unit 26 uses the action estimated by the action estimation unit 24 and the reach range of the person's hand estimated by the range estimation unit 25, and determines with a previously set threshold value to recognize it as an action of reaching out the hand. Also, these determination items may be quantified as feature amounts, and the Mahalanobis distance or the like may be calculated and the result of machine learning may be used for determination.

[0043] The hand stretching position calculation unit 27 is a processing unit that calculates the position information of the outstretched hand using skeleton information, shelf ROI 30, etc. in an image (frame of video) recognized as an action of stretching the hand by the hand stretching action recognition unit 26.

[0044] In addition, the hand stretching position calculation unit 27 can calculate the position information of the hand at the time when the hand stretching action is recognized and access the corresponding product. For example, the hand stretching position calculation unit 27 identifies products that a person has picked up or is about to pick up by comparing the position information of the products within the shelf ROI 30 with the position information of the outstretched hand. In this way, when a person is identified from each frame of the video data, the hand stretching position calculation unit 27 identifies and aggregates the products that the person has stretched their hand towards.

[0045] Here, the action recognition and position calculation will be specifically described. FIG. 8 is a diagram for explaining the recognition and position calculation of a person's hand stretching action. As shown in FIG. 8, the hand stretching action recognition unit 26 recognizes whether it is a hand stretching action based on the score (e.g., coincidence rate) between the estimated result of the reach range of the hand and the estimated result of the person's movement and a predetermined condition.

[0046] For example, the hand stretching action recognition unit 26 recognizes a hand stretching action when it matches a predetermined number or more of conditions, or when it matches a predetermined percentage of all conditions. Note that the conditions can also be generated and set in advance using experimental data, etc., and can be defined and judged separately for the conditions for the estimated result of the reach range of the hand and the conditions for the estimated result of the person's movement.

[0047] Subsequently, the hand stretching position calculation unit 27 identifies the image data recognized as the hand stretching action, and calculates the hand position information using the skeleton information of the person identified using the image data, the shelf ROI 30, the position information of the product, and the like. For example, the hand stretching position calculation unit 27 calculates the coordinates of the shelf or product that overlaps with the hand in the image data as the hand position information using the coordinates of the shelf and the product. In addition, the hand stretching position calculation unit 27 can also input the coordinates of the shelf, the coordinates of the product, the recognition result of the hand stretching action, and the skeleton information of the person into a trained machine learning model to obtain the hand position information of the person.

[0048] Then, the hand stretching position calculation unit 27 identifies the product accessed by the person using the position information of the product and the hand position information of the person. For example, the hand stretching position calculation unit 27 identifies a product that coincides with the hand coordinates, a product within a predetermined range from the hand coordinates, or a product on the extension line from the hand coordinates in the shelf direction as the accessed product.

[0049] [Flow of processing] Next, the flow of the above-described action recognition processing will be described. Here, the overall flow of processing and the flow of processing by each processing unit will be described.

[0050] (Overall flow of processing) FIG. 9 is a flowchart showing the overall flow of the action recognition processing according to the first embodiment. As shown in FIG. 9, when the information processing apparatus 10 starts processing (S101: Yes), it acquires video data from the imaging device (S102).

[0051] Subsequently, the information processing apparatus 10 detects a person from each frame (image data) in the video data (S103). If a person cannot be detected (S104: No), it acquires the next video data. On the other hand, when the information processing apparatus 10 can detect a person (S104: Yes), it detects the skeleton of the detected person (S105).

[0052] Then, the information processing device 10 estimates the detected person's actions (S106) and estimates the reaching range of the detected person's outstretched hand (S107). Here, if the estimated reaching range of the hand does not reach the product or the shelf (S108: No), the information processing device 10 acquires the next image. On the other hand, if the range reaches the product or the shelf (S108: Yes), the information processing device 10 performs action recognition of the person (S109) and calculates the position of the person's hand (S110).

[0053] (Flow of person detection process) FIG. 10 is a flowchart showing the flow of the person detection process according to the first embodiment. As shown in FIG. 10, the person detection unit 22 acquires video data (S201), and performs person detection by applying the first machine learning model to each frame (image data) in the video data (S202).

[0054] Then, if the score (probability) of the output value of the first machine learning model is less than the threshold (S203: No), the person detection unit 22 determines that the prediction accuracy of the first machine learning model is low and acquires the next video data. On the other hand, if the score (probability) of the output value of the first machine learning model is greater than or equal to the threshold (S203: Yes), the person detection unit 22 determines that the prediction accuracy of the first machine learning model is high and outputs the position information of the detected person to the storage unit 13 or the like (S204).

[0055] (Flow of skeleton detection process) FIG. 11 is a flowchart showing the flow of the skeleton detection process according to the first embodiment. As shown in FIG. 11, the skeleton detection unit 23 acquires video data (S301) and acquires the position information of the person detected by the person detection unit 22 (S302).

[0056] Then, the skeleton detection unit 23 inputs the image data in which the person in the video data is detected and the position information of the person into the second machine learning model, acquires the skeleton information of the detected person (S303), and outputs it to the storage unit 13 or the like (S304). Here too, similar to FIG. 10, determination by score may be executed.

[0057] (Flow of Action Estimation Process) FIG. 12 is a flowchart showing the flow of the action estimation process according to the first embodiment. As shown in FIG. 12, the action estimation unit 24 acquires information (area information) of a target location including information on a shelf or the like in the imaging area from the target location information DB 15 (S401). Further, the action estimation unit 24 acquires the skeletal information of the person from the skeleton detection unit 23 (S402).

[0058] Subsequently, the action estimation unit 24 estimates the action that the person performs on the target object based on the skeletal information of the person and the information of the target location. Specifically, the action estimation unit 24 estimates the reaching action (S403), the looking action (S404), and the turning action (S405). Note that the action estimation unit 24 outputs information regarding the estimated action to the storage unit 13 or the like.

[0059] (Flow of Range Estimation Process) FIG. 13 is a flowchart showing the flow of the range estimation process according to the first embodiment. As shown in FIG. 13, the range estimation unit 25 acquires information (area information) of a target location including information on a shelf or the like in the imaging area from the target location information DB 15 (S501). Further, the action estimation unit 24 acquires the skeletal information of the person from the skeleton detection unit 23 (S502).

[0060] Subsequently, the range estimation unit 25 estimates the shortest distance from the person to the target location based on the skeletal information of the person and the information of the target location (shelf) (S503). Then, the range estimation unit 25 estimates the reachable range of the hand based on the estimated shortest distance and the skeletal information of the person (S504). Note that the range estimation unit 25 outputs information regarding the estimated reachable range of the hand to the storage unit 13 or the like.

[0061] (Flow of Recognition Process for Reaching Action) FIG. 14 is a flowchart showing the flow of the recognition process for the reaching action according to the first embodiment. As shown in FIG. 14, the reaching action recognition unit 26 acquires the action estimation information estimated by the action estimation unit 24 (S601), and acquires the estimation result of the reachable range of the hand estimated by the range estimation unit 25 (S602).

[0062] Subsequently, the reaching action recognition unit 26 calculates a score for the action of a person reaching their hand towards the target object (shelf) using the acquired information (S603). Here, if the score is less than the threshold value (S603: No), the reaching action recognition unit 26 returns to S601 and repeats the subsequent processing. On the other hand, if the score is equal to or greater than the threshold value (S603: Yes), the reaching action recognition unit 26 recognizes it as a reaching action (S604). Note that the reaching action recognition unit 26 outputs the recognition result of the recognized reaching action to the storage unit 13 or the like.

[0063] (Flow of calculation process for reaching position) FIG. 15 is a flowchart showing the flow of the calculation process for the reaching position according to the first embodiment. As shown in FIG. 15, the reaching position calculation unit 27 acquires the recognition result of the reaching action from the reaching action recognition unit 26 (S701), acquires the skeletal information of the person (S702), and calculates the position information of the outstretched hand using the recognition result, skeletal information, etc. (S703). Note that the reaching position calculation unit 27 outputs the calculated hand position information to the storage unit 13 or the like.

[0064] [Effect] As described above, the information processing apparatus 10 can accurately calculate the position where the hand is extended based on the human body characteristics from the skeletal information acquired from the video data and the shelf position information defined in advance. As a result, the information processing apparatus 10 can generate and output useful information for use in work state analysis, purchase analysis, etc. Also, by further analyzing the time required from reaching in front of the shelf to extending the hand, the reaching speed, trajectory, posture, and number of times, it is also possible to grasp the psychological state of the worker or purchaser.

[0065] In addition, since the information processing apparatus 10 performs person detection and skeleton detection using a machine learning model, it is possible to improve the detection accuracy by periodically performing retraining of the machine learning model and the like. Further, the information processing apparatus 10 prepares skeleton information of a pre-assumed motion in advance, and can estimate the motion by comparing the detected skeleton information with them. Therefore, it is possible to achieve high-speed detection processing without depending on the image quality of the image data and without requiring advanced analysis of the image data.

[0066] In addition, the recognition technology that uses only sensors uses only the positional relationship acquired by the sensors, so it determines that the hand is extended when the hand position or body orientation overlaps with the product. For this reason, when a large shelf such as the product shelf described in the first embodiment is used as an object, false detection often occurs, and even in a place where the hand cannot reach, false detection occurs when there is an overlap in the image. On the other hand, the information processing apparatus 10 employs a mechanism that dynamically changes the reachable range of the hand according to the position of the person (skeleton information) and the position of the shelf, and does not detect when the hand and the product (or the entire shelf) overlap outside the reachable range of the hand, making it possible to improve the detection accuracy.

Embodiment

[0067] By the way, the information processing apparatus 10 can generate and output more useful information that can be used for work state analysis, purchase analysis, etc. by estimating the attributes of a person.

[0068] FIG. 16 is a flowchart showing the overall flow of the action recognition process according to the second embodiment. As shown in FIG. 16, when the information processing apparatus 10 starts the process (S801: Yes), it acquires video data from the imaging device (S802).

[0069] Subsequently, the information processing apparatus 10 performs person detection from each frame (image data) in the video data (S803). If a person cannot be detected (S804: No), it acquires the next video data. On the other hand, when the information processing apparatus 10 detects a person (S804: Yes), it detects the skeleton of the detected person (S805).

[0070] Then, the information processing apparatus 10 estimates the detected person's actions (S806) and estimates the detected person's reach range (S807). Here, if the reach range estimated by the information processing apparatus 10 is not within the reach of the product or the shelf (S808: No), the information processing apparatus 10 acquires the next image. On the other hand, if the reach range is within the reach of the product or the shelf (S808: Yes), the information processing apparatus 10 performs action recognition of the person (S809).

[0071] In parallel with these, the information processing apparatus 10 estimates the attributes of the person (S810). For example, an attribute estimation unit (not shown) of the information processing apparatus 10 inputs the detected person's skeletal information into a machine learning model generated by deep learning or the like to perform attribute estimation. As the attributes of a person, in addition to estimating gender and age, it is also possible to estimate store employee detection, suspicious person detection, etc. Note that parameters such as thresholds can be used to determine the likelihood of attributes, and the process can move on to the next frame video acquisition process.

[0072] After that, the information processing apparatus 10 calculates the hand position information at the time when the reaching action is recognized, adds the estimated attribute information, and outputs which attribute of the person accessed the corresponding product (S811).

[0073] In this way, the information processing apparatus 10 can use video data for a predetermined period such as one day or one week to further estimate the attributes of the person, and then aggregate and output to the user what kind of person picks up what kind of product at what time. FIG. 17 is a diagram for explaining an example of the output result screen. As shown in FIG. 17, the information processing apparatus 10 can generate and output a conversion analysis result 50 including various information generated using attribute information in addition to the information obtained by aggregating the number of people passing in front of the shelf and the number of people staying in front of the shelf in each time period of the store. The conversion analysis result 50 includes the acquired video 51, the ratio of genders that picked up the product with the left hand and the aggregation result 52 for each time period, the ratio of ages that picked up the product with the left hand and the aggregation result 53 for each time period, and the like.

Example

[0074] Now, although the embodiments of the present invention have been described so far, the present invention may be implemented in various different forms other than the above-described embodiments.

[0075] [Numerical values, etc.] The numerical examples, operation examples, screen examples, attribute examples, etc. used in the above embodiments are merely examples and can be arbitrarily changed. Also, a cloud system can be adopted. For example, the result processed by the edge terminal can be uploaded, and the result can be displayed on a browser via the cloud. Also, camera video can be uploaded, processed in the cloud, and the result can be displayed on a browser.

[0076] Also, the data to be detected is not limited to video data, and may be image data or video data. Also, although an example of recognizing an action of reaching out a hand and calculating the position of the outstretched hand has been described, it is not limited to this, and various actions of a person can be detected. For example, the information processing apparatus 10 can also detect mischief on a shelf by detecting an action of stretching a foot. Also, the information processing apparatus 10 can recognize an action of activating a sensor of a vehicle trunk with a foot in a parking lot, collect information useful for vehicle development, and output it.

[0077] [Action recognition example] For example, when a continuous action is estimated from a plurality of frames (image data) in a video frame, the information processing apparatus 10 can estimate a reaching-out hand action, thereby suppressing false detection and improving the estimation accuracy. FIG. 18 is a diagram for explaining an action estimation process using action transition. Specifically, as shown in FIG. 18, the information processing apparatus 10 can also estimate a specific action to transition to the recognition of a reaching-out hand action when it continuously detects (1) an action of coming in front of a shelf, (2) an action of looking in front of the shelf, standing or squatting in front of the shelf, or putting a hand forward, and (3) an action of stretching a hand within a predetermined number of frames.

[0078] For example, in the first frame, the information processing apparatus 10 detects an operation of "coming in front of the shelf" which is at a position more than a predetermined distance away from the shelf ROI 30 but in the passage ROI 40 in front of the shelf ROI 30. Next, in the second frame acquired within a predetermined number of frames from the first frame, the information processing apparatus 10 detects an operation of "looking at the front of the shelf" which is at a position less than the predetermined distance from the shelf ROI 30. Thereafter, in the third frame acquired within a predetermined number of frames from the second frame, the information processing apparatus 10 detects an operation of "extending the hand". In this way, when the information processing apparatus 10 detects a pre - assumed continuous operation, it can also execute the recognition of the hand - extending action.

[0079] Here, an example where a continuous operation is estimated from a plurality of frames (image data) in a video frame has been described, but it is not limited to this. For example, when a plurality of pre - specified operations are estimated in the pre - specified order, the information processing apparatus 10 can also execute the recognition of the hand - extending action.

[0080] [System] Regarding the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above - mentioned documents and drawings, they can be arbitrarily changed unless otherwise specified.

[0081] Also, each component of each illustrated device is a functional concept, and it is not necessarily physically configured as shown in the figure. That is, the specific form of the distribution and integration of each device is not limited to that shown in the figure. In other words, all or part of it can be functionally or physically distributed and integrated in any unit according to various loads, usage situations, etc.

[0082] Furthermore, each processing function performed by each device can be realized in whole or in any part by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware by wired logic.

[0083] [Hardware] FIG. 19 is a diagram for explaining a hardware configuration example. As shown in FIG. 19, the information processing apparatus 19 includes a communication device 10a, a HDD (Hard Disk Drive) 10b, a memory 10c, and a processor 10d. Further, each part shown in FIG. 2 is interconnected by a bus or the like.

[0084] The communication device 10a is a network interface card or the like and communicates with other devices. The HDD 10b stores programs and databases for operating the functions shown in FIG. 2.

[0085] The processor 10d reads out a program for executing the same processing as each processing unit shown in FIG. 2 from the HDD 10b or the like and expands it in the memory 10c, thereby operating a process for executing each function described in FIG. 2 and the like. For example, this process executes the same functions as each processing unit included in the information processing apparatus 10. Specifically, the processor 10d reads out a program having the same functions as a video acquisition unit 21, a person detection unit 22, a skeleton detection unit 23, a motion estimation unit 24, a range estimation unit 25, an arm stretching action recognition unit 26, an arm stretching position calculation unit 27, etc. from the HDD 10b or the like. Then, the processor 10d executes a process for executing the same processing as the video acquisition unit 21, the person detection unit 22, the skeleton detection unit 23, the motion estimation unit 24, the range estimation unit 25, the arm stretching action recognition unit 26, the arm stretching position calculation unit 27, etc.

[0086] In this way, the information processing apparatus 10 operates as an information processing apparatus that executes an action recognition method by reading and executing a program. Further, the information processing apparatus 10 can also read the program from a recording medium by a medium reader and execute the read program to realize the same functions as those in the above-described embodiments. Note that the program in this other embodiment is not limited to being executed by the information processing apparatus 10. For example, the present invention can be similarly applied when another computer or server executes the program, or when these cooperate to execute the program.

[0087] This program can be distributed via a network such as the Internet. Also, this program can be recorded on a computer-readable recording medium such as a hard disk, a flexible disk (FD), a CD-ROM, a MO (Magneto-Optical disk), a DVD (Digital Versatile Disc), etc., and can be executed by being read from the recording medium by a computer.

Explanation of Signs

[0088] 10 Information processing apparatus 11 Communication unit 12 Output unit 13 Storage unit 14 Machine learning model DB 15 Target location information DB 20 Control unit 21 Video acquisition unit 22 Person detection unit 23 Skeleton detection unit 24 Motion estimation unit 25 Range estimation unit 26 Arm-reaching action recognition unit 27 Arm-reaching position calculation unit

Claims

1. Cause a computer to detect skeletal information of a person from an image acquired by an imaging device, acquire region information regarding the position of an object located at the position where the person is imaged, estimate an action that the person performs on the object based on the skeletal information of the person and the region information, recognize the behavior of the person based on the distance between the person and the object and the estimated action, execute a process of estimating, as the distance between the person and the object, the range reachable by the hand of the person from the skeletal information of the person and the region information, The process of estimating the range reachable by the hand of the person detects, as the position where the person stands, the center point of both feet specified by the skeletal information, calculates the shortest distance from the center point of both feet to the object, calculates a perpendicular line in the height direction of the object with respect to the shortest distance, and estimates, as the range reachable by the hand of the person, the position of the perpendicular line corresponding to the height from the foot to the shoulder specified by the skeletal information with respect to the perpendicular line. An action recognition program characterized by this.

2. The process of estimating estimates an action that the person performs on the object based on whether any one of the conditions of an action of stretching a hand toward the object, an action of looking at the object, or an action of facing the object is satisfied, The process of recognizing recognizes that the person stretches a hand toward the object when the distance between the person and the object is less than a threshold value and an action that satisfies any one of the conditions is estimated. The action recognition program according to claim 1, characterized by this.

3. The process of recognizing recognizes that the person stretches a hand toward the object when the distance between the person and the object is less than a threshold value and a plurality of pre-specified actions are estimated in a pre-specified order within a predetermined time. The action recognition program according to claim 2, characterized by this.

4. The process of recognizing recognizes an action of the person stretching a hand toward the object using the estimated action and the range reachable by the hand of the person. The action recognition program according to any one of claims 1 to 3, characterized by this.

5. The process of estimating the range reachable by the hand of the person calculates the shoulder width of the person based on the skeletal information, and estimates, as the range reachable by the hand of the person, the region corresponding to the shoulder width at the position of the perpendicular line corresponding to the height from the foot to the shoulder specified by the skeletal information. The action recognition program according to claim 1, characterized in that...

6. When the action that the person reaches out to the object is recognized, based on the skeleton information and the region information, calculate the position of the hand, Based on the positional relationship between the position of the hand and the object, cause the computer to execute a process of specifying the object accessed by the hand, the action recognition program according to any one of claims 1 to 5.

7. Cause the computer to execute a process of estimating the attributes of the person based on the skeleton information, The recognition process is, Output by associating the result of recognizing that the person reaches out to the object with the attributes of the person, The action recognition program according to any one of claims 1 to 6, characterized in that...

8. The computer, Detect the skeleton information of a person from an image acquired by an imaging device, Obtain region information regarding the position of an object at the position where the person is imaged, Based on the skeleton information of the person and the region information, estimate the action that the person performs on the object, Based on the distance between the person and the object and the estimated action, recognize the action of the person, As the distance between the person and the object, execute a process of estimating the range that the hand of the person can reach from the skeleton information of the person and the region information, The process of estimating the range that the hand of the person can reach is, Detect the center point of both feet specified by the skeleton information as the position where the person stands, Calculate the shortest distance from the center point of both feet to the object, Calculate a perpendicular line in the height direction of the object with respect to the shortest distance, Estimate the position of the perpendicular line corresponding to the height from the foot to the shoulder specified by the skeleton information as the range that the hand of the person can reach with respect to the perpendicular line. An action recognition method characterized by...

9. Detect the skeleton information of a person from an image acquired by an imaging device, Obtain region information regarding the position where the person is imaged, the target region, and the position of the object, Based on the skeleton information of the person and the region information, estimate the action that the person performs on the object, Based on the distance between the person and the object and the estimated action, recognize the action of the person, Have a control unit that estimates the range that the hand of the person can reach from the skeleton information of the person and the region information as the distance between the person and the object, The control unit is, Detect the center points of both feet specified by the skeletal information as the position where the person is standing, Calculate the shortest distance from the center points of both feet to the object, Calculate a perpendicular line in the height direction of the object with respect to the shortest distance, Estimate the position of the perpendicular line corresponding to the height from the foot to the shoulder specified by the skeletal information with respect to the perpendicular line as the reachable range of the person's hand. An information processing apparatus characterized by this.

Citation Information

Patent Citations

  • Monitoring method, monitoring device, and monitoring program

    JP2015176227A

  • Customer behavior analysis system and customer behavior analysis method

    JP2019139321A

  • Autonomous store tracking system

    JP2020053019A

  • JPP6854959B