Identification Program, Identification Method, and Information Processing Apparatus
By correcting person region sizes using skeletal information, the system addresses the challenge of obscured individuals in video analysis, enhancing identification and tracking accuracy.
Patent Information
- Application Number
- JP2021082943
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-05-17
AI Technical Summary
Existing video analysis systems struggle with accurate person identification and tracking when at least a part of a person is hidden in the shadow of an obstacle, leading to non-uniform person regions and reduced identification accuracy.
The system corrects the size of person regions based on skeletal information, using techniques like OpenPose and Mask R-CNN, to ensure uniformity and accuracy in identifying and tracking individuals, even when partial body parts are obscured by obstacles.
This approach enhances the accuracy of person identification and tracking by ensuring consistent region sizes, improving the reliability of similarity evaluations and reducing errors in identifying the same person across frames.
Smart Images

Figure 0007707643000001 
Figure 0007707643000002 
Figure 0007707643000003
Abstract
Description
[Technical field]
[0001] The present invention relates to an identification program, an identification method, and an information processing device. [Background technology]
[0002] In recent years, the retail industry has been required to automate and streamline store operations due to, for example, changes in lifestyles and labor shortages. For example, analysis of purchasing behavior and suspicious behavior using video data captured in stores by surveillance cameras has been considered.
[0003] For example, attempts are being made to analyze consumer shopping behavior in a store from video data and use the results to acquire new customers and improve the efficiency of store operations. In addition, attempts are being made to use the results of behavioral analysis using video data to detect suspicious behavior, such as leaving a store without scanning items in stores that use unmanned registers.
[0004] In this regard, techniques relating to the detection and identification of people from an image are known (for example, Patent Documents 1 to 3). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2010-239992 A [Patent Document 2] JP 2019-144830 A [Patent Document 3] JP 2020-198053 A Summary of the Invention [Problem to be solved by the invention]
[0006] When analyzing the actions of a person shown in a video, for example, the person is detected from the video, the person is identified, and the movement is tracked. However, at least a part of the person shown in the video may be hidden in the shadow of an obstacle, and as a result, the person identification and tracking may fail.
[0007] On one aspect, the present invention aims to perform person identification with high accuracy even when at least a part of the person shown in the video is hidden in the shadow of an obstacle.
Means for Solving the Problem
[0008] The identification program according to one aspect of the present invention detects a person area from a plurality of frame images included in video data The process to be performed, Performs skeleton detection on the person shown in the person area detected from the plurality of frame images to obtain information on the skeleton of the person The process to be performed, Corrects the person area according to the full body size of the person based on the information on the skeleton of the person to generate a person area of full body size, and corrects the person area according to the half body size of the person based on the information on the skeleton of the person to generate a person area of half body size The process, , Among the plurality of frame images, in a first person region detected from a first frame image and a second person region detected from a second frame image, At least one of Both, The similarity is When a predetermined condition is satisfied, assuming that the person shown in the first person region and the person shown in the second person region are the same person, the first and second person regions are regarded as person regions of the same person, Specify The process, , and cause a computer to execute The specifying process evaluates the similarity between the first person region and the second person region based on the similarity between the half-body-sized person region corresponding to the first person region and the half-body-sized person region corresponding to the second person region, and the similarity between the full-body-sized person region corresponding to the first person region and the full-body-sized person region corresponding to the second person region, when at least one of the first person region detected from the first frame image and the second person region detected from the second frame image is hidden by an obstacle that shields a part of the object behind so that a part of the object behind can be seen.
Effect of the Invention
[0009] It is possible to perform person identification with high accuracy even when at least a part of the person shown in the video is hidden in the shadow of an obstacle.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
MODE FOR CARRYING OUT THE INVENTION
[0011] Hereinafter, some embodiments of the present invention will be described in detail with reference to the drawings. In the plurality of drawings, the corresponding elements are denoted by the same reference numerals.
[0012] FIG. 1 is a diagram illustrating a motion detection system 100 according to an embodiment. The motion detection system 100 may include, for example, an information processing apparatus 101 and a photographing apparatus 102. The information processing apparatus 101 may be a computer having a computing function such as a server computer, a personal computer (PC), a mobile PC, or a tablet terminal.
[0013] The photographing apparatus 102 is, for example, a camera that generates video data. In one example, the photographing apparatus 102 may be, for example, a surveillance camera installed in a store in the retail industry or the like, and may photograph a person who has visited the store or the like and generate video data. Then, the information processing apparatus 101 may perform motion detection of a person on the video data photographed by the photographing apparatus 102. Note that a plurality of photographing apparatuses 102 may be installed in a store, for example. Further, the information processing apparatus 101 may receive video data from the photographing apparatus 102, or may acquire the video data photographed by the photographing apparatus 102 via another apparatus.
[0014] FIG. 2 is a diagram illustrating a block configuration of the information processing apparatus 101 according to an embodiment. The information processing apparatus 101 includes, for example, a control unit 201, a storage unit 202, and a communication unit 203. The control unit 201 includes, for example, a detection unit 211, an acquisition unit 212, a generation unit 213, a specification unit 214, and the like, and may include other functional units. The storage unit 202 of the information processing apparatus 101 stores, for example, video data photographed by the photographing apparatus 102 and information such as obstacle information 800. The communication unit 203 communicates with other apparatuses according to an instruction from the control unit 201. For example, the control unit 201 may acquire video data from the photographing apparatus 102 via the communication unit 203. Details of each of these units and details of the information stored in the storage unit 202 will be described later.
[0015] As described above, when analyzing the actions of a person shown in a video, for example, tracking the movement of the person shown in the video is performed. In tracking the movement of a person, the control unit 201 executes, for example, object detection of the person for the image of each frame of the video. In one example, the control unit 201 may detect a human region from the image of each frame of the video using a machine learning-based technique such as deep learning. For example, the object detection of the person may be performed using techniques such as SSD (Single Shot MultiBox Detector), YOLO (You Only Look Once), and R-CNN (Region Convolutional Neural Network). Also, in one example, the human region is a bounding box.
[0016] Then, the control unit 201 may, for example, identify the human region of the same person using an identical person determination model (for example, Person re-identification) or the like in a plurality of temporally consecutive frame images of the video, and track the movement of the person. Note that the identification of the same person may be performed, for example, by obtaining an evaluation value indicating a similarity or a distance measure indicating how similar the human region detected from a certain frame image is to the comparison target human region detected from another frame image. In one example, the control unit 201 identifies, based on the evaluation value, the human regions that satisfy a predetermined condition and are similar in a certain frame image and another frame image as the human regions of the same person, and may generate a track representing the trajectory of the human region in time series.
[0017] For example, the control unit 201 may input the person area of a certain frame image and the person area of another frame image into the same person determination model. Note that the same person determination model may be generated by learning with deep learning so as to output the similarity (or distance measure) of the images of two person areas. Then, the control unit 201 calculates, for example, the similarity between the person area of a certain frame image and the person area of another frame image using the same person determination model. Subsequently, the control unit 201 may generate a track indicating the movement trajectory of the person by associating the person areas of the same person through weighted matching such as the Hungarian method using the similarity.
[0018] However, for example, at least a part of the person shown in the video may be hidden behind an obstacle, and the range of the person included in the generated person area may become non-uniform. As a result, the accuracy of person identification and tracking may decrease.
[0019] FIG. 3 is a diagram showing a frame image of video data captured by the exemplary imaging device 102. As shown in FIG. 3, a plurality of persons are shown in the frame image. Also, among the persons shown in the frame image, there are persons whose entire body is shown, and there are also persons whose body part is hidden by obstacles such as a shelf and a cart. Then, for example, when person object detection is performed on such a frame image, a person area surrounding various areas of the person may be detected.
[0020] FIG. 4 is a diagram showing the detection result of the exemplary person area 401. In FIG. 4(a), since the entire body of the person is visible, the person area 401 is set to surround the entire body of the person.
[0021] On the other hand, in FIGS. 4(b) and 4(c), a part of the person's body is hidden by the shelf 402, and as a result, the person area 401 is set for the upper body of the person. Note that the size of the person area 401 set for the person may differ depending on the visible range of the person's body, the person's posture, etc. For example, in FIGS. 4(b) and 4(c), person areas 401 of different sizes are set.
[0022] Also, in FIGS. 4(d) and 4(e), a part of the person's body is hidden by the cart 403. The cart 403 has a framework structure, and a part of the person's body can be seen through the cart 403. For example, in this way, it is assumed that the person's body is not completely shielded, but is hidden by an obstacle that shields in such a way that a part of the object behind can be seen. In this case, depending on how the person's body appears through the obstacle, the size of the person area 401 set for the person may change. For example, in FIG. 4(d), the person area 401 is set for the entire body of the person, while in FIG. 4(e), the person area 401 is set for the upper body of the person.
[0023] And if the person area 401 is set unevenly for various ranges of the person's body in this way, when identifying the same person between frame images, even if it is the same person, differences may appear in the feature amounts, and the identification may fail.
[0024] FIG. 5 is a diagram showing an exemplary person identification. FIG. 5(a) exemplifies the evaluation of the similarity of the person areas 401 of the same person detected in a certain frame image i and another frame image i+1. For example, the control unit 201 inputs two images, the image of the person area 401 detected in a certain frame image i and the image of the person area 401 detected in another frame image i+1 to be compared, into the same person determination model. The same person determination model generates a feature amount vector representing the features of the person shown in each image. Note that the same person determination model may be executed by, for example, the control unit 201. And the same person determination model may determine that the persons shown in the two person areas 401 are the same person when the similarity of the obtained feature amount vectors satisfies a predetermined condition and is similar.
[0025] In the example of Fig. 5(a), a person region 401 that includes the entire body is detected in both a certain frame image i and another frame image i + 1. Therefore, when extracting a feature vector that characterizes a person from the images within the person region 401, similar feature vectors can be obtained if they are of the same person. Accordingly, it is possible to associate person regions 401 with a high degree of similarity of feature vectors between a certain frame image i and another frame image i + 1 as the same person.
[0026] However, for example, when a part of a person is hidden by an obstacle such as a cart, it may be possible to detect the entire body as the person region 401, or it may be possible to detect only a part of the body as the person region 401. For example, in Fig. 5(b), even though there is a cart in a certain frame image i, the entire body can be detected as the person region 401, but in another frame image i + 1, due to the body being hidden by the cart, the person region 401 is detected for only a part of the body. For example, even when the person region 401 is detected for only a part of the body in this way, since the feature vector is generated based on the person region 401, there may be differences in the feature vectors even for the same person. As a result, the similarity between the two person regions may decrease even for the same person, and tracking of the same person may not work well. Thus, the non-uniformity of the range of the person included in the person region 401 may have an adverse effect on person identification.
[0027] Therefore, in the embodiment described below, the control unit 201 corrects the size of the person region based on the skeletal information of the person. For example, when the inside of a store is photographed by the photographing device 102, the photographing device 102 is installed near the ceiling of the store and often photographs the inside of the store obliquely downward from above. Also, for example, obstacles inside the store such as shelves and carts often hide the lower body of a person, and the upper body of the person can often be photographed.
[0028] In this case, for example, by performing person detection and skeleton detection on a frame image in which the upper body is shown, it is possible to obtain the skeleton information of the upper body of the person. Then, from the skeleton information of the upper body, it is possible to correct the size of the person region so that a predetermined range of the body of the upper body is included. Also, for example, from the skeleton information of the upper body, it is also possible to correct the size of the person region so that the entire body of the person is included.
[0029] And in this way, by correcting the size of the person region so that the person region surrounds a predetermined range of the body of the person, it becomes possible to make the range of the body of the person included in the person region uniform and evaluate the similarity. Therefore, it is possible to improve the accuracy of person identification and person tracking.
[0030] Hereinafter, the embodiments will be described in more detail. In the embodiment, the control unit 201 corrects the size of the person region based on the skeleton information detected from the person region.
[0031] [Correction of Person Region] The control unit 201 can identify the positions of the skeleton such as the position of the waist and the position of the head of the person, for example, by performing skeleton detection on the detected person region. Note that the skeleton detection may be performed using methods such as OpenPose, Mask R-CNN, DeepPose, PoseNet, etc.
[0032] FIG. 6 is a diagram illustrating the correction of the size of the person region according to the embodiment. In FIG. 6(a), an example is shown in which the lower body of the person is hidden behind the cart, and the person region is detected in a region including the upper body and a part of the lower body. Also, the information on the positions of the skeleton as a result of performing skeleton detection on the person region is indicated by circles, and from the information on the positions of the skeleton, the position 601 of the head and the position 602 of the waist can be identified.
[0033] In the embodiment, the control unit 201 corrects the size of the human region based on the detected position of the skeleton. For example, the length of the upper body of a person in the height direction is defined as the length from the position 601 of the head to the position 602 of the waist. In this case, for example, as shown in FIG. 6(b), by correcting the size of the human region so that the length of the person's body in the height direction is within the range of the upper body from the position 601 of the head to the position 602 of the waist, the human region can be set according to the size of the upper body of the person. Thereby, the range of the person's body included in the human region can be made uniform, and the comparison accuracy of the human region can be improved. Note that the height direction of the person's body may be, for example, the direction of height.
[0034] Further, the control unit 201 may correct the size of the human region so that the human region is set for the whole body. For example, it is estimated that the length of the upper body and the length of the lower body are approximately equal. Therefore, for example, the control unit 201 obtains the length from the position 601 of the head to the position 602 of the waist as the length of the upper body, and doubles that length to extend the human region in the direction of the lower body, so that a human region with a size corresponding to the whole body of the person may be set. For example, as shown in FIG. 6(c), by setting the size of the human region to twice the size of the range from the position 601 of the head to the position 602 of the waist and extending the human region in the direction of the lower body of the person, the human region can be set according to the size of the whole body of the person.
[0035] As described above, the control unit 201 can correct the human region based on the skeleton information so that the range of the person's body included in the human region is within a predetermined range such as the upper body and the whole body. Thereby, the range of the person's body included in the human region used for person identification can be made uniform, and the accuracy of person identification can be improved.
[0036] Note that the correction of the human region is not limited to the upper body and the whole body. The control unit 201 may correct the size of the human region for the person based on the skeleton information so as to surround other ranges. For example, in another embodiment, the size of the human region may be set according to the size of other body parts such as the lower body instead of the upper body.
[0037] In the embodiment, the control unit 201 may, for example, evaluate the similarity using the corrected human area and perform person identification. In one example, the control unit 201 may identify a human area of the same person as a certain human area based on the human area of the whole body size obtained by correction and the human area of the upper body size.
[0038] In this case, the control unit 201 may determine the corrected human area used for the similarity evaluation according to the relationship between the two human areas for which the similarity is evaluated and the obstacle.
[0039] FIG. 7 is a diagram illustrating evaluation target information 700 according to the embodiment. In the evaluation target information 700, information for designating two human areas for which the similarity is evaluated and the corrected human area used for the similarity evaluation according to the relationship with the obstacle is registered. In the evaluation target information 700, the query human area may be, for example, a human area in which a certain person is reflected among a plurality of human areas detected from a certain frame image. Here, a certain frame image may be called, for example, the first frame image. Also, the query human area may be called, for example, the first human area.
[0040] The comparison target human area may be a comparison target human area for evaluating the similarity with a human area in which a certain person is reflected. The comparison target human area may be, for example, one human area among a plurality of human areas detected from another frame image. Here, another frame image may be called, for example, the second frame image. Also, the comparison target human area may be called, for example, the second human area.
[0041] Then, for example, the similarity between the two human areas, that is, the query human area and the comparison target human area, is evaluated. Here, when neither of the two human areas is hidden by an obstacle and the whole body of both is visible (without an obstacle in FIG. 7), the control unit 201 may evaluate the similarity using the human area of the whole body size as shown in the evaluation target information 700.
[0042] On one hand, for example, assume that at least one person area is hidden by an obstacle through which a person can be seen partially, such as a cart. In this case, information about the person who can be seen partially through the obstacle may be useful for identifying the same person. On the other hand, depending on how the area that can be seen partially through the obstacle appears, it may be preferable to evaluate the similarity using only the upper body without including the lower body. Therefore, when a person is hidden by an obstacle that shields a part of an object behind, such as a cart, so that a part of the object can be seen, the control unit 201 may evaluate the similarity using both the whole body and the partial body (e.g., the upper body) of the person area.
[0043] Also, for example, when at least one person area is hidden by an obstacle that completely shields an object behind, such as a shelf, the control unit 201 may evaluate the similarity using only the partial body (e.g., the upper body) of the person area.
[0044] In this way, by setting the person area used for evaluating the similarity according to the relationship between the person area serving as the query, the person area of the comparison target, and the obstacle, it is possible to evaluate the similarity using the person areas that include the same area of the person. Therefore, the accuracy of identifying the same person can be improved.
[0045] Also, for example, in the embodiment, when a person is hidden by an obstacle that shields a part of an object behind so that a part of the object can be seen, the similarity may be evaluated using both the whole body and the partial body (e.g., the upper body) of the person area. Thereby, the accuracy of identifying the same person can be improved.
[0046] Hereinafter, an example of the same person identification process according to the embodiment will be described with reference to FIGS. 8 to 15.
[0047] FIG. 8 is a diagram illustrating obstacle information 800 according to an embodiment. Information indicating an area where an obstacle appears in the shooting range of the imaging device 102 may be registered in the obstacle information 800. In the example of FIG. 8, in the obstacle information 800, for example, an obstacle ID (identifier) for identifying an obstacle and information indicating an area within the frame image in which the obstacle appears are registered in association with each other. Note that, for example, information about obstacles that are rarely moved, such as shelves, may be registered in the obstacle information 800. Further, information such as a front-back relationship indicating whether it is in front of or behind another area such as a passageway may be registered in the obstacle information 800.
[0048] Note that, for example, for an obstacle accompanied by movement such as a cart, the control unit 201 may specify the position of the obstacle using another method such as object detection.
[0049] FIG. 9 is a diagram showing two exemplary frame images. FIG. 9 shows two frame images, namely, a certain frame image i (FIG. 9(a)) included in the video data and another frame image i+1 (FIG. 9(b)). The control unit 201 may perform object detection of a person on these two frame images and detect a person area.
[0050] FIG. 10 shows the detected person areas as a result of performing object detection of a person on the two frame images of FIG. 9. In FIG. 10(a), the control unit 201 assigns IDs a to d to the person areas in order to identify each person area. In FIG. 10(b), the control unit 201 assigns IDs 1 to 4 to the person areas in order to identify each person area.
[0051] Note that, in the example of FIG. 10, the frame image includes a person whose part of the body is hidden by obstacles such as a shelf and a cart. As a result, person regions are set for various ranges of the person's body, such as the whole body of the person or the upper body of the person. For example, in a certain frame image i of FIG. 10(a), for the persons with IDs = b and d, since their lower bodies are hidden by the shelf or the cart, person regions are detected for their upper bodies. Also, for example, in another frame image i+1 of FIG. 10(b), for the person with ID = 4, the lower body is hidden by the shelf and not completely visible, and a person region is detected for the upper body.
[0052] For example, as described above, when the ranges of the persons' bodies included in the detected person regions are non-uniform, even if an attempt is made to evaluate the similarity and identify the same person as it is, there may be a case where a wrong person is identified as the same person. Therefore, in the embodiment, based on the information on the skeletons of the persons included in the person regions, the range in which the person regions are set for the persons' bodies is corrected.
[0053] FIG. 11 is an example in which the size of the person region is corrected based on the skeleton information so that the whole body of the person is included in the person region. Also, FIG. 12 is an example in which the size of the person region is corrected based on the skeleton information so that the upper body of the person is included in the person region.
[0054] Then, as shown in the evaluation target information 700, the control unit 201 obtains the similarity using at least one of the person region with the whole body size and the person region with the upper body size based on the relationship between the two person regions to be evaluated for similarity and the obstacles.
[0055] FIG. 13 is a diagram showing an example of similarity evaluation according to an embodiment. In FIG. 13(a), the control unit 201 uses the person area: a of a certain frame image i as the query person area, and obtains the similarity with each person area of ID: 1 to ID: 4 of another frame image i+1 as the comparison target person area. Note that the query person area: a does not overlap with the position of the obstacle, and the whole body of the person is included in the person area. On the other hand, the person area of ID: 1 as the comparison target also does not have an overlap with the obstacle. Therefore, as shown in the evaluation target information 700, the control unit 201 may evaluate the similarity between the two person areas to be evaluated using the person areas of the whole body size generated in FIG. 11.
[0056] Also, since the person area of ID: 2 as the comparison target does not have an overlap with the obstacle, the control unit 201 may similarly evaluate the similarity using the person areas of the whole body size generated in FIG. 11.
[0057] Also, the person area of ID: 3 as the comparison target has an overlap with the cart obstacle. Therefore, as shown in the evaluation target information 700, the control unit 201 may evaluate the similarity using both the similarity between the person areas of the whole body size generated in FIG. 11 and the similarity between the person areas of the upper body size generated in FIG. 12. In one example, as shown in FIG. 13(a), the control unit 201 may use the value obtained by averaging the similarity between the person areas of the whole body size and the similarity between the person areas of the upper body size as the similarity between the query person area: a and the comparison target person area: 3.
[0058] Also, the person area of ID: 4 as the comparison target has an overlap with the shelf obstacle. Therefore, the control unit 201 may evaluate the similarity using the person areas of the upper body size generated in FIG. 12.
[0059] Similarly, the similarity can be evaluated for the query person area: b to obtain the similarity in FIG. 13(b). The similarity can be evaluated for the query person area: c to obtain the similarity in FIG. 13(c). The similarity can be evaluated for the query person area: d to obtain the similarity in FIG. 13(d).
[0060] Then, by performing weighted matching using the similarity between the obtained query person region and the comparison target person region, the person region of the same person as the person region detected in a certain frame image i can be specified from among the person regions detected in another frame image i + 1.
[0061] FIG. 14 is a diagram illustrating person regions of the same person linked using weighted matching. As shown in FIG. 14, in the person regions of the same person linked using weighted matching, the correct ID can be specified. In this way, by correcting the size of the person region so that the same range of the person's body is included in the query person region and the comparison target person region, and evaluating the similarity, the identification accuracy of the same person can be improved.
[0062] Also, when evaluating the similarity using only the upper body, the information is reduced by the amount of information in the lower body. In this case, for example, if the upper body clothing happens to be similar, there is a risk that the identification accuracy of the same person will decrease. Also, for example, in the case of an obstacle such as a cart that does not completely block the person's figure but hides part of the body by blocking so that a part of the object behind can be seen, using the information in the visible area may also improve the identification accuracy of the same person. In the above-described embodiment, when the person's body is hidden by an obstacle that blocks so that a part of the object behind, such as a cart, can be seen, the similarity is evaluated by integrating the similarity evaluated using the person region of the full body size and the similarity evaluated using the upper body size. Therefore, the identification accuracy of the same person can be improved.
[0063] As described above, according to the embodiment, even when at least a part of the person shown in the video is hidden in the shadow of an obstacle, person identification can be performed with high accuracy.
[0064] FIG. 15 is a diagram illustrating the operation flow of the person movement detection process according to the embodiment. For example, when an execution instruction for the person movement detection process is input, the control unit 201 may start the process of the operation flow in FIG. 15.
[0065] In step 1501 (hereinafter, steps are described as "S", for example, denoted as S1501), the control unit 201 executes, for example, person detection. The control unit 201 may execute person detection on the image of each frame of the captured data captured by the imaging device 102 to identify a person area from the image. Note that the person detection may be executed using techniques such as SSD (Single Shot MultiBox Detector), YOLO (You Only Look Once), and R-CNN (Region Convolutional Neural Network).
[0066] In S1502, the control unit 201 executes skeleton detection on the person area detected from the image of each frame. For example, the control unit 201 may detect the skeleton of the person shown in the person area using methods such as OpenPose, Mask R-CNN, DeepPose, and PoseNet.
[0067] In S1503, the control unit 201 corrects the detected person area. For example, as illustrated in FIG. 6, the control unit 201 may correct the person area to match the upper body of the person or correct the person area to match the whole body of the person based on the information of the skeleton of the person detected from the person area.
[0068] In S1504, the control unit 201 identifies the position of the obstacle. For example, for an obstacle fixed at a predetermined position such as a shelf, the control unit 201 may acquire the position of the obstacle from the information registered in the obstacle information 800.
[0069] Also, for example, for a moving obstacle such as a cart, the control unit 201 may acquire the position of the obstacle using a technique such as object detection. For object detection, techniques such as SSD, YOLO, and R-CNN may be used. For example, the control unit 201 may acquire the area where the obstacle appears from each frame image of the video data using a learned model that is machine-learned to detect a detection target obstacle such as a cart.
[0070] In S1505, the control unit 201 identifies the presence or absence of an obstacle with respect to the human area. For example, when the human area detected in S1501 is at a position hidden by an obstacle in front of it, the control unit 201 may determine that the person in the human area is hidden by the obstacle. On the other hand, when the human area detected in S1501 is not at a position hidden by an obstacle in front of it, the control unit 201 may determine that the person in the human area is not hidden by the obstacle.
[0071] In S1506, the control unit 201 evaluates the similarity between the human area of a certain frame image and the human area of another frame image according to the presence or absence of the identified obstacle. For example, as shown in the evaluation target information 700, the control unit 201 may evaluate the similarity between the human areas not hidden by the obstacle by using the human area of the whole body size. Further, when at least one of the two human areas to be evaluated for similarity is hidden by an obstacle that completely hides the person behind an obstacle such as a shelf, the control unit 201 may evaluate the similarity by using the upper body size human area. Also, for example, assume that at least one of the two human areas to be evaluated for similarity is hidden by an obstacle where a part of the person behind an obstacle such as a cart is visible. In this case, the control unit 201 may evaluate the similarity using both the whole body size human area and the upper body size human area, and integrate them by averaging the evaluation values, etc., and use the integrated value as the similarity.
[0072] In S1507, the control unit 201 associates the human areas of the same person by weighted matching based on the similarity, thereby performing tracking of the person, and this operation flow ends.
[0073] As described above, according to the embodiment, even when at least a part of the person shown in the video is hidden in the shadow of an obstacle, person identification can be performed with high accuracy.
[0074] For example, the control unit 201 can improve the specific accuracy of the same person because it corrects the size of the human area so that the same area of the person's body is included in the human area of the query and the human area of the comparison target, and evaluates the similarity.
[0075] In addition, for example, when a person's body is hidden by an obstacle that shields a part of an object behind, such as a cart, so that a part of the object behind can be seen, the control unit 201 integrates the evaluation results of the similarity evaluated in each person region of the whole body and the upper body to obtain the similarity. Thereby, the identification accuracy of the same person can be improved.
[0076] Also, in the above-described embodiment, depending on at least one of the two person regions that are the evaluation targets of the similarity, or depending on what kind of obstacle is hiding it, the person region used for the evaluation of the similarity is changed. Therefore, the similarity can be evaluated with high accuracy according to the appearance of the person, and the identification accuracy of the same person can be improved.
[0077] Also, in the above-described embodiment, based on the skeleton information, the size of the person region is corrected according to the size of the upper body and the whole body of the person. Therefore, the size of the person region can be set with high accuracy according to the size of the person's body.
[0078] In the above, the embodiments have been illustrated, but the embodiments are not limited to this. For example, the above-described operation flow is an illustration, and the embodiments are not limited to this. If possible, the operation flow may be executed by changing the order of processing, may include additional processing separately, or some processing may be omitted.
[0079] For example, in one embodiment, when identifying the same person from video data in an environment where the lower body of the person is frequently hidden, the control unit 201 may use only the upper body of the person to evaluate the similarity and execute the identification of the same person. In this case, in the process of S1503, the control unit 201 may correct the person region according to the upper body based on the skeleton information, and the processes of S1504 and S1505 may be omitted. Also, in another embodiment, after the operation flow of FIG. 15, processing such as the behavior analysis of the person may be further executed on the track that tracks the movement of the person determined to be the same person.
[0080] In addition, in the above-described embodiment, an example of using similarity to determine whether two human regions are of the same person has been described, but the embodiment is not limited thereto. For example, in another embodiment, the control unit 201 may determine whether the human regions are of the same person based on a distance metric.
[0081] In addition, in the above-described embodiment, as an example of the half body of a person for adjusting the size of the human region, the upper body has been described as an example, but the embodiment is not limited thereto. In another embodiment, the control unit 201 may correct the size of the human region according to other predetermined body regions such as the lower body of the person, and use it for the evaluation of similarity.
[0082] Also, for example, when the whole body of a person is visible in two human regions to be compared during person identification, the determination result of the similarity is presumed to be more reliable than when a part of the person's body is not visible due to an obstacle. Similarly, when a person is hidden by an obstacle that completely hides the person on the back side of the obstacle, the case where the person is hidden by an obstacle that allows a part of the object behind to be seen has more information and the similarity is presumed to be more reliable. Therefore, when the control unit 201 associates human regions of the same person between frame images by weighted matching or the like in S1507, as shown below, the similarity evaluated using the whole body may be prioritized over other similarities to associate the human regions. Similarly, when the control unit 201 associates human regions of the same person between frame images by weighted matching or the like in S1507, the similarity evaluated using the whole body + upper body human region may be prioritized over the similarity evaluated using only the upper body to associate the human regions. Whole body Priority: High Whole body + upper body Priority: Medium Upper body only Priority: Low
[0083] In the above-described embodiment, for example, in the process of S1501, the control unit 201 operates as the detection unit 211. Also, for example, in the process of S1502, the control unit 201 operates as the acquisition unit 212. In the process of S1503, the control unit 201 operates as the generation unit 213. In the processes of S1504 to S1507, the control unit 201 operates as the specifying unit 214.
[0084] FIG. 16 is a diagram illustrating a hardware configuration of a computer 1600 for realizing the information processing apparatus 101 according to the embodiment. The hardware configuration for realizing the information processing apparatus 101 in FIG. 16 includes, for example, a processor 1601, a memory 1602, a storage device 1603, a reading device 1604, a communication interface 1606, and an input / output interface 1607. Note that the processor 1601, the memory 1602, the storage device 1603, the reading device 1604, the communication interface 1606, and the input / output interface 1607 are connected to each other via, for example, a bus 1608.
[0085] The processor 1601 may be, for example, a single processor, a multi-processor, or a multi-core. The processor 1601 provides some or all of the functions of the above-described control unit 201 by executing a program that describes the above-described operation flow procedure using the memory 1602. For example, the processor 1601 of the information processing apparatus 101 operates as the detection unit 211, the acquisition unit 212, the generation unit 213, and the specifying unit 214 by reading and executing a program stored in the storage device 1603.
[0086] The memory 1602 is, for example, a semiconductor memory and may include a RAM area and a ROM area. The storage device 1603 is, for example, a semiconductor memory such as a hard disk or a flash memory, or an external storage device. Note that RAM is an abbreviation for Random Access Memory. Also, ROM is an abbreviation for Read Only Memory.
[0087] The reading device 1604 accesses the removable storage medium 1605 according to the instruction of the processor 1601. The removable storage medium 1605 is realized by, for example, a semiconductor device, a medium in which information is input and output by magnetic action, a medium in which information is input and output by optical action, and the like. Note that the semiconductor device is, for example, a USB (Universal Serial Bus) memory. In addition, the medium in which information is input and output by magnetic action is, for example, a magnetic disk. The medium in which information is input and output by optical action is, for example, a CD-ROM, a DVD, a Blu-ray Disc, etc. (Blu-ray is a registered trademark). CD is an abbreviation for Compact Disc. DVD is an abbreviation for Digital Versatile Disk.
[0088] The storage unit 202 includes, for example, the memory 1602, the storage device 1603, and the removable storage medium 1605. For example, the storage device 1603 of the information processing apparatus 101 stores, for example, video data captured by the imaging device 102 and the obstacle information 800.
[0089] The communication interface 1606 communicates with other devices such as the imaging device 102 according to the instruction of the processor 1601, for example. The communication interface 1606 is an example of the communication unit 203 described above.
[0090] The input / output interface 1607 may be an interface between an input device and an output device, for example. The input device is a device such as a keyboard, a mouse, or a touch panel that receives an instruction from a user, for example. The output device is a display device such as a display and an audio device such as a speaker, for example.
[0091] Each program according to the embodiment is provided to the information processing apparatus 101 in the following form, for example. (1) It is pre-installed in the storage device 1603. (2) It is provided by the removable storage medium 1605. (3) It is provided from a server such as a program server.
[0092] Note that the hardware configuration of the computer 1600 for realizing the information processing apparatus 101 described with reference to FIG. 16 is an example, and the embodiments are not limited thereto. For example, a part of the above-described configuration may be deleted, or a new configuration may be added. Further, in another embodiment, for example, some or all of the functions of the above-described control unit 201 may be implemented as hardware by an FPGA, an SoC, an ASIC, a PLD, or the like. Note that FPGA is an abbreviation for Field Programmable Gate Array. SoC is an abbreviation for System-on-a-chip. ASIC is an abbreviation for Application Specific Integrated Circuit. PLD is an abbreviation for Programmable Logic Device.
[0093] As described above, several embodiments have been described. However, the embodiments are not limited to the above-described embodiments, and should be understood to include various modifications and alternative forms of the above-described embodiments. For example, it will be understood that each of the various embodiments can be embodied by modifying the components without departing from the spirit and scope thereof. Also, it will be understood that various embodiments can be implemented by appropriately combining a plurality of components disclosed in the foregoing embodiments. Furthermore, it will be understood by those skilled in the art that various embodiments can be implemented by deleting some components from all the components shown in the embodiments or adding some components to the components shown in the embodiments.
Explanation of Signs
[0094] 100 Motion detection system 101 Information processing apparatus 102 Imaging apparatus 201 Control unit 202 Storage unit 203 Communication unit 211 Detection unit 212 Acquisition unit 213 Generation unit 214 Specific part 401 Human area 402 Shelf 403 Cart 700 Information to be evaluated 800 Obstacle information 1600 Computer 1601 Processor 1602 Memory 1603 Storage device 1604 Reading device 1605 Removable storage medium 1606 Communication interface 1607 Input / output interface 1608 Bus
Claims
1. A process of detecting a human region from a plurality of frame images included in video data, a process of performing skeleton detection on a person shown in the human region detected from the plurality of frame images to obtain information on the skeleton of the person, a process of correcting the human region according to the size of the whole body of the person based on the information on the skeleton of the person to generate a human region of the whole body size, and a process of correcting the human region according to the size of the half body of the person based on the information on the skeleton of the person to generate a human region of the half body size, a process of specifying the first and second human regions as human regions of the same person on the assumption that the person shown in the first human region and the person shown in the second human region are the same person when at least one of the similarities of the human region of the whole body size and the human region of the half body size between the first human region detected from the first frame image among the plurality of frame images and the second human region detected from the second frame image satisfies a predetermined condition, and causing a computer to execute the process, The specifying process is an identification program that evaluates the similarity between the first human region and the second human region based on the similarity between the half-body-size human region corresponding to the first human region and the half-body-size human region corresponding to the second human region, and the similarity between the full-body-size human region corresponding to the first human region and the full-body-size human region corresponding to the second human region when, in at least one of the first human region detected from the first frame image and the second human region detected from the second frame image, a person is hidden by an obstacle that shields a part of the object behind so that it can be seen.
2. The specifying process is the identification program according to Claim 1, which evaluates the similarity between the first human region and the second human region based on the similarity between the full-body-size human region corresponding to the first human region and the full-body-size human region corresponding to the second human region when neither the first human region detected from the first frame image nor the second human region detected from the second frame image is hidden by an obstacle.
3. The specific processing is, in at least one of the first human region detected from the first frame image and the second human region detected from the second frame image, when a person is hidden by an obstacle that completely shields an object behind, based on the similarity between the half-body sized human region corresponding to the first human region and the half-body sized human region corresponding to the second human region, evaluate the similarity between the first human region and the second human region. The identification program according to claim 1.
4. The generating process is to set the length in the height direction of the half-body sized human region to be the length from the head to the waist of the person specified based on the information of the skeleton. The identification program according to any one of claims 1 to 3.
5. The generating process is to set the length in the height direction of the full-body sized human region to be twice the length from the head to the waist of the person specified based on the information of the skeleton. The identification program according to any one of claims 1 to 4.
6. An identification method executed by a computer, wherein the computer detects a human region from a plurality of frame images included in video data, performs skeleton detection on the person reflected in the human region detected from the plurality of frame images to obtain information on the skeleton of the person, corrects the human region according to the full-body size of the person based on the information of the skeleton of the person to generate a full-body sized human region, and corrects the human region according to the half-body size of the person based on the information of the skeleton of the person to generate a half-body sized human region, in the first human region detected from the first frame image and the second human region detected from the second frame image among the plurality of frame images, when the similarity of at least one of the full-body sized human region and the half-body sized human region satisfies a predetermined condition, identify the first and second human regions as the human regions of the same person on the assumption that the person reflected in the first human region and the person reflected in the second human region are the same person. In the specific procedure, when a person is hidden by an obstacle that shields at least a part of an object behind the person in at least one of the first person region detected from the first frame image and the second person region detected from the second frame image, the similarity between the first person region and the second person region is evaluated based on the similarity between the half-body sized person region corresponding to the first person region and the half-body sized person region corresponding to the second person region, and the similarity between the full-body sized person region corresponding to the first person region and the full-body sized person region corresponding to the second person region. Identification method.
7. A detection unit that detects a person region from a plurality of frame images included in video data, An acquisition unit that performs skeleton detection on the person reflected in the person region detected from the plurality of frame images to obtain information on the skeleton of the person, A generation unit that corrects the person region according to the full body size of the person based on the information on the skeleton of the person to generate a full body size person region, and corrects the person region according to the half body size of the person based on the information on the skeleton of the person to generate a half body size person region, In the first person region detected from the first frame image and the second person region detected from the second frame image among the plurality of frame images, when the similarity of at least one of the full body size person region and the half body size person region satisfies a predetermined condition, the first person region and the second person region are identified as person regions of the same person. A specifying unit, When a person is hidden by an obstacle that shields at least a part of an object behind the person in at least one of the first person region detected from the first frame image and the second person region detected from the second frame image, the specifying unit evaluates the similarity between the first person region and the second person region based on the similarity between the half-body sized person region corresponding to the first person region and the half-body sized person region corresponding to the second person region, and the similarity between the full-body sized person region corresponding to the first person region and the full-body sized person region corresponding to the second person region. Information processing apparatus.
Citation Information
Patent Citations
Human body image comparison method and device
CN105243395A
Information display device
JP2007127437A
Apparatus, method, and program for detecting object from image
JP2009075868A
Person identification device, person identification method, and person identification program
JP2010239992A
Person detection device
JP2015191338A