Information processing device, information processing method, and information processing program

Using skeleton points from a pose estimation algorithm to calculate ground contact points addresses the inaccuracy of obscured object detection, ensuring precise tracking and path estimation.

JP7803994B2Active Publication Date: 2026-01-21SOFTBANK CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024057474
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2026-01-21
Estimated Expiration
2044-03-29

AI Technical Summary

Technical Problem

Existing object detection models inaccurately estimate ground contact points when parts of a target object are obscured due to training on unobstructed parts, necessitating costly new training datasets.

Method used

Utilize skeleton points from a pose estimation algorithm to accurately calculate ground contact points by determining the height distance between highest and lowest skeletal points in an image.

Benefits of technology

Accurately estimates ground contact points even when parts of the object are hidden, preventing loss of tracking and ensuring correct movement path calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007803994000001
    Figure 0007803994000001
  • Figure 0007803994000002
    Figure 0007803994000002
  • Figure 0007803994000003
    Figure 0007803994000003
Patent Text Reader

Abstract

To achieve object detection with high accuracy by utilizing skeleton points.SOLUTION: An information processing device includes an acquisition part and a calculation part. The acquisition part acquires skeleton points corresponding to a prescribed subject included in an image. The calculation part calculates a ground contact point on a ground surface of the prescribed subject on the basis of a height distance between a first skeleton point at the highest position and a second skeleton point at the lowest position in a coordinate system of the image among the skeleton points.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]

[0002] There are known methods for detecting a specific object (e.g., a person) from an image and tracking the detected object. For example, Patent Document 1 discloses a method for estimating a contact point on a ground surface of a subject and a movement line formed by the movement of the subject. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-236569 Summary of the Invention [Problem to be solved by the invention]

[0004] Existing object detection models used to estimate ground contact points have a problem in that the accuracy of estimating ground contact points decreases when a part of the target object (e.g., a person) is hidden by an obstruction. This problem occurs because existing object detection models are trained to recognize only the unobstructed parts of the target object.

[0005] To improve existing object detection models so that they can accurately detect the positions of occluded parts, a new training dataset is required, and training is costly. Therefore, a method that can accurately estimate ground contact points using existing object detection models is needed.

[0006] For example, there is a pose estimation algorithm that uses an object detection model, and by utilizing the skeleton points estimated by this algorithm, it may be possible to accurately estimate the contact points on the ground surface of the subject. However, the above-mentioned conventional technology does not disclose an object detection method that utilizes skeleton points.

[0007] Therefore, the present invention provides an information processing device, an information processing method, and an information processing program that can utilize skeleton points to accurately estimate contact points. [Means for solving the problem]

[0008] In order to solve the above problem, one form of information processing device according to the present invention includes an acquisition unit that acquires skeleton points corresponding to a specified subject included in an image, and a calculation unit that calculates a ground contact point on the ground surface of the specified subject based on the height distance between a first skeleton point that is located at the highest position in a coordinate system of the image and a second skeleton point that is located at the lowest position among the skeleton points. [Effects of the Invention]

[0009] According to the present invention, the contact point can be estimated with high accuracy by utilizing the skeleton points. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram for explaining the problem underlying the present invention. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of an information processing system according to the embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of the configuration of the information processing device according to the embodiment. [Figure 4] FIG. 4 is a diagram showing an example of a method for determining how a body appears in an image. [Figure 5] FIG. 5 is a diagram showing the contact point calculation method (1). [Figure 6] FIG. 6 is a diagram showing the contact point calculation method (2). [Figure 7]FIG. 7 is a sequence diagram illustrating an example of operation of the information processing system according to the embodiment. [Figure 8] FIG. 8 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of the information processing device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.

[0012] One or more embodiments (including examples, modifications, and application examples) described below can be implemented independently. However, at least a portion of the embodiments described below may be implemented in appropriate combination with at least a portion of another embodiment. These embodiments may include novel features that are different from each other. Therefore, these embodiments may contribute to solving different purposes or problems and may produce different effects from each other.

[0013] Furthermore, in the following embodiments, it is assumed that the subject (detection target) to be detected among the subjects included in an image is a person, and a person detection model is exemplified as an example of an object detection model. However, the proposed technology of the present invention is not limited to systems that detect people, and can also be applied to systems that detect non-people (for example, animals such as dogs and cats). For this reason, in the following embodiments, the expression "part of a person's body" corresponds to "part of a subject," and the expression "the entire body of a person" is an example expression that corresponds to "the entire subject."

[0014] The person detection model here refers to a commonly used existing person detection algorithm, and the position of the person output as the detection result is indicated by a rectangular area, which is represented by position coordinates corresponding to the coordinate system of the image containing the person as a subject.

[0015] (Embodiment) 1. Introduction There is an application that uses a person detection model to predict a rectangular area (detect a person) from an image captured by a fixed camera, estimate the contact point on the person's surface based on the rectangular area, and then perform a projective transformation (bird's-eye view) on the captured image containing information on the contact point to track the same person and record their movement line (path of movement).

[0016] However, in a captured image in which part of a person's body (for example, the feet) is hidden by some object, existing person detection models cannot detect the position coordinates that correctly indicate the person's entire body, and the application may calculate the contact point at an incorrect position. In such a case, the application may lose track of the person whose movement it was originally tracking, and may not be able to obtain a path that correctly tracks the same person.

[0017] To cite a specific example, when a person's body is partially obscured, existing person detection models can only detect the part of the person's body that is visible in the captured image (i.e., the part that is not obscured), i.e., the part that is visible to the viewer of the captured image. In this case, the application calculates the ground contact point at a position different from the intended position (e.g., the position of the part that is visible in the captured image), and when projected onto a bird's-eye view, an incorrect movement path is obtained, in which the ground contact point appears to have moved far away from its actual position. In other words, in a situation where a scene in which a person passes behind some object is captured by a fixed camera, if the person enters behind the object and part of their body is obscured, the application may lose track of the person that it had been tracking and erroneously determine that the movement path is that of a completely different person.

[0018] The above problem will be explained using FIG. 1. FIG. 1 is a diagram for explaining the problem underlying the present invention. FIG. 1 shows a scene in which a fixed camera CA fixed at a predetermined position in a certain space M continuously captures images of a person U moving, and the person U is tracked based on the captured images. In FIG. 1, the problem will be explained by focusing on one captured image of the person U passing directly behind a desk, among the images successively acquired by the continuous capture. In this captured image, part of the person U's body is hidden by an obstruction (the desk).

[0019] For example, when a captured image is input, an existing person detection model (hereinafter referred to as "person detection model M1") detects person U from the input captured image (predicting a rectangular area indicating where person U is located in the captured image). As shown in Figure 1, if part of person U (the lower body) is hidden by an obstruction, person detection model M1 cannot detect the position of the hidden part, and only detects the position of the part that is not hidden and is displayed in the captured image (visible part). As a result, the application will draw a rectangular area AR that surrounds the visible part based on the detected position.

[0020] The application calculates the ground contact point of person U in the captured image based on the rectangular area AR, and so if the rectangular area AR surrounds the entire body of person U, it can calculate the correct position G1 as the ground contact point. On the other hand, if the rectangular area AR surrounds only a part of the body of person U, the application will incorrectly calculate a position G2 different from position G1 as the ground contact point, as shown in Figure 1. For example, the application may calculate a position G2 within the rectangular area AR above position G1 as the ground contact point.

[0021] If an incorrect position is calculated as the touchdown point in this way, the application will not be able to correctly obtain the movement line of person U, and will not be able to track person U as the same person. This point will be explained further using Figure 1.

[0022] The application depicts the movement path of person U in an overhead view obtained by projectively transforming the captured image for which the ground contact point has been calculated. When projectively transforming the captured image for which the correct position G1 has been calculated as the ground contact point, an overhead view is obtained in which an appropriate position within space M is determined to be the position of person U. Specifically, an overhead view is obtained in which the position within space M directly behind the obstruction is the position of person U. As a result, the application can correctly track the movement of person U without losing sight of person U as the camera CA continuously captures the movement of person U.

[0023] On the other hand, if a captured image in which the incorrect position G2 is calculated as the ground contact point is projectively transformed, the resulting bird's-eye view may indicate that the position of a different person is different from person U. Specifically, the resulting bird's-eye view may indicate that the different person is located in another space outside space M, far away from directly behind the obstruction. As a result, while camera CA is continuously capturing images of person U moving, the application may lose track of the person it had been tracking while person U is passing through the obstruction, and may erroneously determine that a different person is moving to a completely different position.

[0024] The reason for such erroneous judgment is thought to be that person detection model M1 is trained to estimate only the position of the visible part of a person's body when a photographic image in which part of the person's body is hidden is input. Specifically, this is thought to be because person detection model M1 is generated based on training data in which a rectangular area surrounding the visible part is labeled as the correct answer area.

[0025] For this reason, if an image having a rectangular area enclosing the entire body of a person is generated based on an image in which part of the person's body is hidden, and that image is used as training data to train person detection model M1, it is possible to generate a new person detection model that can detect the entire body of a person even if an image in which part of the person's body is hidden is input.

[0026] However, it is necessary to generate a new training dataset to improve the person detection model M1, which is time-consuming and costly. Therefore, there is a need for a method that can accurately estimate ground contact points even when using the existing person detection model M1 as is.

[0027] Therefore, the inventors of the present invention focused on a pose estimation algorithm. Here, the pose estimation algorithm is one that estimates the skeletal points of a living organism. Based on this, the inventors of the present invention focused on the fact that by utilizing the skeletal points obtained by the pose estimation algorithm, it is possible to accurately estimate the ground contact points of a person even from an image in which part of the person's body is hidden. That is, in consideration of the accuracy issues inherent in ground contact point calculation algorithms that use a person detection model, the proposed technology of the present invention proposes a ground contact point calculation algorithm that uses a pose estimation model. Details of the pose estimation algorithm will be described later. The pose estimation algorithm may incorporate an object detection model to detect a person from an image as preprocessing for pose estimation, and estimate the skeletal points of the detected person. Alternatively, the pose estimation algorithm may omit the step of detecting a person and directly estimate the skeletal points from the image.

[0028] An information processing device according to an embodiment described in this specification acquires skeleton points corresponding to a specified subject (person) included in an image, and calculates a ground contact point on the ground surface of the specified subject based on the height distance between the first skeleton point that is highest and the second skeleton point that is lowest in the coordinate system of the image among the acquired skeleton points.

[0029] [2. System Configuration Overview] The configuration of the information processing system 1 will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the configuration of the information processing system 1 according to the embodiment. As shown in Fig. 2, the information processing system 1 includes an imaging system 2 and an information processing device 100. The imaging system 2 and the information processing device 100 are connected to each other via a predetermined communication network (network N) so as to be able to communicate with each other via wired or wireless communication. Note that the information processing system 1 shown in Fig. 2 may include a plurality of imaging systems 2 and a plurality of information processing devices 100.

[0030] As shown in FIG. 2, the imaging system 2 may be configured with an imaging device 10, a display control device 11, and a display device 12.

[0031] The imaging device 10 is an imaging device (camera) installed to capture images of a specific indoor space (e.g., inside a store) in order to track the movement of people in the space and estimate their traffic lines. The imaging device 10 may be, for example, an AI camera.

[0032] The display control device 11 superimposes a rectangular area or a grounding point on the captured image acquired by the imaging device 10, and controls the display device 12 to display the superimposed captured image.

[0033] The display device 12 has a screen using, for example, a liquid crystal display, an electroluminescence (EL), a cathode ray tube (CRT), etc. The display device 12 may be compatible with 4K or 8K, or may be formed by a plurality of display devices. The display device 12 displays the captured image controlled to be displayed by the display control device 11.

[0034] The information processing device 100 executes information processing (mainly calculation of ground contact points) according to the proposed technique of the present invention. The information processing device 100 may be implemented as either a local server or a cloud server incorporating a learning function (AI software).

[0035] 3. Configuration of Information Processing Device The information processing device 100 according to the embodiment will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of the configuration of the information processing device 100 according to the embodiment. As shown in Fig. 3, the information processing device 100 includes a communication unit 110, a storage unit 120, and a control unit 130.

[0036] <Communication Unit 110> The communication unit 110 is realized by, for example, a network interface card (NIC), etc. For example, the communication unit 110 transmits and receives information to and from the imaging system 2.

[0037] <Storage section 120> The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 120 may store, for example, data and programs related to the information processing according to the embodiment.

[0038] <Control unit 130> The control unit 130 is realized by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like executing various programs (for example, the information processing program according to the embodiment) stored in a storage device inside the information processing device 100 using RAM as a work area. The control unit 130 is also realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0039] As shown in Fig. 3, the control unit 130 has an estimation unit 131, an acquisition unit 132, a determination unit 133, a calculation unit 134, and a processing unit 135, and realizes or executes the functions and actions of information processing described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in Fig. 3, and may be other configurations as long as they perform the information processing described below. Furthermore, the connection relationship between the processing units included in the control unit 130 is not limited to the connection relationship shown in Fig. 3, and may be other connection relationships.

[0040] In the control unit 130, the acquisition unit 132 acquires skeleton points corresponding to a person included in an image. The calculation unit 134 calculates a ground contact point on the ground surface of the person included in the image based on the height distance between a first skeleton point that is the highest in the coordinate system of the image and a second skeleton point that is the lowest among the skeleton points. Specifically, each time a captured image is obtained by continuous shooting with the imaging device 10, the estimation unit 131 estimates a skeleton point corresponding to the person for each captured image. Therefore, the acquisition unit 132 sequentially acquires the estimated skeleton points in response to the continuous shooting. The calculation unit 134 calculates the height distance for each captured image in real time when continuous shooting is performed, and calculates the ground contact point of the person based on the calculated height distance.

[0041] <Estimation part 131> The estimation unit 131 estimates the posture of a person included in a captured image. The estimation unit 131 estimates skeleton points corresponding to the person using a subject image corresponding to the person in the captured image. In the following embodiment, the subject image will be described as a partial image included in a rectangular area enclosed as a result of person detection in the captured image. On the other hand, the posture estimation algorithm may omit the step of detecting a person and directly estimate skeleton points from the captured image, so the subject image may be the captured image itself. Alternatively, the subject image may be a partial image extracted from the captured image based on the skeleton points. In the following embodiment, the posture estimation algorithm will be described as detecting a person from an image as preprocessing for posture estimation, and the subject image will be described as a partial image included in a rectangular area enclosed as a result of person detection in the captured image.

[0042] That is, the estimation unit 131 extracts a partial image included in a rectangular area enclosed as a result of person detection from the captured image, and estimates skeleton points corresponding to the detected person using the extracted partial image. The estimation unit 131 estimates 17 key points (skeleton points) present on the person's head, joints, etc. using the above-mentioned posture estimation algorithm (hereinafter referred to as "posture estimation model M2"). That is, posture estimation is a task of estimating 17 key points from the person included in the image.

[0043] The estimation unit 131 can use YOLO (You Only Look Once) as the posture estimation model M2. YOLO (registered trademark) is one of the algorithms used for object detection, and can perform object detection using rectangular regions as preprocessing for posture estimation. For this reason, YOLO handles key points such as "nose," "left-eye," "right-eye," "left-ear," "right-ear," "left-shoulder," and "right-shoulder," and estimates the position coordinates of these key points for a person detected using a rectangular region. Then, lines connecting the key points are drawn and output as a posture estimation result. Note that the posture estimation model M2 is not limited to YOLO and may be any model capable of estimating posture (skeleton points). In the following embodiment, an example in which 17 key points are estimated is shown, but the type and number of key points are not limited to this example.

[0044] <Judgment unit 133> The determination unit 133 determines whether a part of the body of a person included in a partial image extracted in a rectangular area is hidden or whether the entire body of the person is shown, based on the number of key points corresponding to the person. For example, if the number of key points is 17, the determination unit 133 may determine that the entire body of the person included in the partial image is shown, and if the number of key points is less than 17, the determination unit 133 may determine that a part of the body of the person included in the partial image is hidden.

[0045] An example of the operation of the determination unit 133 will be specifically described with reference to Fig. 4. Fig. 4 is a diagram showing an example of a method for determining how a body appears in an image. Fig. 4 shows a scene in which it is determined whether a part of the body of a person included in one captured image acquired at a certain timing during continuous shooting is hidden or whether the entire body is visible.

[0046] In the determination process for determining whether a part of the body is hidden or the whole body is visible, posture estimation is performed using posture estimation model M2 (YOLO), but person detection is first performed as preprocessing for posture estimation. Specifically, estimation unit 131 inputs the original image to posture estimation model M2. As shown in FIG. 4(a), posture estimation model M2 performs person detection as preprocessing, in which a rectangular area (position coordinates of the person) AR indicating the position of the person in the input original image is predicted.

[0047] Next, as shown in FIG. 4(b), the posture estimation model M2 estimates 17 key points from the person included in the partial image (the captured image within the rectangular area AR) extracted from the rectangular area AR. Specifically, the posture estimation model M2 estimates the position of each key point for the person included in the partial image. As a result, the acquisition unit 132 can acquire key points corresponding to the person included in the partial image.

[0048] When the posture estimation model M2 obtains the key point estimation result, the determination unit 133 determines whether the person detected in the partial image has a part of their body hidden or their entire body shown, based on the number of key points included in the partial image. If the number of key points is 17, the determination unit 133 determines that the person included in the partial image has their entire body shown. On the other hand, if the number of key points is less than 17, the determination unit 133 determines that the person included in the partial image has a part of their body hidden.

[0049] <Calculation unit 134> 3, when it is determined that part of the body of the person included in the partial image is hidden, the calculation unit 134 calculates the height distance between the first skeleton point at the highest position in the partial image and the second skeleton point at the lowest position in the partial image. Specifically, the calculation unit 134 estimates the height distance of the entire body in the partial image (height in the partial image) based on a first ratio, which is the ratio between a general average value corresponding to the height distance between the first skeleton point and the second skeleton point and the average height corresponding to the people included in the partial image, and the height distance calculated in the partial image, and calculates the ground contact point of the person included in the partial image based on the estimated height distance.

[0050] If the calculation unit 134 satisfies condition C1 that the number of key points estimated in the captured image of the ground contact point estimation target is less than a first threshold (for example, 17 points) and is equal to or greater than a second threshold (for example, 5 points), the calculation unit 134 estimates the height distance (height) of a person included in a partial image extracted from the captured image of the ground contact point estimation target, based on the height distance calculated in the captured image of the ground contact point estimation target and the first ratio. Note that an example of a situation in which the number of key points is less than the first threshold (for example, 17 points) and equal to or greater than the second threshold (for example, 5 points) is a situation in which the person's head to shoulders is visible in the partial image, and the part below the shoulders is hidden.

[0051] On the other hand, if the condition C2 that the number of key points estimated in the captured image of the ground contact point estimation target is less than a second threshold (for example, 5 points) is satisfied, the calculation unit 134 estimates the height distance (height) of a person included in a partial image extracted from the captured image of the ground contact point estimation target based on a second ratio, which is the ratio between the height distance calculated in the captured image of the ground contact point estimation target and the height distance calculated in a captured image that satisfies condition C1 among captured images taken before the captured image of the ground contact point estimation target, and the first ratio. Note that an example of a situation in which the number of key points is less than the second threshold (for example, 5 points), is a situation in which only the head of a person is visible in the partial image, and the rest below is hidden.

[0052] <Processing Unit 135> The processing unit 135 performs various processes using the ground contact points calculated by the calculation unit 134. For example, the processing unit 135 may generate a screen depicting a movement path (traffic line) of a person based on an overhead view obtained by projectively transforming the captured image in which the ground contact points have been calculated.

[0053] [4. Grounding point calculation method (1)] Fig. 5 is a diagram showing the ground contact point calculation method (1). Fig. 5 shows a scene in which a ground contact point is calculated for one captured image img1 acquired at the latest timing t1 among the timings of continuous shooting. More specifically, Fig. 5 shows a scene in which a ground contact point G1 in the current captured image img1 is calculated for a person U1 included in a partial image Pimg1 extracted from a rectangular area AR1 that indicates the person detection result in the captured image img1.

[0054] 5 illustrates a method for calculating ground contact points when the number of key points corresponding to person U1 satisfies condition C1. Specifically, in the example of FIG. 5, five key points, namely key points P1 to P5, are estimated by pose estimation, and therefore it is determined that part of the body of person U1 is hidden and condition C1 is satisfied. In this example, the calculation unit 134 acquires, from among key points P1 to P5, a first key point (first skeleton point) that is located at the highest position in the coordinate system of the captured image img1 and a second key point (second skeleton point) that is located at the lowest position. Here, it is assumed that the calculation unit 134 acquires key point P2 as the highest first key point and key point P5 as the lowest second key point.

[0055] The calculation unit 134 calculates the height distance D1 between the key point P2 and the key point P5. The height distance D1 here refers to the vertical distance from the key point P5 to the key point P2.

[0056] Here, the calculation unit 134 calculates an average value AV1 corresponding to the height distance D1 (a general average value corresponding to the height distance between the first skeleton point and the second skeleton point). For example, if the key point P2 is the "eye" and the key point P5 is the "elbow," the calculation unit 134 can determine that the "height distance D1" corresponds to the "distance from the eye to the elbow." In this example, the calculation unit 134 can use information on the "average value of the distance from the eye to the elbow" corresponding to the person U1 as the average value AV1. For example, the calculation unit 134 can acquire information on the "average value of the distance from the eye to the elbow" from a human body dimension database. For example, the calculation unit 134 may acquire the "average value of the distance from the eye to the elbow" corresponding to the age and gender of the person U1. Then, the calculation unit 134 uses the acquired "average value of the distance from the eye to the elbow" as the average value AV1.

[0057] Furthermore, to calculate the ground contact point G1 of person U1 in the captured image img1, the calculation unit 134 needs to calculate the height distance X1 of the entire body of person U1 in the captured image img1. Here, the height distance X1 of the entire body indicates the vertical distance from the ground contact point G1 to the key point P2 and corresponds to the height of person U1. However, the height of person U1 is unknown. Therefore, the calculation unit 134 can use the average value AV2 of the height (i.e., height) of the entire body of person U1 (a general average value corresponding to the height distance of the entire predetermined subject). For example, the calculation unit 134 can acquire information on the "average height" corresponding to person U1 from a human body dimension database as the average value AV2. For example, the calculation unit 134 may acquire the "average height" corresponding to the age and gender of person U1. Then, the calculation unit 134 uses the acquired "average height" as the average value AV2.

[0058] In this state, the calculation unit 134 can calculate the height distance X1 of the whole body of the person U1 in the captured image img1 as the height distance Xn of the whole body by solving the equation (1) using the first ratio, which is the ratio between the average value AV1 (cm) and the average value AV2 (cm). In the example of FIG. 5, the height distance D nIn [px], the height distance D1 in the partial image Pimg1 (photographed image img1) is used. Also, the height distance D n The average value AV corresponding to n For [cm], the average value AV1 of the partial image Pimg1 (photographed image im1) is used. Also, the average value AV of the whole body height is used. U For [cm], the average value AV2 corresponding to person U1 is used.

[0059] Then, the calculation unit 134 can calculate the position coordinates of the ground contact point G1 based on the height distance X1 obtained from equation (1) and the highest key point P2.

[0060] [5. Grounding point calculation method (2)] FIG. 6 is a diagram illustrating the ground contact point calculation method (2). FIG. 6 shows a scene in which a ground contact point is calculated for one captured image img3 acquired at the latest timing t3 among the timings of continuous shooting. More specifically, FIG. 6 shows a scene in which a ground contact point G3 in the current captured image img3 is calculated for a person U1 included in a partial image Pimg3 extracted from a rectangular area AR3 that indicates the person detection result in the captured image img3. The captured image img3 is an image captured at timing t3, which is later than timing t1 in FIG. 5.

[0061] 6 illustrates a method for calculating ground contact points when the number of key points corresponding to person U1 does not satisfy condition C1 but satisfies condition C2. Specifically, in the example of FIG. 6, three key points, namely key points P1 to P3, have been estimated by pose estimation, and therefore it is determined that part of the body of person U1 is hidden and condition C2 is satisfied. In this example, the calculation unit 134 acquires, from among key points P1 to P3, a first key point (first skeleton point) that is located at the highest position in the coordinate system of the captured image img3 and a second key point (second skeleton point) that is located at the lowest position. Here, it is assumed that the calculation unit 134 acquires key point P2 as the highest first key point and key point P3 as the lowest second key point.

[0062] The calculation unit 134 calculates the height distance D3 between the key point P2 and the key point P3. The height distance D3 here refers to the vertical distance from the key point P3 to the key point P2.

[0063] 6, the calculation unit 134 needs to calculate the height distance X3 of the entire body of the person U1 in the captured image img3 to find the ground contact point G3, but when the number of key points is small as in condition 2, it may not be possible to accurately calculate the height distance using formula (1). Therefore, it is considered to use information on a captured image in which the person U1 is in a posture closer to the posture of the person U1 at time t3, from among captured images captured at a time before the current time t3 and for which the number of key points satisfies condition C1.

[0064] Specifically, the calculation unit 134 searches for a captured image that was acquired more recently (most recently) among captured images that were acquired before the timing t3 and for which the number of key points satisfies the condition C1. In the example of Fig. 6, the calculation unit 134 obtains the captured image img1 acquired at the timing t1 shown in Fig. 5 as the captured image that was acquired more recently among captured images that were acquired before the timing t3 and for which the number of key points satisfies the condition C1. Since the posture of the person U1 at the timing t3 and the posture of the person U1 at the timing t1 are expected to be relatively similar, it is considered that the height distance X3 can be calculated more accurately than, for example, simply using the height (e.g., height) of the entire body of the person U1 when standing upright.

[0065] In the example of Fig. 6, the height distance X3 of the whole body of the person U1 in the captured image im3 can be calculated as the height distance Xn of the whole body by solving the equation (2) using the second ratio, which is the ratio between the height distance D3 calculated in the current partial image Pimg3 and the height distance D1 calculated in the previous partial image Pimg1. bIn [px], the height distance D1 in the partial image Pimg1 (photographed image img1) is used. Also, the height distance D n In [px], the height distance D3 in the partial image Pimg3 (photographed image img3) is used. Also, the height distance D b The average value AV corresponding to b For [cm], the average value AV1 of the partial image Pimg1 (photographed image im1) is used. Also, the average value AV of the whole body height is used. U For [cm], the average value AV2 corresponding to person U1 is used.

[0066] Then, the calculation unit 134 can calculate the position coordinates of the ground contact point G1 based on the height distance X3 obtained from equation (2) and the highest key point P2.

[0067] [6. Example of operation of the entire system] Fig. 7 is a sequence diagram illustrating an example of operation of the information processing system 1 according to the embodiment. Fig. 7 shows a scene in which the information processing system 1 calculates a ground contact point and displays the calculated ground contact point on the screen.

[0068] 7, the imaging device 10 is fixed at a predetermined position in the space M and is installed in a manner that allows it to capture images of people moving through the space M. The imaging device 10 continuously captures images and sequentially determines whether a captured image has been acquired (step S701). While the imaging device 10 has not been able to acquire a captured image (step S701; No), the imaging device 10 waits until a captured image can be acquired.

[0069] On the other hand, if the imaging device 10 has acquired a captured image (step S701; Yes), the imaging device 10 uploads the acquired captured image to the information processing device 100 (step S702). The information processing device 100 accepts the upload of the captured image (step S703).

[0070] The estimation unit 131 performs inference to detect a person from the captured image acquired this time, and estimates key points corresponding to the detected person using a partial image included in the rectangular area enclosed as the person detection result (step S704).

[0071] Based on the number of key points, the determination unit 133 determines whether the detected person's body is partially hidden or the entire body is visible (step S705).

[0072] First, the processing route (a) when it is determined that the entire body is captured will be described. In the processing route (a), the calculation unit 134 calculates a grounding point on the ground surface of the person detected from the captured image based on the person detection result, i.e., the position coordinates of the person (step S706a). For example, the calculation unit 134 can calculate the grounding point from the aspect ratio of a rectangular area corresponding to the position coordinates of the person.

[0073] On the other hand, in the processing route (b) when it is determined that a part of the body is hidden, the calculation unit 134 calculates the ground contact points on the ground surface of the person detected from the captured image using a method according to the condition (condition 1 or condition 2) that is satisfied by the number of key points (step S706b). This method is as described in FIGS. 5 and 6.

[0074] The processing unit 135 transmits the position coordinates of the person detected from the captured image and information on the ground contact point to the display control device 11 (step S707).

[0075] The display control device 11 depicts a rectangular area indicating the position coordinates of the person and information indicating the grounding point for the currently acquired photographed image (step S708). Furthermore, the display control device 11 controls the display device 12 to display the photographed image including the rectangular area and the grounding point (step S709).

[0076] The display device 12 displays the captured image including the rectangular area and the grounding point in accordance with the control of the display control device 11 (step S710). When this series of processes is completed, the processes from step S701 onwards are repeated again. That is, every time a captured image is acquired by the imaging device 10, the processes from steps S701 to S710 are performed on the acquired captured image.

[0077] FIG. 7 illustrates an example in which the information processing device 100 estimates key points and calculates ground contact points using captured images uploaded by the imaging device 10. However, the key points may be estimated by the imaging device 10. In such an example, the imaging device 10 may include the estimation unit 131 described in FIG. 3, and the information processing device 100 may include components other than the estimation unit 131, namely, the acquisition unit 132, the determination unit 133, the calculation unit 134, and the processing unit 135. In such a configuration, a single device combining the imaging device 10 and the information processing device 100 can be defined as a single information processing device that executes information processing according to the proposed technology of the present invention. In this way, all of the information processing according to the embodiment does not necessarily need to be executed by a cloud computer; part of the information processing according to the embodiment may be executed by an edge computer (the imaging device 10), and the other information processing may be executed by a cloud computer (the information processing device 100).

[0078] 7. Other Embodiments In the above embodiment, an example has been shown in which the cloud-based information processing device 100 calculates the ground contact point of a person detected from an image, that is, an example has been shown in which the information processing device 100 tracks a person by image analysis.

[0079] However, the imaging device 10 may be configured to incorporate an AI processor so that the imaging device 10 itself detects a person from an image and calculates the ground contact point of the detected person. In such an example, the imaging device 10 may function as the information processing device according to the embodiment, and the information processing system 1 may not include the information processing device 100.

[0080] [8. Hardware Configuration] The information processing device 100 according to the embodiment may be realized, for example, by a computer 1000 configured as shown in Fig. 8. Fig. 8 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device 100 according to the embodiment. The computer 1000 has a CPU 1100, a RAM 1200, a ROM 1300, an HDD 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.

[0081] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 starts up, programs that depend on the hardware of the computer 1000, and the like.

[0082] The HDD 1400 stores programs executed by the CPU 1100, data used by these programs, etc. The communication interface 1500 receives data from other devices via a predetermined communication network and sends the data to the CPU 1100, and transmits data generated by the CPU 1100 to other devices via the predetermined communication network.

[0083] The CPU 1100 controls an output device such as a display and an input device such as a keyboard via the input / output interface 1600. The CPU 1100 acquires data from the input device via the input / output interface 1600. The CPU 1100 also outputs generated data to the output device via the input / output interface 1600.

[0084] Media interface 1700 reads a program or data stored in recording medium 1800 and provides it to CPU 1100 via RAM 1200. CPU 1100 loads the program or data from recording medium 1800 onto RAM 1200 via media interface 1700 and executes the loaded program. Recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0085] For example, when the computer 1000 functions as the information processing device 100 according to the embodiment, the CPU 1100 of the computer 1000 executes programs loaded onto the RAM 1200 to realize the functions of the control unit 130. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800, but as another example, the CPU 1100 may obtain these programs from another device via a predetermined communication network.

[0086] [9. Other] Furthermore, among the processes described in each of the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0087] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0088] Furthermore, the above-described embodiments can be combined as appropriate within the scope of not causing any contradiction in the processing content.

[0089] Although some of the embodiments of the present application have been described in detail above with reference to the drawings, these are merely examples, and the present invention can be implemented in other forms that include the aspects described in the "present invention" section and that have been modified and improved in various ways based on the knowledge of those skilled in the art. [Explanation of symbols]

[0090] 1. Information Processing Systems 10. Imaging device 11 Display control device 12 Display device 100 Information processing device 130 control section 131 Estimation Department 132 Acquisition Department 133 Judgment section 134 Calculation Unit 135 Processing section

Claims

1. an acquisition unit that acquires a subject image, which is an image of a portion of a rectangular area detected from an image using an object detection model, the rectangular area indicating the position of a predetermined subject in the image; an estimation unit that estimates a skeleton point of the predetermined subject included in the subject image; a determination unit that determines, based on the estimated skeleton points, whether the image shows a state in which the ground surface at the feet of the predetermined subject is hidden or whether the image shows a state in which the entire predetermined subject is included; a calculation unit that calculates a ground contact point on the ground contact surface of the predetermined subject based on a height distance between a first skeleton point that is at the highest position in the subject image and a second skeleton point that is at the lowest position in the subject image, among the estimated skeleton points, when it is determined that the ground contact surface is captured in a hidden state; Equipped with The calculation unit estimates a height distance of the entire predetermined subject included in the subject image based on a first ratio, which is a ratio between a general average value corresponding to a height distance between the first skeleton point and the second skeleton point and a general average value corresponding to a height distance of the entire predetermined subject, and the height distance calculated in the subject image, and calculates a ground contact point of the predetermined subject included in the subject image based on the estimated height distance of the entire predetermined subject. Information processing device.

2. The calculation unit calculates the ground contact point of the predetermined subject for a subject image determined to be a subject image in which the ground contact surface is hidden among the subject images sequentially acquired by continuous real-time photography. The information processing device according to claim 1 .

3. When the number of skeleton points is less than a first threshold and is equal to or greater than a second threshold, the calculation unit estimates a height distance of the entire predetermined subject included in the subject image of the target based on the height distance calculated in the subject image of the target determined to have the number of skeleton points less than the first threshold and equal to or greater than the second threshold and the first ratio. The information processing device according to claim 1 .

4. When the number of skeleton points is less than the second threshold, the calculation unit estimates the height distance of the entire predetermined subject included in the subject image of the target based on a second ratio, which is a ratio between the height distance calculated for the subject image of the target in which it is determined that the number of skeleton points is less than the second threshold, and the height distance calculated for a previous subject image, among subject images taken before the subject image of the target, in which it is determined that the number of skeleton points is less than the first threshold and equal to or greater than the second threshold, and based on the first ratio. The information processing device according to claim 3 .

5. An information processing method executed by an information processing device, an acquisition step of acquiring a subject image, which is an image of a portion of a rectangular area detected from an image using an object detection model, the rectangular area indicating the position of a predetermined subject in the image; an estimation step of estimating a skeleton point of the predetermined subject included in the subject image; a determining step of determining, based on the estimated skeleton points, whether the photograph shows a state in which the ground surface at the feet of the predetermined subject is hidden or whether the photograph shows a state in which the entire predetermined subject is included; a calculation step of calculating a ground contact point on the ground contact surface of the predetermined subject based on a height distance between a first skeleton point that is at the highest position in the subject image and a second skeleton point that is at the lowest position in the subject image, among the estimated skeleton points, when it is determined that the ground contact surface is photographed in a hidden state; Including, The calculation step estimates a height distance of the entire predetermined subject included in the subject image based on a first ratio, which is a ratio between a general average value corresponding to a height distance between the first skeleton point and the second skeleton point and a general average value corresponding to a height distance of the entire predetermined subject, and the height distance calculated in the subject image, and calculates a ground contact point of the predetermined subject included in the subject image based on the estimated height distance of the entire predetermined subject. Information processing methods.

6. An information processing program executed by an information processing device, an acquisition step of acquiring a subject image, which is an image of a portion of a rectangular area detected from an image using an object detection model, the rectangular area indicating the position of a predetermined subject in the image; an estimation step of estimating a skeleton point of the predetermined subject included in the subject image; a determination step of determining, based on the estimated skeleton points, whether the photograph shows a state in which the ground surface at the feet of the predetermined subject is hidden or whether the photograph shows a state in which the entire predetermined subject is included; a calculation step of calculating a ground contact point on the ground contact surface of the predetermined subject based on a height distance between a first skeleton point at the highest position and a second skeleton point at the lowest position in the subject image among the estimated skeleton points when it is determined that the ground contact surface is hidden in the image; causing the information processing device to execute the above; The calculation step estimates a height distance of the entire predetermined subject included in the subject image based on a first ratio, which is a ratio between a general average value corresponding to a height distance between the first skeleton point and the second skeleton point and a general average value corresponding to a height distance of the entire predetermined subject, and the height distance calculated in the subject image, and calculates a ground contact point of the predetermined subject included in the subject image based on the estimated height distance of the entire predetermined subject. Information processing program.

Citation Information

Patent Citations

  • Ground point estimation device, ground point estimation method, flow line display system, and server

    JP2009236569A

  • Identification program, identification method, and information processor

    JP2023031227A

  • Position information recording device, position information recording method, program, and storage media

    JP2024022921A

  • System and Method for Providing Multi-Camera 3D Body Part Labeling and Performance Metrics

    US20220129669A1