Information processing apparatus, program, system and information processing method

A dual-camera system with one camera capturing from directly above and another from a different angle addresses the challenge of recognizing objects from direct overhead views by generating annotated images, thereby improving recognition accuracy and reducing manual annotation requirements.

JP2025172952APending Publication Date: 2025-11-26SOFTBANK CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025149507
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-11-26

AI Technical Summary

Technical Problem

Existing image processing systems struggle to accurately recognize objects, particularly humans, when captured from directly above due to hidden body parts and significant shape differences, making it difficult to automatically assign metadata and requiring manual annotation of large image datasets.

Method used

Employing a dual-camera system with a main camera capturing images from directly above and an auxiliary camera from a different angle, such as the side or diagonally, to generate annotated images and improve recognition accuracy through projective transformation and metadata assignment.

Benefits of technology

Enables automatic generation of annotated images, enhancing the ability to recognize objects like humans from challenging angles, reducing the need for manual annotation and improving recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025172952000001_ABST
    Figure 2025172952000001_ABST
Patent Text Reader

Abstract

To appropriately annotate a person area in an image that captures a person from above or the like.SOLUTION: An information processing device of a system 10 includes: an image acquisition unit which acquires a first captured image 500 captured by an imaging unit 100 which captures an object and a ground surface from a position other than directly above the object, and a second captured image 600 captured by an imaging unit 200 which captures the object from directly above or the like; a rectangular range specification unit which specifies a first rectangular range showing a range of the object in the first captured image; a ground surface range specification unit which specifies a ground surface range showing the range of the object within the ground surface in the first captured image on the basis of the first rectangular range; a projective transformation unit which uses a plurality of predetermined reference points on the ground surface in the first captured image and a plurality of reference points on the ground surface in the second captured image to perform projective transformation of the ground surface range in the first captured image into the second captured image; and a rectangular range determination unit which determines a second rectangular range showing the range of the object in the second captured image on the basis of the ground surface range projective-transformed into the second captured image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, a program, a system, and an information processing method. [Background technology]

[0002] Patent document 1 describes an image display device that cuts out a portion of a forward image and performs projective transformation on the cut-out image to generate a downward image corresponding to a rectangular road surface area located below a vehicle VC at a second timing. [Prior art document] [Patent documents] [Patent Document 1] JP 2023-141792 A Summary of the Invention [Means for solving the problem]

[0003] According to one embodiment of the present invention, there is provided an information processing device. The information processing device may include an image acquisition unit that acquires a first captured image of an object and a ground surface on which the object is grounded by a first imaging unit that captures an image of the object and the ground surface from a position other than directly above the object, and a second captured image including the object and the ground surface that is captured by a second imaging unit that captures an image of the object from directly above the object. The information processing device may include a rectangular area identification unit that analyzes the first captured image and identifies a first rectangular area indicating the range of the object in the first captured image. The information processing device may include a ground surface area identification unit that identifies a ground surface area indicating the range of the object within the ground surface in the first captured image based on the first rectangular area. The information processing device may include a projection transformation unit that projectively transforms the ground surface area in the first captured image into the second captured image using a plurality of predetermined reference points on the ground surface in the first captured image and the plurality of reference points on the ground surface in the second captured image. The information processing device may include a rectangular range determination unit that determines a second rectangular range indicating the range of the object in the second captured image based on the contact surface range projected into the second captured image by the projection transformation unit.

[0004] In the information processing device, the image acquisition unit may acquire the first captured image of the object and the ground surface captured by the first imaging unit, which captures images of the object and the ground surface from a position directly above the object and a position diagonally above that is substantially equal to directly above the object, and the second captured image of the object and the ground surface captured by a second imaging unit, which captures images of the object from a position directly above the object or a position diagonally above that is substantially equal to directly above the object. In any of the information processing devices, the ground contact area specification unit may specify a first point and a second point that are two vertices of a lower side of the first rectangular area, a third point within the rectangular area that is the horizontal center of the rectangular area specified by the rectangular area specification unit and is located at a length from the top of the rectangular area that is a predetermined percentage of the vertical length of the rectangular area, a fourth point that is a point-symmetrical position of the first point with the third point as the center, and a fifth point that is a point-symmetrical position of the second point with the third point as the center, thereby specifying the ground contact area of ​​a rectangle formed by the first point, the second point, the fourth point, and the fifth point. The projective transformation unit may project the first point, the second point, the third point, the fourth point, and the fifth point in the first captured image into the second captured image using the plurality of reference points of the ground contact area in the first captured image and the plurality of reference points in the second captured image.

[0005] In any of the information processing devices, the rectangular area specification unit may estimate a posture of the subject in the first captured image, and specify a midpoint between a ground contact point of the subject's right foot and a ground contact point of the subject's left foot as a center of gravity of the subject based on the estimated posture of the subject. The projection transformation unit may further projectively transform the center of gravity of the subject into the second captured image. The rectangular area determination unit may determine a size of the second rectangular area from the ground contact area projectively transformed into the second captured image by the projection transformation unit, and may set the center of gravity of the subject projected into the second captured image by the projection transformation unit as a center point of the second rectangular area.

[0006] In any of the information processing devices, the rectangular area determination unit may determine the size of the second rectangular area by setting the larger of the difference between the maximum and minimum values ​​in a horizontal direction and the difference between the maximum and minimum values ​​in a vertical direction of the coordinates of the first point, the second point, the fourth point, and the fifth point as the length of one side. The rectangular area identification unit may estimate the subject's skeleton including a point indicating the subject's right ankle, a point indicating the right knee, a point indicating the left ankle, and a point indicating the left knee, and identify the subject's center of gravity by adjusting the vertical position of the midpoint between the point indicating the right ankle and the point indicating the left ankle based on a shin length calculated from the point indicating the right ankle and the point indicating the right knee or the point indicating the left ankle and the point indicating the left knee.

[0007] In any of the information processing devices, the image acquisition unit may select, as the first captured image, an image in which the distance between the right foot and left foot of the subject is longer from among multiple images successively captured by the first imaging unit.

[0008] Any of the information processing devices may include an annotation unit that assigns metadata representing the object to the rectangular area determined by the rectangular area determination unit. In any of the information processing devices, the projective transformation unit may identify the plurality of reference points in the first captured image and the plurality of reference points in the second captured image by detecting a plurality of markers arranged on the ground surface. The information processing device may perform machine learning using training data including captured images to which metadata has been assigned by the annotation unit. The learning execution unit may use a plurality of the training data to generate a learning model that, when a captured image includes a person captured from directly above, is capable of outputting that the part of the captured image is a person.

[0009] According to one embodiment of the present invention, there is provided an information processing device. The information processing device may include an image acquisition unit that acquires a first captured image of a person captured by a first imaging unit with the person's torso, limbs, and head separated, and a second captured image of the person captured by a second imaging unit that captures the person from a direction different from the first imaging unit with the person's torso, limbs, and head not separated. The information processing device may include a rectangular area identification unit that analyzes the first captured image and identifies a first rectangular area indicating the area of ​​the person in the first captured image. The information processing device may include a ground plane area identification unit that identifies a ground plane area indicating the area of ​​the person within the ground plane of the person in the first captured image based on the first rectangular area. The information processing device may include a projection transformation unit that projectively transforms the ground plane area in the first captured image into the second captured image using a plurality of predetermined reference points of the ground plane in the first captured image and the plurality of reference points of the ground plane in the second captured image. The information processing device may include a rectangular range determination unit that determines a second rectangular range indicating the range of the person in the second captured image based on the contact surface range projected into the second captured image by the projection transformation unit.

[0010] According to one embodiment of the present invention, there is provided a program for causing a computer to function as the information processing device when executed by the computer.

[0011] According to one embodiment of the present invention, there is provided a system including the information processing device, the first imaging unit, and the second imaging unit.

[0012] According to one embodiment of the present invention, there is provided an information processing method executed by a computer. The information processing method may include an image acquisition step in which a first imaging unit, which images an object and a ground surface on which the object is in contact with the ground from a position other than directly above the object, acquires a first captured image of the object and the ground surface, and a second imaging unit, which images the object from directly above the object, acquires a second captured image including the object and the ground surface. The information processing method may include a rectangular area identification step in which the first captured image is analyzed to identify a first rectangular area indicating the range of the object in the first captured image. The information processing method may include a ground surface area identification step in which, based on the first rectangular area, a ground surface area indicating the range of the object within the ground surface in the first captured image is identified. The information processing method may include a projection transformation step of projecting the ground surface range in the first captured image into the second captured image using a plurality of predetermined reference points of the ground surface in the first captured image and the plurality of reference points of the ground surface in the second captured image, and a rectangular range determination step of determining a second rectangular range indicating the range of the object in the second captured image based on the ground surface range projected into the second captured image in the projection transformation step.

[0013] According to one embodiment of the present invention, there is provided an information processing method executed by a computer. The information processing method may include an image acquisition step of acquiring a first captured image of a person captured by a first imaging unit with the person's torso, limbs, and head separated, and a second captured image of the person captured by a second imaging unit, which captures the person from a direction different from that of the first imaging unit, with the person's torso, limbs, and head not separated. The information processing method may include a rectangular area identification step of analyzing the first captured image to identify a first rectangular area indicating the area of ​​the person in the first captured image. The information processing method may include a ground surface area identification step of identifying a ground surface area indicating the area of ​​the person within the ground surface of the person in the first captured image based on the first rectangular area. The information processing method may include a projective transformation step of projectively transforming the ground surface area in the first captured image into the second captured image using a plurality of predetermined reference points of the ground surface in the first captured image and the plurality of reference points of the ground surface in the second captured image. The information processing method may include a rectangular range determination step of determining a second rectangular range indicating the range of the person in the second captured image based on the ground plane range projected into the second captured image in the projective transformation step.

[0014] The above summary of the invention does not list all of the features of the present invention, and subcombinations of these features may also be inventions. [Brief explanation of the drawings]

[0015] [Figure 1] An example of a system 10 is shown schematically. [Figure 2] 10 is an explanatory diagram for explaining how the information processing device 300 identifies the contact surface range. FIG. [Figure 3] 10 is an explanatory diagram for explaining projective transformation of the contact surface range by the information processing device 300. FIG. [Figure 4] FIG. 10 is an explanatory diagram illustrating how the information processing device 300 determines the size of a rectangular area. [Figure 5] 2 shows an example of a functional configuration of an information processing device 300. [Figure 6] 10 is an explanatory diagram for explaining how the information processing device 300 identifies the center of gravity of an object. FIG. [Figure 7] 10 shows an example of a processing flow of the information processing device 300. [Figure 8] An example of the hardware configuration of a computer 1200 that functions as the information processing device 300 is shown in schematic form. DETAILED DESCRIPTION OF THE INVENTION

[0016] The present invention will be described below through embodiments of the invention, but the following embodiments do not limit the scope of the invention according to the claims. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.

[0017] When a person is imaged from directly above, from a diagonal direction substantially equal to directly above, from directly below, or from a diagonal direction substantially equal to directly below (sometimes referred to as "directly above, etc."), it is difficult to automatically recognize that the person is a person. While many conventional annotated images are composed of images captured from the side of a person, capturing the head, torso, legs, etc., when a person is imaged from directly above, etc., at least a portion of the torso is hidden by the head, or at least a portion of the legs are hidden by the torso, etc., and the head, torso, legs, etc. are not captured. As a result, the shape of the person captured in the captured image differs significantly from the shape of the person captured in the annotated image. Machine learning can be performed using a large number of annotated images in which metadata indicating that a region of a person in an image captured from directly above, etc., is added to the region of the person. This can improve the accuracy of detecting people from directly above, etc. However, because it is difficult to automatically assign metadata to images of a person captured from directly above, annotated images must be prepared manually, making it practically difficult to prepare a large number of annotated images. In the system 10 according to the present embodiment, for example, in addition to a main camera that captures a person from directly above, an auxiliary camera that captures a person from a position other than directly above is used. For example, the system 10 uses an auxiliary camera that captures a person from the side. For example, the system 10 uses an auxiliary camera that captures a person from an obliquely above. As is well known, if a person is captured from the side or from an obliquely above, it is possible to recognize that the person is included in the captured image. Therefore, by using an auxiliary camera in addition to the main camera, a person can be recognized from the image captured by the auxiliary camera, and metadata representing the person can be assigned to an area of ​​the image captured by the main camera that includes the person. As described above, the main camera is not limited to being installed so as to capture a person from directly above, but may also be installed so as to capture a person from directly below, from an obliquely above that is substantially equal to directly above, or from an obliquely below that is substantially equal to directly below. The diagonally upward position, which is substantially the same as directly above, may be within a range in which it is difficult to automatically recognize a person when an image of the person is captured from that position.The diagonal downward angle, which is substantially the same as directly below, may be within a range in which it is difficult to automatically recognize a person when a person is captured from that angle. This contributes to the automatic generation of annotated images. Note that the system 10 may generate annotated images not limited to people, but for any object that is difficult to automatically recognize from directly above, but is automatically recognizable from other angles. In this embodiment, the main example is a case in which the object is a person and the main camera is installed to capture the person from directly above.

[0018] FIG. 1 schematically illustrates an example of a system 10. The system 10 includes an imaging unit 100, an imaging unit 200, and an information processing device 300. The imaging unit 100 may be an example of a first imaging unit. The imaging unit 100 may be used as an auxiliary camera. The imaging unit 200 may be an example of a second imaging unit. The imaging unit 200 may be used as a main camera.

[0019] The information processing device 300 communicates with the image capturing unit 100 and the image capturing unit 200. The information processing device 300 communicates with the image capturing unit 100 and the image capturing unit 200 via a network 20, for example.

[0020] The network 20 may include the Internet. The network 20 may include a cloud. The network 20 may include a LAN (Local Area Network). The network 20 may include a mobile communication network. The mobile communication network may be compliant with a 5G (5th Generation) communication system. The mobile communication network may be compliant with an LTE (Long Term Evolution) communication system. The mobile communication network may be compliant with a 3G (3rd Generation) communication system. The mobile communication network may be compliant with a 6G (6th Generation) communication system or later.

[0021] The information processing device 300 may be connected to the network 20 by wire. The information processing device 300 may be connected to the network 20 wirelessly. The information processing device 300 may be connected to the network 20 via a wireless base station. The information processing device 300 may be connected to the network 20 via a Wi-Fi (registered trademark) access point.

[0022] The imaging unit 100 may be connected to the network 20 by wire. The imaging unit 100 may be connected to the network 20 wirelessly. The imaging unit 100 may be connected to the network 20 via a wireless base station. The imaging unit 100 may be connected to the network 20 via a Wi-Fi access point.

[0023] The imaging unit 200 may be connected to the network 20 by wire. The imaging unit 200 may be connected to the network 20 wirelessly. The imaging unit 200 may be connected to the network 20 via a wireless base station. The imaging unit 200 may be connected to the network 20 via a Wi-Fi access point.

[0024] The image capturing unit 100 and the information processing device 300 may be directly connected via a communication cable. The image capturing unit 200 and the information processing device 300 may be directly connected via a communication cable.

[0025] The imaging unit 100 may be any device that can capture an image of the target person 50 and provide the captured image 500 to the information processing device 300. For example, the imaging unit 100 may be a camera device such as a surveillance camera or a network camera. The imaging unit 100 may also be a smartphone or a tablet terminal. For example, the imaging unit 100 may provide the information processing device 300 with the captured image 500, in which the target person 50's torso, limbs, and head are separated. For example, the imaging unit 100 is disposed in a position where it can capture an image of the target person 50 from a position other than directly above, and captures an image of the target person 50 and the ground surface 60 on which the target person 50 is standing. For example, the imaging unit 100 is disposed in a position where it can capture an image of the target person 50 from diagonally above, and captures an image of the target person 50 and the ground surface 60 from diagonally above. The captured image 500 may be an example of a first captured image.

[0026] The imaging unit 200 may be any device that can capture an image of the target person 50 and provide the captured image 600 to the information processing device 300. For example, the imaging unit 200 may be a camera device such as a surveillance camera or a network camera. The imaging unit 200 may also be a smartphone or a tablet terminal. The imaging unit 200 captures an image of the target person 50 from a direction different from that of the imaging unit 100. The imaging unit 200 may provide the information processing device 300 with a captured image 600 in which the target person 50's torso, limbs, and head are not separated. For example, the imaging unit 200 is disposed in a position where it can capture an image of the target person 50 from directly above, and captures an image from above downward to capture the target person 50 and the ground surface 60 from directly above. The captured image 600 may be an example of a second captured image.

[0027] The information processing device 300 acquires a captured image 500, which is an image of the target person 50 and the ground surface 60 captured by the imaging unit 100, from the imaging unit 100. The information processing device 300 acquires a captured image 600, which is an image of the target person 50 and the ground surface 60 captured by the imaging unit 200, from the imaging unit 200. The information processing device 300, for example, analyzes the captured image 500 to identify a rectangular area indicating the range of the target person 50 in the captured image 500, identifies a ground surface area indicating the range of the target person 50 within the ground surface 60 in the captured image 500 from the rectangular area, and performs projective transformation of the ground surface area in the captured image 500 onto the captured image 600. The information processing device 300 determines a rectangular area indicating the range of the target person 50 in the captured image 600 based on the ground surface area projectively transformed onto the captured image 600, and assigns metadata representing the target person 50 to the determined rectangular area. Even if the rectangular area of ​​the target person 50 in the captured image 600 cannot be directly identified from the captured image 600, it can be indirectly identified by using the captured image 500, and an annotated image in which metadata of the target person 50 is added to an appropriate area of ​​the captured image 600 in which the target person 50 is captured from directly above can be generated. The information processing device 300 may be any device that can execute the above-described processing. The information processing device 300 may be, for example, a PC (Personal Computer), a server device, a smartphone, etc.

[0028] 2 is an explanatory diagram for explaining the identification of the ground surface range 518 by the information processing device 300. The captured image 500 is captured by the imaging unit 100, and includes a target person 50 and a ground surface 60 within the captured range.

[0029] A plurality of reference points are arranged on the ground surface 60. In the example shown in FIG. 2, reference points 71, 72, 73, and 74 are arranged on the ground surface 60. The plurality of reference points may be arranged in any manner as long as they are arranged so that their positions can be identified in the captured image. For example, the plurality of reference points may be indicated by tape, such as a so-called marking tape. The shapes of the plurality of reference points may be the same or different from each other.

[0030] 2, the target person 50 included in the captured image 500 is seen from behind with his left foot forward while walking. The information processing device 300 analyzes the captured image 500 and identifies a rectangular bounding box 510 that indicates the range of the target person 50 in the captured image 500. The bounding box 510 may be an example of a first rectangular range.

[0031] In this embodiment, a rectangle may be a shape with four right-angled vertices, and may include a rectangle and a square. In this embodiment, a bounding box may be a partial region surrounding an object in an image or video. The information processing device 300 may identify a bounding box 510 surrounding the target person 50 in the captured image 500 using well-known technology.

[0032] The information processing device 300 may use the bounding box 510 to identify the contact surface range 518. For example, the information processing device 300 identifies points 511 and 512, which are the two vertices of the lower side of the bounding box 510. The point 511 may be an example of a first point. The point 512 may be an example of a second point. The information processing device 300 may identify point 513 within the bounding box 510, which is the center of the horizontal length of the bounding box 510 and is a position at a distance from the top of the bounding box 510 that is a predetermined percentage of the vertical length of the bounding box 510. The point 513 may be an example of a third point. The predetermined percentage may be set in advance and may be changeable. The example shown in FIG. 2 illustrates a case where 0.95 is set as the predetermined percentage.

[0033] The information processing device 300 may identify point 514, which is a position symmetrical to point 511 with point 513 as the center. The information processing device 300 may identify point 515, which is a position symmetrical to point 512 with point 513 as the center. Point 514 may be an example of a fourth point. Point 515 may be an example of a fifth point. The information processing device 300 may identify points 514 and 515 by vector calculation. The information processing device 300 may identify a quadrangular contact surface range 518 formed by points 511, 512, 514, and 515.

[0034] Fig. 3 is an explanatory diagram for describing the projective transformation of the ground plane range 518 by the information processing device 300. In the example shown in Fig. 3, for the sake of explanation, four reference points 71, 72, 73, and 74 are shown with different shapes.

[0035] Captured image 600 is captured by imaging section 200, and includes within its imaging range a target person 50 and a ground surface 60. Captured image 600 includes reference points 71 to 74, which are also included within the imaging range of captured image 500, at positions captured from directly above.

[0036] The information processing device 300 uses reference points 71, 72, 73, and 74 in the captured image 500 and reference points 71, 72, 73, and 74 in the captured image 600 to project five points, 511, 512, 513, 514, and 515, in the captured image 500 onto the captured image 600.

[0037] 4 is an explanatory diagram for explaining determination of the size of a rectangular area by the information processing device 300. The information processing device 300 may determine the size of a rectangular area surrounding the target person 50 in the captured image 600 by using points 511, 512, 514, and 515 that are projectively transformed onto the captured image 600.

[0038] For example, the information processing device 300 determines the size of the rectangular information by taking the larger of the difference between the maximum and minimum values ​​in the horizontal direction of the coordinates of points 511, 512, 514, and 515 and the difference between the maximum and minimum values ​​in the vertical direction as the length of one side. In the example shown in FIG. 4, the information processing device 300 determines the size of the rectangular information by taking the x-coordinate (x min ) to the x-coordinate (x max ) is set as the difference 612. The information processing device 300 may generate rectangle information 622 with the difference 612 as one side. The information processing device 300 may use the y coordinate (y min ) to the y coordinate of point 511, which is the maximum value of the y coordinate (y max ) may be used as the difference 614. The information processing device 300 may generate rectangle information 624 with the difference 614 as one side. The information processing device 300 determines the size of the rectangle information by taking the larger value of the difference 612 or the difference 614 as the length of one side. In the example shown in FIG. 4, the information processing device 300 sets the size of the rectangle information 622 to the size of a rectangular range that surrounds the target person 50 in the captured image 600.

[0039] The information processing device 300 may arrange the rectangular information 622 within the captured image 600. For example, the information processing device 300 estimates the posture of the target person 50 in the captured image 500, and determines the center of gravity of the target person 50 on the ground surface 60 based on the estimated posture of the target person 50. For example, the information processing device 300 determines the midpoint between the ground contact point of the target person 50's right foot and the ground contact point of the target person 50's left foot as the center of gravity of the target person 50. A method for determining the center of gravity will be described later. The information processing device 300 may determine a rectangular range that indicates the range of the target person 50 in the captured image 600 by projectively transforming the determined center of gravity onto the captured image 600 and arranging the rectangular information 622 so that the center of gravity coincides with the center of gravity. In addition, the information processing device 300 may determine a rectangular range indicating the range of the target person 50 in the captured image 600 by positioning the rectangular information 622 so that the center point of the rectangular information 622 coincides with the point 513 projected onto the captured image 600.

[0040] 1 to 4, the image capturing unit 200 has been mainly described as being installed so as to capture an image of the target person 50 from directly above, but as described above, the image capturing unit 200 may be installed so as to capture an image of the person from an oblique angle substantially equivalent to directly above. In this case, an annotated image can be generated using a method similar to that described in FIGS. 1 to 4.

[0041] Furthermore, the imaging unit 200 may be installed so as to capture an image of a person from directly below. In this case, the imaging unit 100 and the imaging unit 200 may be placed under a transparent floor. As a specific example, the imaging unit 100 and the imaging unit 200 are placed under a glass or acrylic floor. A plurality of markers are placed on the transparent floor.

[0042] The imaging unit 200 is placed in a position where it can image the person from directly below, and images the target person 50 and the ground surface 60 from directly below in an upward direction. The imaging unit 100 is placed in a position where it can image the target person 50 from a position other than directly below, and images the target person 50 and the ground surface 60. The imaging unit 100 is placed in a position where it can image the target person 50 from diagonally below, for example, and images the target person 50 and the ground surface 60 from diagonally below.

[0043] The information processing device 300 acquires a captured image 500, which is an image of the target person 50 and the ground surface 60 captured by the imaging unit 100, from the imaging unit 100. The information processing device 300 acquires a captured image 600, which is an image of the target person 50 and the ground surface 60 captured by the imaging unit 200, from the imaging unit 200. The information processing device 300, for example, analyzes the captured image 500 to identify a rectangular area indicating the range of the target person 50 in the captured image 500, identifies a ground surface area indicating the range of the target person 50 within the ground surface 60 in the captured image 500 from the rectangular area, and performs projective transformation of the ground surface area in the captured image 500 onto the captured image 600. The information processing device 300 determines a rectangular area indicating the range of the target person 50 in the captured image 600 based on the ground surface area projectively transformed onto the captured image 600, and assigns metadata representing the target person 50 to the determined rectangular area. The information processing device 300 may identify the ground surface range 518, perform projective transformation of the ground surface range 518, and determine the size of the rectangular range using methods similar to those described in Figures 2 to 4.

[0044] As described above, the image capturing unit 200 may be installed so as to capture an image of a person from a diagonally downward direction substantially equivalent to directly downward. In this case, the same method as described above can be used.

[0045] 5 schematically illustrates an example of the functional configuration of the information processing device 300. The information processing device 300 includes a storage unit 302, an image acquisition unit 304, a rectangular area identification unit 306, a ground surface area identification unit 308, a projective transformation unit 310, a rectangular area determination unit 312, an annotation unit 314, and a learning execution unit 316. Note that it is not essential for the information processing device 300 to include all of these components.

[0046] The image acquisition unit 304 may acquire a captured image 500 obtained by the imaging unit 100 capturing an image of an object and a ground surface on which the object is grounded. The image acquisition unit 304 stores the acquired captured image 500 in the storage unit 302. The image acquisition unit 304 may acquire a captured image 600 obtained by the imaging unit 200 capturing an image of an object and a ground surface on which the object is grounded. The image acquisition unit 304 stores the acquired captured image 600 in the storage unit 302. The object may be a person. The object is not limited to a person, and may be any object that is difficult to automatically recognize from directly above but can be automatically recognized from other than directly above.

[0047] When the target is a person, the image acquisition unit 304 may acquire a captured image 500 in which the imaging unit 100 captures the person in a manner in which the person's torso, limbs, and head are separated. The image acquisition unit 304 stores the acquired captured image 500 in the storage unit 302. The image acquisition unit 304 may acquire a captured image 600 in which the imaging unit 200 captures the person in a manner in which the person's torso, limbs, and head are not separated. The image acquisition unit 304 stores the acquired captured image 600 in the storage unit 302.

[0048] The rectangular area identification unit 306 analyzes the captured image 500 and identifies a rectangular area (sometimes referred to as a first rectangular area) that indicates the range of the object in the captured image 500. The rectangular area identification unit 306 may identify a bounding box 510 that indicates the range of the object in the captured image 500.

[0049] The rectangular area identification unit 306 may further identify the center of gravity of the target on the ground contact surface in the captured image 500. For example, if the target is a person, the rectangular area identification unit 306 estimates the posture of the target in the captured image 500, and based on the estimated posture of the target, identifies the midpoint between the ground contact points of the target's right foot and the target's left foot as the target's center of gravity.

[0050] The ground plane range identifying unit 308 identifies a ground plane range 518 indicating the range of an object within the ground plane in the captured image 500 based on the bounding box 510. For example, the ground plane range identifying unit 308 identifies the ground plane range 518 indicating the range of an object within the ground plane in the captured image 500 by using points 511 and 512, which are the two lower vertices of the bounding box 510. For example, the ground plane range identifying unit 308 identifies a point 513 within the bounding box 510, which is the horizontal center of the bounding box 510 and is a position a predetermined percentage of the vertical length of the bounding box 510 from above the bounding box 510; a point 514 which is a point-symmetrical position of point 511 with point 513 as the center; and a point 515 which is a point-symmetrical position of point 512 with point 513 as the center, and identifies a quadrangular ground plane range 518 formed by points 511, 512, 514, and 515.

[0051] The projection transformation unit 310 uses a plurality of predetermined reference points on the ground surface in the captured image 500 and a plurality of reference points on the ground surface in the captured image 600 to projectively transform a ground surface range 518 in the captured image 500 onto the captured image 600. For example, the projection transformation unit 310 uses a plurality of reference points on the ground surface in the captured image 500 and a plurality of reference points on the ground surface in the captured image 600 to projectively transform points 511, 512, 513, 514, and 515 in the captured image 500 onto the captured image 600.

[0052] The projection transformation unit 310 may further projectively transform the center of gravity of the object identified by the rectangular area identification unit 306 into the captured image 600. The projection transformation unit 310 may identify multiple reference points in the captured image 500 and multiple reference points in the captured image 600 by detecting multiple markers placed on the ground surface. As described above, the multiple markers may be tape, such as a marker tape. The multiple markers may be of any shape as long as they are arranged in the captured image so that their positions can be identified. It is desirable that the multiple markers be on the same plane as much as possible, which is preferably the ground surface. For example, a mark such as a sticker may be placed on the ground surface to serve as a marker. The shapes of the multiple markers may be the same or different from each other.

[0053] The rectangular area determination unit 312 determines a rectangular area (sometimes referred to as a second rectangular area) indicating the area of ​​the object in the captured image 600, based on the ground plane area 518 that has been projectively transformed into the captured image 600 by the projection transformation unit 310. The rectangular area determination unit 312 may determine the size of the second rectangular area from the ground plane area 518 that has been projectively transformed into the captured image 600 by the projection transformation unit 310. For example, the rectangular area determination unit 312 determines the size of the second rectangular area by taking the larger of the difference between the maximum and minimum values ​​in the horizontal direction and the difference between the maximum and minimum values ​​in the vertical direction of the coordinates of points 511, 512, 514, and 515 as the length of one side. The rectangular area determination unit 312 may determine the center of gravity of the object that has been projectively transformed into the captured image 600 by the projection transformation unit 310 as the center point of the second rectangular area. The rectangular area determination section 312 may determine the point 513 that has been projectively transformed into the captured image 600 by the projective transformation section 310 as the center point of the second rectangular area.

[0054] The annotation unit 314 assigns metadata representing the object to the second rectangular area of ​​the captured image 600 determined by the rectangular area determination unit 312.

[0055] The learning execution unit 316 executes machine learning using learning data including captured images 600 to which metadata has been added by the annotation unit 314. The learning execution unit 316 uses a plurality of pieces of learning data to generate a learning model that, using a captured image as input, can output that a part of the captured image captured from directly above contains a person. The learning execution unit 316 stores the generated learning model in the storage unit 302. Note that the information processing device 300 does not necessarily have to include the learning execution unit 316.

[0056] 6 is an explanatory diagram for explaining how the information processing device 300 identifies the center of gravity of the target. The information processing device 300 may estimate the posture of the target, and identify the center of gravity of the target on the ground surface based on the estimated posture of the target. Here, a captured image 520, which is an example of the captured image 500, will be used as an example for explanation. The captured image 520 includes a target person 50.

[0057] The rectangular area identification unit 306 analyzes the captured image 520 to identify a bounding box 530 indicating the area of ​​the target person 50 in the captured image 520. The rectangular area identification unit 306 performs posture estimation of the target person 50. By performing posture estimation, the rectangular area identification unit 306 may estimate the skeleton of the target person 50, as shown in FIG. 6 . The rectangular area identification unit 306 may estimate the skeleton of the target person 50, including points for each joint of the target person 50. The rectangular area identification unit 306 may estimate a skeleton 540 of the target person 50, including a point 542 indicating the right ankle of the target person 50, a point 546 indicating the right knee, a point 544 indicating the left ankle, and a point 548 indicating the left knee. The rectangular area identification unit 306 may estimate the skeleton of the target person 50 using well-known techniques.

[0058] The rectangular area identification unit 306 may identify a midpoint 552 between points 542 and 544 and adjust the vertical position of the midpoint 552 based on a shin length 550 calculated from points 542 and 546 or points 544 and 548, thereby identifying the adjustment point 554. For example, the rectangular area identification unit 306 may compare the distance between points 542 and 546 and the distance between points 544 and 548 in the captured image 520 and determine the longer distance as the shin length 550. In the example shown in FIG. 6 , the distance between points 542 and 546 may be determined as the shin length 550. The rectangular area identification unit 306 may identify the adjustment point 554 by moving the midpoint 552 downward by a length that is a predetermined percentage of the identified shin length 550. The predetermined percentage may be set in advance or may be changeable. An example of the predetermined ratio is, but is not limited to, 30%. The rectangular area identification unit 306 may determine the adjustment point 554 as the center of gravity of the target person 50's ground surface. Note that the rectangular area identification unit 306 may also determine the midpoint 552 as the center of gravity of the target person 50's ground surface.

[0059] When the rectangular area specification unit 306 is unable to specify the ground contact points of both feet of the target person 50 but can specify the positions of both ankles through posture estimation of the target person 50, it may use the adjustment point 554 as the center of gravity of the ground contact surface of the target person 50. The midpoint 552 is the midpoint of both ankles and is therefore spaced upward from the ground contact surface, but the adjustment point 554 is brought closer to the ground contact surface through adjustment. Because multiple reference points are disposed on the ground contact surface, using the adjustment point 554 provides higher accuracy than using the midpoint 552 when performing projective transformation using multiple reference points. Therefore, using the adjustment point 554 as the center of gravity can improve the accuracy of the projective transformation and increase the degree of agreement between the range of the target person 50 in the captured image 600 and the second rectangular area. If the ground contact points of the right foot and the left foot of the target person 50 can be identified by estimating the posture of the target person 50, the rectangular area identification unit 306 may use the midpoint 552 as the center of gravity of the target person 50's ground contact surface.

[0060] When the rectangular area identification unit 306 identifies the center of gravity of the target person 50 using the above-described method, accurate identification may be difficult if the distance between the right foot and left foot of the target person 50 in the captured image 520 is too close. In response to this, the image acquisition unit 304 may select a captured image to be processed from multiple captured images of the target captured continuously by the imaging unit 100. For example, the paired image acquisition unit 304 selects a captured image in which the distance between the right foot and left foot of the target is greatest. Specifically, the image acquisition unit 304 selects a captured image in which the distance between the right foot and left foot of the target is greatest from among multiple captured images of the target captured continuously by the imaging unit 100. This improves the accuracy with which the rectangular area identification unit 306 identifies the center of gravity of the target person 50.

[0061] 7 schematically illustrates an example of a processing flow of the information processing device 300. Here, a processing flow will be described from when the information processing device 300 acquires the captured image 500 and the captured image 600 to when the information processing device 300 performs annotation on the captured image 600. The information processing device 300 may execute the processing illustrated in FIG. 7 every time it acquires the captured image 500 and the captured image 600.

[0062] In step (step may be abbreviated as S) 102, the image acquisition section 304 acquires the captured image 500 and the captured image 600. The image acquisition section 304 may acquire the captured image 500 from the imaging section 100, and may acquire the captured image 600 from the imaging section 200.

[0063] In S104, the rectangular area identification unit 306 identifies a first rectangular area in the captured image 500 and identifies the center of gravity of the target person 50. In S106, the ground surface area identification unit 308 identifies the ground surface area using the first rectangular area identified in S104.

[0064] In S108, the projection transformation unit 310 projectively transforms the installation surface range identified in S106 and the center of gravity identified in S104 into the captured image 600. In S110, the rectangular area determination unit 312 determines a second rectangular area that indicates the range of the target person 50 in the captured image 600. The rectangular area determination unit 312 determines the size of the second rectangular area using the contact surface range projectively transformed onto the captured image 600, and sets the center of gravity projectively transformed onto the captured image 600 as the center point of the second rectangular area.

[0065] In S112, the annotation unit 314 assigns metadata indicating a person to the second rectangular area of ​​the captured image 600 determined in S110, and then the process ends.

[0066] 8 schematically illustrates an example of the hardware configuration of a computer 1200 functioning as the information processing device 300. A program installed on the computer 1200 can cause the computer 1200 to function as one or more "units" of an apparatus according to the present embodiment, or can cause the computer 1200 to execute operations associated with the apparatus according to the present embodiment or one or more "units," and / or can cause the computer 1200 to execute a process according to the present embodiment or steps of the process. Such a program can be executed by the CPU 1212 to cause the computer 1200 to perform specific operations associated with some or all of the blocks in the flowcharts and block diagrams described herein.

[0067] The computer 1200 according to this embodiment includes a CPU 1212, a RAM 1214, and a graphics controller 1216, which are interconnected by a host controller 1210. The computer 1200 also includes input / output units such as a communications interface 1222, a storage device 1224, a DVD drive, and an IC card drive, which are connected to the host controller 1210 via an input / output controller 1220. The DVD drive may be a DVD-ROM drive, a DVD-RAM drive, or the like. The storage device 1224 may be a hard disk drive, a solid-state drive, or the like. The computer 1200 also includes a ROM 1230 and legacy input / output units such as a keyboard, which are connected to the input / output controller 1220 via an input / output chip 1240.

[0068] The CPU 1212 operates according to programs stored in the ROM 1230 and the RAM 1214, thereby controlling each unit. The graphics controller 1216 acquires image data generated by the CPU 1212 into a frame buffer or the like provided in the RAM 1214 or into the graphics controller itself, and causes the image data to be displayed on the display device 1218.

[0069] The communication interface 1222 communicates with other electronic devices via a network. The storage device 1224 stores programs and data used by the CPU 1212 in the computer 1200. The DVD drive reads programs or data from a DVD-ROM or the like and provides them to the storage device 1224. The IC card drive reads programs and data from an IC card and / or writes programs and data to an IC card.

[0070] The ROM 1230 stores therein a boot program or the like that is executed by the computer 1200 upon activation, and / or programs that depend on the hardware of the computer 1200. The input / output chip 1240 may also connect various input / output units to the input / output controller 1220 via a USB port, a parallel port, a serial port, a keyboard port, a mouse port, etc.

[0071] The programs are provided by a computer-readable storage medium such as a DVD-ROM or an IC card. The programs are read from the computer-readable storage medium, installed in the storage device 1224, RAM 1214, or ROM 1230, which are also examples of computer-readable storage media, and executed by the CPU 1212. Information processing described in these programs is read by the computer 1200, and causes cooperation between the programs and the various types of hardware resources described above. An apparatus or method may be configured by implementing operations or processing of information in accordance with the use of the computer 1200.

[0072] For example, when communication is performed between the computer 1200 and an external device, the CPU 1212 may execute a communication program loaded into the RAM 1214 and instruct the communication interface 1222 to perform communication processing based on the processing described in the communication program. Under the control of the CPU 1212, the communication interface 1222 reads transmission data stored in a transmission buffer area provided in the RAM 1214, the storage device 1224, a DVD-ROM, or a recording medium such as an IC card, and transmits the read transmission data to the network, or writes reception data received from the network to a reception buffer area or the like provided on the recording medium.

[0073] Furthermore, the CPU 1212 may cause all or a necessary portion of a file or database stored in an external recording medium such as the storage device 1224, a DVD drive (DVD-ROM), an IC card, etc. to be read into the RAM 1214, and may perform various types of processing on the data on the RAM 1214. The CPU 1212 may then write back the processed data to the external recording medium.

[0074] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and may undergo information processing. The CPU 1212 may perform various types of processing on data read from the RAM 1214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search / replacement, etc., as described throughout this disclosure and specified by the instruction sequences of the programs, and write the results back to the RAM 1214. The CPU 1212 may also search for information in a file, database, etc. on the recording medium. For example, if multiple entries, each having an attribute value of a first attribute associated with an attribute value of a second attribute, are stored on the recording medium, the CPU 1212 may search for an entry whose attribute value of the first attribute matches a specified condition from among the multiple entries, read the attribute value of the second attribute stored in the entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.

[0075] The above-described programs or software modules may be stored in a computer-readable storage medium on or near the computer 1200. A recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can also be used as a computer-readable storage medium, thereby providing the programs to the computer 1200 via the network.

[0076] The blocks in the flowcharts and block diagrams in the present embodiments may represent stages of a process in which an operation is performed or "parts" of an apparatus responsible for performing the operation. Particular stages and "parts" may be implemented by dedicated circuitry, programmable circuitry provided with computer-readable instructions stored on a computer-readable storage medium, and / or a processor provided with computer-readable instructions stored on a computer-readable storage medium. The dedicated circuitry may include digital and / or analog hardware circuits, including integrated circuits (ICs) and / or discrete circuits. The programmable circuitry may include reconfigurable hardware circuits, such as field programmable gate arrays (FPGAs) and programmable logic arrays (PLAs), including AND, OR, XOR, NAND, NOR, and other logical operations, flip-flops, registers, and memory elements.

[0077] A computer-readable storage medium may include any tangible device capable of storing instructions that are executed by an appropriate device, such that a computer-readable storage medium having instructions stored thereon comprises an article of manufacture, including instructions that can be executed to create means for performing the operations specified in the flowcharts or block diagrams. Examples of computer-readable storage media may include electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, etc. More specific examples of computer-readable storage media may include floppy disks, diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), Blu-ray disc, memory stick, integrated circuit card, etc.

[0078] The computer readable instructions may include either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk®, JAVA®, C++, etc., and conventional procedural programming languages ​​such as the “C” programming language or similar programming languages.

[0079] Computer-readable instructions may be provided to a general-purpose computer, a special-purpose computer, or another programmable data processing device, or a programmable circuit, either locally or via a local area network (LAN) or a wide area network (WAN) such as the Internet, so that the processor of the programmable data processing device, such as a computer, or the programmable circuit executes the computer-readable instructions to generate means for performing the operations specified in the flowcharts or block diagrams. Here, the computer may be a personal computer (PC), a tablet computer, a smartphone, a workstation, a server computer, a general-purpose computer, a special-purpose computer, or the like, or may be a computer system in which multiple computers are connected. Such a computer system in which multiple computers are connected is also called a distributed computing system, and is a broad definition of computers. In a distributed computing system, multiple computers collectively execute a program by each executing a portion of the program and passing data between computers as needed during program execution.

[0080] Examples of processors include computer processors, central processing units, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc. A computer may have one processor or multiple processors. In a multiprocessor system with multiple processors, each processor executes a portion of a program and passes data between processors as needed during program execution, allowing the multiple processors to collectively execute a program. For example, in multitasking, each of the multiple processors may execute a portion of each task in small chunks by switching tasks at time slice intervals. In this case, which portion of a program each processor executes changes dynamically. Which portion of a program each of the multiple processors executes may also be statically determined by multiprocessor-aware programming.

[0081] Although the present invention has been described above using the embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.

[0082] It should be noted that the execution order of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a later process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order.

[0083] Although the present invention has been described above using the embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.

[0084] It should be noted that the execution order of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a later process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order. [Explanation of symbols]

[0085] 10 system, 20 network, 50 target person, 60 ground surface, 71, 72, 73, 74 reference point, 100 imaging unit, 200 imaging unit, 300 information processing device, 302 memory unit, 304 image acquisition unit, 306 rectangular range specification unit, 308 ground surface range specification unit, 310 projective transformation unit, 312 rectangular range determination unit, 314 annotation unit, 316 learning execution unit, 500 captured image, 510 bounding box, 511, 512, 513, 514, 515 point, 518 ground surface range, 520 captured image, 530 bounding box, 540 skeleton, 542, 544, 546, 548 point, 550 shin length, 552 midpoint, 554 adjustment point, 600 captured image, 612, 614 Difference, 622, 624 Rectangle information, 1200 Computer, 1210 Host controller, 1212 CPU, 1214 RAM, 1216 Graphics controller, 1218 Display device, 1220 Input / output controller, 1222 Communication interface, 1224 Storage device, 1230 ROM, 1240 Input / output chip

Claims

1. an image acquisition unit that acquires a first captured image in which a first imaging unit captures an image of a person and a ground surface on which the person is standing in a manner in which the person's torso, limbs, and head are captured, and a second captured image in which a second imaging unit that captures an image of the person from a direction different from that of the first imaging unit captures an image of the person in a manner in which the shape of the person captured in the first captured image is different; a rectangular area specifying unit that analyzes the first captured image and specifies a first rectangular area that indicates an area of ​​the person in the first captured image; a ground surface range specifying unit that specifies a ground surface range indicating a range of the person within the ground surface in the first captured image based on the first rectangular range; a projection transformation unit that performs projection transformation of the ground contact area in the first captured image into the second captured image using a plurality of predetermined reference points of the ground contact area in the first captured image and the plurality of reference points of the ground contact area in the second captured image; a rectangular area determination unit that determines a second rectangular area indicating an area of ​​the person in the second captured image based on the ground plane area projectively transformed into the second captured image by the projective transformation unit; An information processing device comprising:

2. 2. The information processing device according to claim 1, wherein the second imaging unit captures the person from a direction in which the object to be imaged cannot be automatically recognized as a person by capturing the second captured image in a manner that is different in shape from the person captured in the first captured image.

3. an image acquisition unit that acquires a first captured image in which a first imaging unit captures an image of the imaging object and a ground surface on which the imaging object is grounded, in a manner in which the imaging object can be automatically recognized based on an annotated image that is an image of an object of the same type as the imaging object and that has been annotated in advance as learning data, and a second captured image in which a second imaging unit captures an image of the imaging object in a manner in which the imaging object cannot be automatically recognized based on the annotated image; a rectangular area specifying unit that analyzes the first captured image and specifies a first rectangular area that indicates a range of the image capture target object in the first captured image; a ground surface range specifying unit that specifies a ground surface range indicating a range of the image capture object within the ground surface in the first captured image based on the first rectangular range; a projection transformation unit that performs projection transformation of the ground contact area in the first captured image into the second captured image using a plurality of predetermined reference points of the ground contact area in the first captured image and the plurality of reference points of the ground contact area in the second captured image; a rectangular area determination unit that determines a second rectangular area indicating the area of ​​the image capture target object in the second captured image based on the contact surface area projected into the second captured image by the projective transformation unit; An information processing device comprising:

4. an image acquisition unit that acquires a first captured image of the object and the ground surface by a first imaging unit that images the object and the ground surface from a position other than directly above the object and a diagonally upward position substantially equal to directly above the object, and acquires a second captured image of the object and the ground surface by a second imaging unit that images the object from a diagonally upward position substantially equal to directly above the object; a rectangular area specifying unit that analyzes the first captured image and specifies a first rectangular area that indicates a range of the object in the first captured image; a ground surface range specifying unit that specifies a ground surface range indicating a range of the object within the ground surface in the first captured image based on the first rectangular range; a projection transformation unit that performs projection transformation of the ground contact area in the first captured image into the second captured image using a plurality of predetermined reference points of the ground contact area in the first captured image and the plurality of reference points of the ground contact area in the second captured image; a rectangular area determination unit that determines a second rectangular area indicating the area of ​​the object in the second captured image based on the contact surface area projected into the second captured image by the projection transformation unit; An information processing device comprising:

5. A program for causing a computer to function as the information processing device according to any one of claims 1 to 4 when executed by the computer.

6. An information processing device according to any one of claims 1 to 4; the first imaging unit; the second imaging unit; A system comprising:

7. 1. A computer-implemented information processing method, comprising: an image acquisition step of acquiring a first captured image in which a first imaging unit captures an image of the person and a ground surface on which the person is standing in a manner in which the person's torso, limbs, and head are captured, and a second captured image in which a second imaging unit, which images the person from a direction different from that of the first imaging unit, captures an image of the person in a manner in which the shape of the person captured in the first captured image is different; a rectangular area specifying step of analyzing the first captured image and specifying a first rectangular area indicating an area of ​​the person in the first captured image; a ground surface area specifying step of specifying a ground surface area indicating an area of ​​the person within the ground surface in the first captured image based on the first rectangular area; a projection transformation step of projecting the ground surface range in the first captured image into the second captured image using a plurality of predetermined reference points of the ground surface in the first captured image and the plurality of reference points of the ground surface in the second captured image; a rectangular area determination step of determining a second rectangular area indicating an area of ​​the person in the second captured image based on the ground plane area projectively transformed into the second captured image in the projective transformation step; An information processing method comprising:

8. 1. A computer-implemented information processing method, comprising: an image acquisition step in which a first imaging unit acquires a first captured image of the imaging target and a ground surface on which the imaging target is grounded, in a manner in which the first imaging unit can automatically recognize the imaging target based on an annotated image that is an image of an object of the same type as the imaging target and that has been annotated in advance as learning data, and a second captured image of the imaging target, in a manner in which the second imaging unit cannot automatically recognize the imaging target based on the annotated image; a rectangular area specifying step of analyzing the first captured image and specifying a first rectangular area indicating an area of ​​the image capture target in the first captured image; a ground surface range specifying step of specifying a ground surface range indicating a range of the image capture target object within the ground surface in the first captured image based on the first rectangular range; a projection transformation step of projecting the ground surface range in the first captured image into the second captured image using a plurality of predetermined reference points of the ground surface in the first captured image and the plurality of reference points of the ground surface in the second captured image; a rectangular area determination step of determining a second rectangular area indicating the area of ​​the object to be imaged in the second captured image based on the ground plane area projectively transformed into the second captured image in the projective transformation step; An information processing method comprising:

9. 1. A computer-implemented information processing method, comprising: an image acquisition step in which a first image capturing unit captures an image of the object and the ground surface on which the object is in contact from a position other than directly above the object or a position diagonally above that is substantially equal to directly above the object, and acquires a first image of the object and the ground surface; and a second image capturing unit captures an image of the object and the ground surface from a position diagonally above that is substantially equal to directly above the object, and acquires a second image of the object and the ground surface; a rectangular area specifying step of analyzing the first captured image and specifying a first rectangular area indicating an area of ​​the object in the first captured image; a ground surface range specifying step of specifying a ground surface range indicating a range of the object within the ground surface of the object in the first captured image based on the first rectangular range; a projection transformation step of projecting the ground surface range in the first captured image into the second captured image using a plurality of predetermined reference points of the ground surface in the first captured image and the plurality of reference points of the ground surface in the second captured image; a rectangular area determination step of determining a second rectangular area indicating the area of ​​the object in the second captured image based on the ground plane area projectively transformed into the second captured image in the projective transformation step; An information processing method comprising: