Generation device, generation method and generation program

The generation device enhances object detection models by generating training data that includes obscured parts, allowing accurate recognition and tracking of hidden object regions, thus correcting tracking errors.

JP2025154458AActive Publication Date: 2025-10-10SOFTBANK CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024057473
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2025-10-10
Estimated Expiration
2044-03-29

AI Technical Summary

Technical Problem

Existing object detection models struggle to accurately estimate the positions of hidden parts of a target object, such as a person, due to their training on unobstructed parts, leading to incorrect tracking and movement path calculations when the object is partially obscured.

Method used

A generation device generates training data by identifying and defining rectangular areas that encompass the entire subject in images where parts are hidden, using techniques like image completion and pseudo-obstacle generation to create datasets for improved object detection models.

Benefits of technology

The solution enables object detection models to accurately recognize and track the entire body of a person even when partially obscured, ensuring correct movement path calculations and preventing loss of tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025154458000001_ABST
    Figure 2025154458000001_ABST
Patent Text Reader

Abstract

To provide an object detection model capable of recognizing a position of a hidden part in a detection object included in an image with high accuracy.SOLUTION: A generation device for generating learning data to be used for learning of an object detection model for detecting a prescribed object from an image includes an acquisition unit, a specification unit, and a learning data generation unit. The acquisition unit acquires a first image including a first subject, a portion of which is hidden. The specification unit specifies a first area surrounding the entire part of the first subject in the first image. The learning data generation unit generates image data in which the first area is determined as a correct answer area showing the entire part of the first subject as learning data.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a generation device, a generation method, and a generation program. [Background technology]

[0002] There are known methods for detecting a specific object (e.g., a person) from an image and tracking the detected object. For example, Patent Document 1 discloses a method for estimating a contact point on a ground surface of a subject and a movement line formed by the movement of the subject. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-236569 Summary of the Invention [Problem to be solved by the invention]

[0004] Existing object detection models used to estimate ground contact points have a problem in that the accuracy of estimating ground contact points decreases when a part of the target object (e.g., a person) is hidden by an obstruction. This problem occurs because existing object detection models are trained to recognize only the unobstructed parts of the target object.

[0005] However, the above-mentioned conventional techniques do not take into consideration the improvement of existing object detection models so that the positions of hidden parts can be detected with high accuracy.

[0006] Therefore, the present invention provides a generation device, a generation method, and a generation program that can generate an object detection model that can accurately recognize the position of hidden parts of a detection target included in an image. [Means for solving the problem]

[0007] In order to solve the above problem, one form of generation device according to the present invention is a generation device that generates training data used to train an object detection model that detects a specified object from an image, and includes: an acquisition unit that acquires a first image including a first subject that is partially hidden; an identification unit that identifies a first region in the first image that surrounds the entire first subject; and a training data generation unit that generates image data as the training data, in which the first region is defined as a correct answer region that indicates the entire first subject. [Effects of the Invention]

[0008] According to the present invention, it is possible to generate an object detection model that can accurately recognize the position of a hidden part of a detection target included in an image. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram for explaining the problem underlying the present invention. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of an information processing system according to the embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of the configuration of a generating device according to an embodiment. [Figure 4] FIG. 4 is a diagram showing an example of a method for determining how a body appears in an image. [Figure 5] FIG. 5 is a diagram showing a specific example of the first method for generating training data. [Figure 6] FIG. 6 is a diagram showing a specific example of the second method for generating training data. [Figure 7] FIG. 7 is a flowchart illustrating an example of the operation of the generating device. [Figure 8] FIG. 8 is a sequence diagram illustrating an example of operation of the information processing system according to the embodiment. [Figure 9] FIG. 9 is a hardware configuration diagram showing an example of a computer that realizes the functions of the generating device 100 according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.

[0011] One or more embodiments (including examples, modifications, and application examples) described below can be implemented independently. However, at least a portion of the embodiments described below may be implemented in appropriate combination with at least a portion of another embodiment. These embodiments may include novel features that are different from each other. Therefore, these embodiments may contribute to solving different purposes or problems and may produce different effects from each other.

[0012] Furthermore, in the following embodiments, it is assumed that the subject (detection target) to be detected among the subjects included in an image is a person, and a person detection model is exemplified as an example of an object detection model. However, the proposed technology of the present invention is not limited to systems that detect people, and can also be applied to systems that detect non-people (for example, animals such as dogs and cats). For this reason, in the following embodiments, the expression "part of a person's body" corresponds to "part of a subject," and the expression "the entire body of a person" is an example expression that corresponds to "the entire subject."

[0013] The person detection model here refers to a commonly used existing person detection algorithm, and the position of the person output as the detection result is indicated by a rectangular area, which is represented by position coordinates corresponding to the coordinate system of the image containing the person as a subject.

[0014] (Embodiment) 1. Introduction There is an application that uses a person detection model to predict a rectangular area (detect a person) from an image captured by a fixed camera, estimate the contact point on the person's surface based on the rectangular area, and then perform a projective transformation (bird's-eye view) on the captured image containing information on the contact point to track the same person and record their movement line (path of movement).

[0015] However, in a captured image in which part of a person's body (for example, the feet) is hidden by some object, existing person detection models cannot detect the position coordinates that correctly indicate the person's entire body, and the application may calculate the contact point at an incorrect position. In such a case, the application may lose track of the person whose movement it was originally tracking, and may not be able to obtain a path that correctly tracks the same person.

[0016] To cite a specific example, when a person's body is partially obscured, existing person detection models can only detect the part of the person's body that is visible in the captured image (i.e., the part that is not obscured), i.e., the part that is visible to the viewer of the captured image. In this case, the application calculates the ground contact point at a position different from the intended position (e.g., the position of the part that is visible in the captured image), and when projected onto a bird's-eye view, an incorrect movement path is obtained, in which the ground contact point appears to have moved far away from its actual position. In other words, in a situation where a scene in which a person passes behind some object is captured by a fixed camera, if the person enters behind the object and part of their body is obscured, the application may lose track of the person that it had been tracking and erroneously determine that the movement path is that of a completely different person.

[0017] The above problem will be explained using FIG. 1. FIG. 1 is a diagram for explaining the problem underlying the present invention. FIG. 1 shows a scene in which a fixed camera CA fixed at a predetermined position in a certain space M continuously captures images of a person U moving, and the person U is tracked based on the captured images. In FIG. 1, the problem will be explained by focusing on one captured image of the person U passing directly behind a desk, among the images successively acquired by the continuous capture. In this captured image, part of the person U's body is hidden by an obstruction (the desk).

[0018] For example, when a captured image is input, an existing person detection model (hereinafter referred to as "person detection model M1") detects person U from the input captured image (predicting a rectangular area indicating where person U is located in the captured image). As shown in Figure 1, if part of person U (the lower body) is hidden by an obstruction, person detection model M1 cannot detect the position of the hidden part, and only detects the position of the part that is not hidden and is displayed in the captured image (visible part). As a result, the application will draw a rectangular area AR that surrounds the visible part based on the detected position.

[0019] The application calculates the ground contact point of person U in the captured image based on the rectangular area AR, and so if the rectangular area AR surrounds the entire body of person U, it can calculate the correct position G1 as the ground contact point. On the other hand, if the rectangular area AR surrounds only a part of the body of person U, the application will incorrectly calculate a position G2 different from position G1 as the ground contact point, as shown in Figure 1. For example, the application may calculate a position G2 within the rectangular area AR above position G1 as the ground contact point.

[0020] If an incorrect position is calculated as the touchdown point in this way, the application will not be able to correctly obtain the movement line of person U, and will not be able to track person U as the same person. This point will be explained further using Figure 1.

[0021] The application depicts the movement path of person U in an overhead view obtained by projectively transforming the captured image for which the ground contact point has been calculated. When projectively transforming the captured image for which the correct position G1 has been calculated as the ground contact point, an overhead view is obtained in which an appropriate position within space M is determined to be the position of person U. Specifically, an overhead view is obtained in which the position within space M directly behind the obstruction is the position of person U. As a result, the application can correctly track the movement of person U without losing sight of person U as the camera CA continuously captures the movement of person U.

[0022] On the other hand, if a captured image in which the incorrect position G2 is calculated as the ground contact point is projectively transformed, the resulting bird's-eye view may indicate that the position of a different person is different from person U. Specifically, the resulting bird's-eye view may indicate that the different person is located in another space outside space M, far away from directly behind the obstruction. As a result, while camera CA is continuously capturing images of person U moving, the application may lose track of the person it had been tracking while person U is passing through the obstruction, and may erroneously determine that a different person is moving to a completely different position.

[0023] The reason for such erroneous judgment is thought to be that person detection model M1 is trained to estimate only the position of the visible part of a person's body when a photographic image in which part of the person's body is hidden is input. Specifically, this is thought to be because person detection model M1 is generated based on training data in which a rectangular area surrounding the visible part is labeled as the correct answer area.

[0024] As described above, the application can correctly calculate the ground contact point if the rectangular area encloses the entire body of the person. Therefore, the inventors of the present invention came up with the idea that if a pseudo-image having a rectangular area enclosing the entire body of the person is generated based on an image in which part of the person's body is hidden, and the pseudo-image is used as training data to train the person detection model M1, a new person detection model can be generated that can detect the entire body of the person even when an image in which part of the person's body is hidden is input. Specifically, the inventors thought that if training data is used in which a rectangular area enclosing the entire body when part of the person's body is hidden is labeled as the correct answer area, a new person detection model can be generated that can detect the entire body of the person even when an image in which part of the person's body is hidden is input.

[0025] Based on this idea, the generation device according to the embodiment described in this specification is a generation device that generates training data used to train an object detection model that detects a specified object from an image.The generation device acquires an image including a partially hidden subject (person), identifies a rectangular area in the acquired image that surrounds the entire subject (the entire person), and generates image data as training data in which the identified rectangular area is defined as a correct answer area that indicates the entire subject.

[0026] [2. System Configuration Overview] The configuration of the information processing system 1 will be described using Fig. 2. Fig. 2 is a diagram showing an example of the configuration of the information processing system 1 according to the embodiment. As shown in Fig. 2, the information processing system 1 includes an imaging system 2 and a generating device 100. The imaging system 2 and the generating device 100 are connected to each other via a predetermined communication network (network N) so as to be able to communicate with each other via wired or wireless communication. Note that the information processing system 1 shown in Fig. 2 may include a plurality of imaging systems 2 and a plurality of generating devices 100.

[0027] As shown in FIG. 2, the imaging system 2 may be configured with an imaging device 10, a display control device 11, and a display device 12.

[0028] The imaging device 10 is an imaging device (camera) installed to capture images of a specific indoor space (e.g., inside a store) in order to track the movement of people in the space and estimate their traffic lines. The imaging device 10 may be, for example, an AI camera.

[0029] The display control device 11 superimposes a rectangular area or a grounding point on the captured image acquired by the imaging device 10, and controls the display device 12 to display the superimposed captured image.

[0030] The display device 12 has a screen using, for example, a liquid crystal display, an electroluminescence (EL), a cathode ray tube (CRT), etc. The display device 12 may be compatible with 4K or 8K, or may be formed by a plurality of display devices. The display device 12 displays the captured image controlled to be displayed by the display control device 11.

[0031] The generation device 100 executes information processing (mainly generating learning data) according to the proposed technique of the present invention. The generation device 100 may be implemented as either a local server or a cloud server incorporating a learning function (AI software).

[0032] 3. Configuration of the Generation Device The generating device 100 according to the embodiment will be described with reference to Fig. 3. Fig. 3 is a diagram illustrating an example of the configuration of the generating device 100 according to the embodiment. As shown in Fig. 3, the generating device 100 includes a communication unit 110, a storage unit 120, and a control unit 130.

[0033] <Communication Unit 110> The communication unit 110 is realized by, for example, a network interface card (NIC), etc. For example, the communication unit 110 transmits and receives information to and from the imaging system 2.

[0034] <Storage section 120> The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 120 may store, for example, data and programs related to the information processing according to the embodiment.

[0035] <Control unit 130> The control unit 130 is realized by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like executing various programs (for example, the generation program according to the embodiment) stored in a storage device inside the generating device 100 using RAM as a work area. The control unit 130 is also realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0036] As shown in Fig. 3, the control unit 130 has an estimation unit 131, a determination unit 132, an image generation unit 133, an acquisition unit 134, an identification unit 135, a learning data generation unit 136, and a processing unit 137, and realizes or executes the functions and actions of information processing described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in Fig. 3, and may have other configurations as long as they perform the information processing described below. Furthermore, the connection relationship between the processing units included in the control unit 130 is not limited to the connection relationship shown in Fig. 3, and may be other connection relationships.

[0037] In the control unit 130, the acquisition unit 134 acquires a first image including a partially hidden first subject. The identification unit 135 identifies a first region in the first image that surrounds the entire first subject. The training data generation unit 136 generates, as training data, image data in which the first region is determined as a correct answer region that indicates the entire first subject. This point will be described in more detail below.

[0038] <Estimation part 131> The estimation unit 131 estimates the posture of a person included in an image (original image) used to generate training data. Specifically, the estimation unit 131 estimates 17 key points (skeleton points) present on the person's head, joints, etc. using an arbitrary posture estimation algorithm (hereinafter referred to as "posture estimation model M2"). In other words, posture estimation is a task of estimating 17 key points from a person included in an image.

[0039] The estimation unit 131 can use YOLO (You Only Look Once) as the posture estimation model M2. YOLO (registered trademark) is one of the algorithms used for object detection, and can perform object detection using rectangular regions as preprocessing for posture estimation. For this reason, YOLO handles key points such as "nose," "left-eye," "right-eye," "left-ear," "right-ear," "left-shoulder," and "right-shoulder," and estimates the position coordinates of these key points for a person detected using a rectangular region. Then, lines connecting the key points are drawn and output as a posture estimation result.

[0040] <Determination unit 132> The determination unit 132 determines whether a person included in the original image has a part of their body hidden or their entire body shown. For example, the determination unit 132 determines whether a person included in a partial image extracted as a rectangular area from the original image has a part of their body hidden or their entire body shown, based on the number of key points estimated for the person. For example, the determination unit 132 may determine that the person included in the original image has their entire body shown if the number of key points is 17, and may determine that the person included in the original image has a part of their body hidden if the number of key points is less than 17.

[0041] An example of the operation of the determination unit 132 will be specifically described with reference to Fig. 4. Fig. 4 is a diagram showing an example of a method for determining the extent to which a body is shown in an image. Fig. 4 shows a scene in which it is determined whether a person included in an original image provided for generating learning data has a part of their body hidden or their entire body shown.

[0042] In the determination process for determining whether a part of the body is hidden or the whole body is visible, posture estimation is performed using posture estimation model M2 (YOLO), but person detection is first performed as preprocessing for posture estimation. Specifically, estimation unit 131 inputs the original image to posture estimation model M2. As shown in FIG. 4(a), posture estimation model M2 performs person detection as preprocessing, in which a rectangular area (position coordinates of the person) AR indicating the position of the person in the input original image is predicted.

[0043] Next, the posture estimation model M2 estimates 17 key points from the person included in the partial image (the original image within the rectangular area AR) extracted from the rectangular area AR, as shown in Fig. 4(b). Specifically, the posture estimation model M2 estimates the position of each key point for the person included in the partial image.

[0044] When the keypoint estimation result is obtained by the posture estimation model M2, the determination unit 132 determines whether the person detected in the original image has a part of their body hidden or their entire body shown, based on the number of keypoints included in the partial image. If the number of keypoints is 17, the determination unit 132 determines that the person included in the original image has their entire body shown. On the other hand, if the number of keypoints is less than 17, the determination unit 132 determines that the person included in the original image has a part of their body hidden.

[0045] <Image generation unit 133> Returning to Figure 3, the image generation unit 133 generates an image according to the determination result as to whether a part of the body of a person detected in the original image is hidden or whether the entire body is visible.

[0046] (Method 1) For example, when the image generation unit 133 determines that the entire body of a person detected in the original image is captured, it generates a composite image by compositing a pseudo obstacle into the original image so as to hide a portion of the person's body in the original image. In this case, the acquisition unit 134 acquires this composite image as a first image (an image in which a portion of the body is hidden), and when the entire body of the person (a person detected in the original image) in the composite image is surrounded by a rectangular area, the identification unit 135 identifies the rectangular area as a first area. Then, the training data generation unit 136 generates training data by determining the first area as a correct answer area for the composite image. In this way, the training data generation unit 136 generates a dataset in which even the hidden portion (i.e., the invisible portion) of the person's body is surrounded by a rectangular area from an image in which a portion of the person's body is hidden by a pseudo obstacle. In the following embodiment, this method will be described as "Method 1."

[0047] (Method 2) When it is determined that a body part of a person detected in the original image is hidden, the acquisition unit 134 acquires the original image as a first image (an image in which a body part is hidden). Then, the image generation unit 133 generates a second image including the person with the hidden part completed from the first image. Specifically, the image generation unit 133 generates a second image from the first image as a pseudo image in which the hidden part is reproduced in a pseudo manner. In this case, the identification unit 135 compares a second region (rectangular region) that surrounds the entire body of the person (the person with the hidden part completed) in the second image with the first image, and identifies a first region (rectangular region) that surrounds the entire body of the person in the first image. Then, the training data generation unit 136 generates training data by determining the first region in the first image in which the first region is identified as a correct region. In this way, the training data generation unit 136 generates a dataset in which even the hidden part (i.e., the invisible part) is surrounded by a rectangular region based on the image in which the hidden part is completed in a pseudo manner. In the following embodiment, this method will be described as "method 2."

[0048] <Processing Unit 137> The processing unit 137 may perform various processes using the training data generated by the training data generation unit 136. Specifically, the processing unit 137 generates a new person detection model MX that is an improvement of an existing person detection model M1 based on the training data. The processing unit 137 may also perform inference processing using the person detection model MX.

[0049] FIG. 3 illustrates an example in which the generation device 100 has both a learning function for generating training data and generating a new person detection model MX from the generated training data, and an inference function for performing object detection using the person detection model MX. However, the generation device 100 may be divided into two devices, such as a device 101 having a learning function and another device 102 having an inference function. In such an example, the device 101 is a learner, and the device 102 is a detector. Furthermore, when the generation device 100 is divided into the device 101 and the device 102, the device 101 does not necessarily need to be included in the information processing system 1 shown in FIG. 1 and may be independent from the information processing system 1. On the other hand, the device 102 may be included in the information processing system 1 and perform object detection on captured images acquired from the imaging system 2.

[0050] [4. Specific example of Method 1] FIG. 5 is a diagram illustrating a specific example of method 1 for generating training data. FIG. 5 illustrates a scene in which training data is generated pseudo-wise using a pseudo obstacle as a result of determining that a person detected in an original image Oimg1 is in full-body view. In method 1, Microsoft's Common Objects in Context dataset (MSCOCO dataset) can be used as the original image Oimg1. The MSCOCO dataset consists of image data and annotations corresponding to the image data. Each image data also includes, for each object, information on a segmentation mask that has the category of the object.

[0051] Furthermore, the original image Oimg1 may already include a rectangular area of ​​the person detection result, but if the rectangular area is not included, the generating device 100 may render the rectangular area in the original image Oimg1 by applying the original image Oimg1 to the person detection model M1. In the example of Figure 5, the original image Oimg1 includes a rectangular area AR1 indicating the position of person U1. Furthermore, in the example of Figure 5, the determination process described in Figure 4 has been performed on the original image Oimg1, and it has been determined that person U1's entire body is captured.

[0052] In this state, the image generation unit 133 generates a composite image Simg1 by combining a pseudo obstacle OB1 with the original image Oimg1 (step S51). The image generation unit 133 may receive an instruction from the user and combine an obstacle OB1 of a size or shape according to the instruction at a position according to the instruction. Alternatively, the image generation unit 133 may automatically combine an obstacle OB1 according to a pre-programmed command without receiving an instruction from the user.

[0053] Furthermore, since the position of person U1 is surrounded by rectangular area AR1 in composite image Simg1, the identification unit 135 identifies this rectangular area AR1 as a first area surrounding the entire body of person U1. Then, the learning data generation unit 136 generates image data to be training data based on composite image Simg1 in which the first area has been identified in this way (step S53). Specifically, the learning data generation unit 136 assigns a label L1 to composite image Simg1 indicating that the correct area representing the entire body of person U1 is rectangular area AR1, and uses the composite image Simg1 (image data) to which label L1 has been assigned as training data.

[0054] [5. Example of Method 2] Fig. 6 is a diagram showing a specific example of method 2 for generating training data. Fig. 6 shows a scene in which a part of the body of a person detected in the original image Oimg2 is determined to be hidden, and training data is generated in a pseudo manner using an image in which the hidden part is pseudo-reproduced. In method 2, the MSCOCO dataset can also be used as the original image Oimg2.

[0055] Furthermore, the original image Oimg2 may already include a rectangular area of ​​the person detection result, but if the rectangular area is not included, the generating device 100 may render a rectangular area in the original image Oimg2 by applying the original image Oimg2 to the person detection model M1. In the example of FIG. 6, the original image Oimg2 includes a rectangular area AR21 indicating the position of person U21 and a rectangular area AR22 indicating the position of person U22. Furthermore, in the example of FIG. 6, the determination process described in FIG. 4 is performed on the original image Oimg2 in this state, and it is determined that a part of the body of each of person U21 and person U22 is hidden.

[0056] In this state, the image generation unit 133 executes a process of masking the region sAR of the obstruction that is hiding a part of the body of each of the persons U21 and U22 (step S61). For example, the user may visually confirm the obstruction in the original image Oimg2 and designate the region of the obstruction as the region sAR to be masked, and the image generation unit 133 may mask the designated region sAR. As a result, the image generation unit 133 masks the portion of the region sAR in the original image Oimg2 using the mask image MK, as shown in FIG. 6.

[0057] Next, the image generation unit 133 predicts the parts of the bodies of each of the persons U21 and U22 that are hidden by the mask image MK (essentially, the parts that are hidden by the obstruction), and performs image completion to complete the predicted parts as images.

[0058] The image generation unit 133 may use any image completion algorithm (hereinafter referred to as "image completion model M3") for image completion, and may also accept input of an instruction statement, i.e., a prompt, to be specified for the image completion model M3 from the user (step S62). FIG. 6 shows a screen IN for accepting input of a prompt to be given to the image completion model M3. An example is shown in which the user inputs a prompt with the content "Complete the masked parts of the human body and draw the full body." The area masked by the mask image MK, i.e., the area sAR, corresponds to the completion area, which is the area where the image is depicted by image completion, and the prompt may also include information indicating the area sAR masked by the mask image MK.

[0059] The image generation unit 133 inputs the masked original image Oimg2 and the prompt received from the user into the image completion model M3, thereby outputting a completed image Cimg2 in which the masked portions are completed (step S63). The completed image Cimg2 corresponds to a second image in which the portions of the bodies of the persons U21 and U22 that are hidden by the obstruction (mask image MK) are completed.

[0060] Next, the identification unit 135 uses the person detection model M1 to detect people from the interpolated image Cimg2 and identifies rectangular areas indicating the detection results as second areas surrounding the entire bodies of the people included in the interpolated image Cimg2 (step S64). According to the example of Fig. 6, the lower bodies of people U21 and U22, which were hidden by an obstruction in the original image Oimg2, are complemented in the interpolated image Cimg2. Furthermore, in the interpolated image Cimg2, the entire body of person U21 is surrounded by a rectangular area AR21, and the entire body of person U22 is surrounded by a rectangular area AR22.

[0061] In this state, the identification unit 135 compares the second regions (rectangular regions AR21 and AR22) with the first image (original image Oimg2) to identify the first regions that surround the entire bodies of the people in the first image (step S65). Specifically, the identification unit 135 fits the second regions to the first image so that the coordinates of the first image match the position coordinates of the second regions, thereby identifying the first regions that surround the entire bodies of people U21 and U22 in the original image Oimg2 as the first image. As a result, in the original image Oimg2, even though parts of the bodies of people U21 and U22 are hidden, the entire bodies of people U21 and U22 are appropriately surrounded by rectangular regions AR31 and AR32.

[0062] The learning data generation unit 136 generates image data to be used as learning data based on the original image Oimg2 in which the first regions (rectangular region AR31, rectangular region AR32) have been identified (step S66). Specifically, the learning data generation unit 136 assigns to the original image Oimg2 a label L31 indicating that the correct region showing the entire body of person U21 is rectangular region AR31, and a label L32 indicating that the correct region showing the entire body of person U22 is rectangular region AR32. Then, the learning data generation unit 136 uses the original image Oimg2 (image data) to which the labels L31 and L32 have been assigned as learning data.

[0063] FIG. 6 illustrates an example in which the generating device 100 masks an image in accordance with a user's instructions and performs image completion on the masked image in steps S61 to S63. However, the generating device 100 may automatically perform steps S61 to S63 according to preprogrammed instructions. For example, the image generating unit 133 determines, based on information about a segmentation mask having an object category, an occluding object that hides a portion of the body of each of persons U21 and U22 among the objects included in the original image Oimg2. The image generating unit 133 may then mask the area sAR corresponding to the determined occluding object with a mask image MK, thereby defining the masked area sAR as a completion area in which the image is depicted by image completion. This automated masking process allows the generating device 100 to efficiently generate large amounts of training data.

[0064] [6. Example of operation of the generating device] Fig. 7 is a flowchart showing an example of the operation of the generating device 100. Fig. 7 shows an example of the operation of the generating device 100 from generating training data to generating a person detection model MX from the generated training data.

[0065] The acquisition unit 134 determines whether or not there are any unprocessed captured images (step S701). If there are any unprocessed captured images (step S701; Yes), the acquisition unit 134 acquires one unprocessed captured image (step S702). The captured image here refers to a person image in which a person has been detected using a rectangular area from among the images included in the MSCOCO dataset.

[0066] The determination unit 132 determines whether a person included in the photographed image has a part of the body hidden or has the entire body shown (step S703).

[0067] First, the processing route (a) when it is determined that the whole body is in view will be described. In the processing route (a), learning data is generated using the method 1 described in FIG.

[0068] The image generating unit 133 synthesizes a pseudo image including a person whose body is partially hidden in a pseudo manner (step S704a).

[0069] The identifying unit 135 identifies an area surrounding the entire body in the synthesized pseudo image (step S705a).

[0070] The learning data generation unit 136 labels the region surrounding the whole body as a correct region, and generates and stores the labeled pseudo image as learning data (step S706a).

[0071] Next, we will explain the processing route (b) when it is determined that a part of the body is hidden. In processing route (b), training data is generated using method 2 described in Figure 5.

[0072] The image generating unit 133 generates a pseudo image in which the part hidden by the obstruction is complemented, using the method 2 described with reference to FIG. 6 (step S704b).

[0073] The identification unit 135 identifies an area surrounding the entire body in the original captured image based on the area surrounding the entire body identified in the generated pseudo image and the original captured image in which part of the body is hidden (step S705b).

[0074] The learning data generation unit 136 labels the region surrounding the whole body as a correct region, and generates and stores the labeled original captured image as learning data (step S706b).

[0075] If there are no unprocessed captured images, that is, if sufficient learning data has been accumulated (step S701; No), the processing unit 137 generates a person detection model MX based on the learning data accumulated so far (step S707). For example, the processing unit 137 may train an existing person detection model M1 on the feature amounts contained in the learning data.

[0076] Then, the processing unit 137 stores the generated human detection model MX in the storage unit 120, and ends the processing (step S708).

[0077] [7. Example of operation of the entire system] Fig. 8 is a sequence diagram illustrating an example of operation of the information processing system 1 according to the embodiment. Fig. 8 shows a scene in which the information processing system 1 calculates a ground contact point and displays the calculated ground contact point on the screen.

[0078] 8, the imaging device 10 is fixed at a predetermined position in the space M and is installed in a manner that allows it to capture images of people moving through the space M. The imaging device 10 continuously captures images and sequentially determines whether or not a captured image has been acquired (step S801). While the imaging device 10 has not been able to acquire a captured image (step S801; No), the imaging device 10 waits until a captured image can be acquired.

[0079] On the other hand, if the imaging device 10 has acquired a captured image (step S801; Yes), the imaging device 10 uploads the acquired captured image to the generating device 100 (step S802). The generating device 100 accepts the upload of the captured image (step S803).

[0080] The processing unit 137 of the generating device 100 executes an inference process for detecting a person from the captured image acquired this time using the person detection model MX (step S804). Specifically, the processing unit 137 inputs the captured image to the person detection model MX, thereby outputting the position coordinates of the person in the captured image.

[0081] The processing unit 137 calculates a grounding point on the ground surface of the person detected from the captured image based on the detection result by the person detection model MX, i.e., the position coordinates of the person (step S805). For example, the processing unit 137 can calculate the grounding point from the aspect ratio of a rectangular area corresponding to the position coordinates of the person.

[0082] The processing unit 137 transmits the position coordinates of the person detected from the captured image and information on the ground contact point to the display control device 11 (step S806).

[0083] The display control device 11 depicts a rectangular area indicating the position coordinates of the person and information indicating the grounding point for the currently acquired photographed image (step S807). The display control device 11 also controls the display device 12 to display the photographed image including the rectangular area and the grounding point (step S808).

[0084] The display device 12 displays the captured image including the rectangular area and the grounding point in accordance with the control of the display control device 11 (step S809). When this series of processes is completed, the processes from step S801 onwards are repeated again. That is, every time a captured image is acquired by the imaging device 10, the processes from steps S801 to S809 are performed on the acquired captured image.

[0085] 8. Other Embodiments In the above embodiment, an example has been shown in which the cloud-based generation device 100 performs model learning and inference using the learned model. That is, an example has been shown in which the generation device 100 performs image analysis on a captured image acquired by the imaging device 10.

[0086] However, the imaging device 10 may be configured to perform image analysis by itself by incorporating an AI processor. In such an example, the generating device 100 may deploy the person detection model MX generated by the information processing according to the embodiment to the imaging device 10.

[0087] Furthermore, the imaging device 10 may be configured to function completely as an edge AI, that is, to perform model learning. In such an example, the imaging device 10 has the same functions as the control unit 130 of the generation device 100, and the information processing system 1 does not need to include the generation device 100.

[0088] [9. Hardware Configuration] The generating device 100 according to the embodiment may be realized, for example, by a computer 1000 configured as shown in Fig. 9. Fig. 9 is a hardware configuration diagram showing an example of a computer that realizes the functions of the generating device 100 according to the embodiment. The computer 1000 has a CPU 1100, a RAM 1200, a ROM 1300, an HDD 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.

[0089] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 starts up, programs that depend on the hardware of the computer 1000, and the like.

[0090] The HDD 1400 stores programs executed by the CPU 1100, data used by these programs, etc. The communication interface 1500 receives data from other devices via a predetermined communication network and sends the data to the CPU 1100, and transmits data generated by the CPU 1100 to other devices via the predetermined communication network.

[0091] The CPU 1100 controls an output device such as a display and an input device such as a keyboard via the input / output interface 1600. The CPU 1100 acquires data from the input device via the input / output interface 1600. The CPU 1100 also outputs generated data to the output device via the input / output interface 1600.

[0092] Media interface 1700 reads a program or data stored in recording medium 1800 and provides it to CPU 1100 via RAM 1200. CPU 1100 loads the program or data from recording medium 1800 onto RAM 1200 via media interface 1700 and executes the loaded program. Recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0093] For example, when the computer 1000 functions as the generating device 100 according to the embodiment, the CPU 1100 of the computer 1000 executes programs loaded onto the RAM 1200, thereby realizing the functions of the control unit 130. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800, but as another example, the CPU 1100 may obtain these programs from another device via a predetermined communication network.

[0094] [10. Other] Furthermore, among the processes described in each of the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0095] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0096] Furthermore, the above-described embodiments can be combined as appropriate within the scope of not causing any contradiction in the processing content.

[0097] Although some of the embodiments of the present application have been described in detail above with reference to the drawings, these are merely examples, and the present invention can be implemented in other forms that include the aspects described in the "present invention" section and that have been modified and improved in various ways based on the knowledge of those skilled in the art. [Explanation of symbols]

[0098] 1. Information Processing Systems 10. Imaging device 11 Display control device 12 Display device 100 generator 130 Control Unit 131 Estimation Department 132 Judgment section 133 Image Generation Unit 134 Acquisition Department 135 Specific part 136 Learning Data Generation Unit 137 Processing section

Claims

1. A generation device for generating training data used to train an object detection model for detecting a predetermined object from an image, comprising: an acquisition unit that acquires a first image including a first object that is partially hidden; an identification unit that identifies a first area in the first image that surrounds the entire first subject; a training data generating unit that generates, as the training data, image data in which the first region is determined as a correct region that represents the entire portion of the first subject; A generating device comprising:

2. a determination unit that determines whether a predetermined subject included in an original image is partially hidden or entirely visible; an image generating unit that generates an image according to a determination result of whether the predetermined subject is partially hidden or entirely captured; Further equipped The generating device of claim 1 .

3. The determination unit determines whether the predetermined subject is partially hidden or entirely visible based on the number of skeleton points corresponding to the predetermined subject detected from the original image. The generating device of claim 2 .

4. the acquisition unit acquires the original image as the first image when it is determined that a part of the predetermined subject is hidden; the image generation unit generates a second image including a second object with the hidden portion complemented from the first image; the identification unit compares a second region in the second image that surrounds the entire second subject with the first image, and identifies a first region in the first image that surrounds the entire first subject; The learning data generation unit generates the learning data by defining the first region as the correct region for the original image. The generating device of claim 2 .

5. The identification unit identifies the second region based on an output result when the second image is applied to the object detection model. The generating device of claim 4 .

6. The image generation unit generates the second image by applying a predetermined image completion algorithm to the first image and an instruction statement including a completion area that is a target area of ​​image completion in the first image and content to be depicted in the completion area, the instruction statement instructing to complete the hidden part that exists in the completion area. The generating device of claim 4 .

7. The image generation unit determines an obstruction that hides a part of the predetermined subject from among objects included in the original image, and masks an area corresponding to the determined obstruction, thereby determining the masked area as the complementary area. The generating device of claim 6.

8. When it is determined that the entire subject is captured, the image generation unit generates a composite image by combining a pseudo obstacle with the original image so as to hide a part of the predetermined subject included in the original image; the acquisition unit acquires the composite image as the first image, the specifying unit specifies an area as the first area when the entire portion of the first subject is surrounded by an area in the composite image; The learning data generation unit generates the learning data by defining the first region as the correct region for the composite image. The generating device of claim 2 .

9. A generation method executed by a generation device that generates training data used to train an object detection model that detects a predetermined object from an image, comprising: an acquisition step of acquiring a first image including a first object that is partially hidden; a specifying step of specifying a first region in the first image that surrounds the entire first subject; a training data generating step of generating, as the training data, image data in which the first region is determined as a correct region that represents the entire portion of the first subject; A generation method including:

10. A generation program executed by a generation device that generates training data used to train an object detection model that detects a predetermined object from an image, an acquisition step of acquiring a first image including a partially obscured first object; an identifying step of identifying a first region in the first image that surrounds the entire first subject; a training data generating step of generating, as the training data, image data in which the first region is determined as a correct region that indicates the entire portion of the first subject; A generation program that causes the generation device to execute the above.

Citation Information

Patent Citations

  • Ground point estimation device, ground point estimation method, flow line display system, and server

    JP2009236569A