Information processing device, information processing method, and recording medium
By learning a pose estimation model with occlusion-inducing processes and fine-tuning, the method enhances the accuracy of posture estimation in images with hidden objects.
Patent Information
- Application Number
- PCT/JP2024/001490
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-19
- Publication Date
- 2025-07-24
Smart Images

Figure JP2024001490_24072025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and recording medium
[0001] The present disclosure relates to the technical fields of an information processing device, an information processing method, and a recording medium.
[0002] Known examples of this type of device include one that estimates the posture of an object included in an image. For example, Patent Document 1 discloses a technology that uses AI to estimate the posture of a person represented by 2D data based on input 2D data.
[0003] International Publication No. 2023 / 095667
[0004] An object of this disclosure is to provide an information processing device, an information processing method, and a recording medium that aim to improve upon the techniques disclosed in prior art documents.
[0005] One aspect of the information processing device disclosed herein includes an acquisition means for acquiring an image including an object, a processing means for executing at least one of a first process of partially masking the image and a second process of copying the object and pasting the objects so that they overlap each other as a predetermined process on the image, and a first learning means for learning an estimation model that estimates the posture of the object in the image by using the image that has been subjected to the predetermined process as learning data.
[0006] One aspect of the information processing method disclosed herein is to use at least one computer to acquire an image including an object, and to perform at least one of a first process of partially masking the image and a second process of copying the object and pasting the object so that the objects overlap each other as a predetermined process on the image, and to use the image that has been subjected to the predetermined process as training data to learn an estimation model that estimates the posture of the object in the image.
[0007] One aspect of the recording medium of this disclosure is a recording medium having recorded thereon a computer program for causing at least one computer to execute an information processing method, which includes acquiring an image including an object, performing at least one of a first process of partially masking the image and a second process of copying the object and pasting the object so that the objects overlap each other as a predetermined process on the image, and using the image that has been subjected to the predetermined process as training data to learn an estimation model that estimates the posture of the object in the image.
[0008] 1 is a block diagram showing a hardware configuration of a first information processing device. FIG. 2 is a block diagram showing a functional configuration of the first information processing device. FIG. 3 is a plan view showing an example of a first process executed in the first information processing device. FIG. 4 is a plan view showing an example of a second process executed in the first information processing device. FIG. 5 is a flowchart showing a flow of a learning operation by the first information processing device. FIG. 6 is a block diagram showing a functional configuration of a second information processing device. FIG. 7 is a flowchart showing a flow of a learning operation by the second information processing device. FIG. 8 is a conceptual diagram showing a specific example of a learning operation in the second information processing device. FIG. 9 is a plan view showing an example of processing of a first output layer in the second information processing device. FIG. 10 is a plan view (part 1) showing an example of processing of a second output layer in the second information processing device. FIG. 11 is a plan view (part 2) showing an example of processing of a second output layer in the second information processing device. FIG. 12 is a conceptual diagram showing a specific example of a second learning operation in the second information processing device. FIG. 13 is a block diagram showing a functional configuration of a modified example of the third information processing device. FIG. 14 is a flowchart showing a flow of a posture estimation operation by the third information processing device.
[0009] Hereinafter, embodiments of an information processing device, an information processing method, and a recording medium will be described with reference to the drawings.
[0010] First Embodiment A first information processing apparatus will be described with reference to FIGS. 1 to 5. FIG.
[0011] (Hardware Configuration) First, the hardware configuration of the first information processing apparatus will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the hardware configuration of the first information processing apparatus.
[0012] 1, the first information processing device 10 includes a processor 11, a RAM (Random Access Memory) 12, a ROM (Read Only Memory) 13, and a storage device 14. The first information processing device 10 may further include an input device 15 and an output device 16. The processor 11, RAM 12, ROM 13, storage device 14, input device 15, and output device 16 are connected to each other via a data bus 17. The data bus 17 may be an interface other than a data bus (for example, a LAN, a USB, etc.).
[0013] The processor 11 loads a computer program. For example, the processor 11 is configured to load a computer program stored in at least one of the RAM 12, the ROM 13, and the storage device 14. Alternatively, the processor 11 may load a computer program stored in a computer-readable storage medium using a storage medium reading device (not shown). The processor 11 may acquire (i.e., load) the computer program from a device (not shown) located outside the first information processing device 10 via a network interface. The processor 11 controls the RAM 12, the storage device 14, the input device 15, and the output device 16 by executing the loaded computer program. In particular, in this embodiment, when the processor 11 executes the loaded computer program, functional blocks for executing processing related to an estimation model that estimates the posture of an object included in an image are realized within the processor 11. In other words, the processor 11 may function as a controller that executes each control in the first information processing device 10.
[0014] The processor 11 may be configured as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or a quantum processor. The processor 11 may be configured as one of these, or may be configured to use multiple processors in parallel.
[0015] The RAM 12 temporarily stores computer programs executed by the processor 11. The RAM 12 temporarily stores data that the processor 11 temporarily uses while it is executing the computer programs. The RAM 12 may be, for example, a dynamic random access memory (D-RAM) or a static random access memory (SRAM). Alternatively, other types of volatile memory may be used instead of the RAM 12.
[0016] The ROM 13 stores computer programs executed by the processor 11. The ROM 13 may also store fixed data. The ROM 13 may be, for example, a programmable read-only memory (PROM) or an erasable read-only memory (EPROM). Alternatively, other types of non-volatile memory may be used instead of the ROM 13.
[0017] The storage device 14 stores data that is to be saved long-term by the first information processing device 10. The storage device 14 may operate as a temporary storage device for the processor 11. The storage device may store computer programs executed by the processor 11. The storage device 14 may include, for example, at least one of a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device.
[0018] The input device 15 is a device that receives input instructions from a user of the first information processing device 10. The input device 15 may include, for example, at least one of a keyboard, a mouse, and a touch panel. The input device 15 may be configured as part of a smartphone, a tablet terminal, an earphone-type terminal, a watch-type terminal, an HMD (Head Mounted Display) terminal, etc. The input device 15 may be, for example, a device that includes a microphone and is capable of voice input.
[0019] The output device 16 is a device that outputs information related to the first information processing device 10 to the outside. For example, the output device 16 may be a display device (e.g., a display or digital signage) that can display information related to the information processing device 10. The output device 16 may also be a speaker or the like that can output information related to the first information processing device 10 as audio.
[0020] 1 may be configured as devices external to the first information processing device 10. For example, the first information processing device 10 may be configured to include a processor 11, a RAM 12, and a ROM 13, and the other devices, such as a storage device 14, an input device 15, and an output device 16, may be configured as devices external to the first information processing device 10. Furthermore, some of the calculation functions of the first information processing device 10 may be realized by an external server, a cloud, or the like.
[0021] (Functional Configuration) Next, the functional configuration of the first information processing device 10 will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the functional configuration of the first information processing device.
[0022] 2, the first information processing device 10 is configured to include, as processing blocks for realizing its functions, an image acquisition unit 110, an image processing unit 120, and a first learning unit 130. Note that each of the image acquisition unit 110, the image processing unit 120, and the first learning unit 130 may be realized by, for example, the above-mentioned processor 11 (see FIG. 1).
[0023] The image acquisition unit 110 is configured to be able to acquire images including an object. Here, an "object" is something whose posture in an image can be estimated. The object may be, for example, a living thing such as a person or an animal, or a non-living thing such as a car or a desk. The image acquired by the image acquisition unit 110 may include multiple objects. For example, the image acquisition unit 110 may acquire an image in which multiple people are depicted. The image acquisition unit 110 may be configured to extract and acquire an appropriate image from multiple images stored in, for example, an image database. Alternatively, the image acquisition unit 110 may be configured to directly acquire an image taken by a camera from the camera. The image acquired by the image acquisition unit 110 is configured to be output to the image processing unit 120.
[0024] The image processing unit 120 is configured to perform a predetermined process on the image acquired by the image acquisition unit 110. The predetermined process is a process that intentionally hides (i.e., makes portions invisible in the image) objects included in the image. Examples of portions that are invisible in the image include portions hidden by overlapping objects, portions hidden due to their posture, and portions of objects that do not fit within the image frame and are cut off. More specifically, the predetermined process includes at least one of a first process that partially masks an image and a second process that copies and pastes objects so that they overlap. Thus, the image processing unit 120 outputs an image on which the first process has been performed, an image on which the second process has been performed, or an image on which both the first process and the second process have been performed. More specific examples of the predetermined process performed by the image processing unit 120 will be described in detail later. The image on which the predetermined process has been performed by the image processing unit 120 is output to the first learning unit 130.
[0025] The first learning unit 130 is configured to be capable of learning a pose estimation model 200 that estimates the pose of an object in an image. Specifically, the first learning unit 130 performs machine learning of the pose estimation model 200 by using images that have been subjected to a predetermined process by the image processing unit 120 as training data. The training data may be a dataset including multiple pieces of image data. In this case, the dataset may include multiple images that have been subjected to different processes by the image processing unit 120. For example, the dataset may include images that have been subjected to a first process, images that have been subjected to a second process, and images that have been subjected to both the first process and the second process. Furthermore, the dataset may include images that have not been subjected to the predetermined process by the image processing unit 120, in addition to images that have been subjected to the predetermined process by the image processing unit 120. In other words, when multiple images are used as training data, it is sufficient that the predetermined process has been applied to some of the multiple images; it is not necessary that the predetermined process be applied to all of the images.
[0026] The posture estimation model 200 may be configured as a neural network that receives, for example, an image as input and outputs information regarding the posture of an object (hereinafter referred to as "posture information" as appropriate). The posture information output by the posture estimation model 200 may include the orientation of the object as a whole, the orientation of each part included in the object, and the positional relationship of each part included in the object. For example, if the object is a person, the posture estimation model 200 may output the orientation and position of the person's face and limbs, the position and overlap of joints, etc. as posture information. If an image contains multiple objects, the posture estimation model 200 may estimate the postures of each of the multiple objects. Alternatively, the estimation model 200 may estimate the posture of one object selected from the multiple objects. The learning method of the estimation model 200 is not particularly limited, and may be, for example, learning using an error backpropagation method. A more specific configuration of the estimation model 200 will be described in detail in other embodiments described below.
[0027] (Image Processing) Next, image processing in the first information processing device 10 (i.e., predetermined processing executed by the image processing unit 120) will be specifically described with reference to Figures 3 and 4. Figure 3 is a plan view showing an example of first processing executed in the first information processing device. Figure 4 is a plan view showing an example of second processing executed in the first information processing device.
[0028] As shown in FIG. 3 , the image processing unit 120 may perform a first process of partially masking an image. The first process may be, for example, a process of filling some pixels of the image with zeros. A specific technique for the first process may be, for example, a masked autoencoder. When the first process is used as training data, the first learning unit 130 may perform learning to restore the masked portion, for example. By learning using images that have been subjected to the first process, it is possible to learn the general shape of an object hidden by masking.
[0029] As shown in FIG. 4 , the image processing unit 120 may execute a second process of copying objects and pasting them so that they overlap each other. The second process may be a process of duplicating one object, as in the example shown in FIG. 4 . Alternatively, the second process may be a process of duplicating multiple types of objects. In this case, the image processing unit 120 may copy objects from each of multiple images and paste them into one image. For example, the image processing unit 120 may copy a first object included in a first image and a second object included in a second image, and generate a new image in which the first object and the second object are pasted so that they overlap. A specific technique for the second process is, for example, copy-paste data augmentation. By learning using images on which the second process has been executed, it is possible to train a model that is robust to occlusion caused by overlapping objects.
[0030] (Learning Operation) Next, the flow of the learning operation in first information processing device 10 (i.e., a series of operations when first learning unit 130 learns posture estimation model 200) will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the flow of the learning operation by the first information processing device.
[0031] 5, when the learning operation by the first information processing device 10 is started, the image acquisition unit 110 first acquires an image including an object (step S101). Then, the image processing unit 120 performs predetermined processing on the image acquired by the image acquisition unit 110 (step S102). Note that the image acquisition unit 110 may acquire multiple images at once. In this case, the image processing unit 120 may perform predetermined processing on the multiple images at once.
[0032] Next, first learning unit 130 trains pose estimation model 200 using the images that have been subjected to predetermined processing by image processing unit 120 (step S103). First learning unit 130 may input images that are training data to pose estimation model 200 and perform training using the output estimation results (i.e., information about the posture of the object). For example, first learning unit 130 may perform training by updating parameters so that a loss function calculated from the estimation results becomes smaller. First learning unit 130 may repeat the above-described parameter updating process and terminate training when processing for all images that are training data has been completed.
[0033] (Technical Effects) Next, technical effects obtained by the first information processing device 10 will be described.
[0034] As described with reference to FIGS. 1 to 5 , the first information processing device 10 trains the pose estimation model 200 using images that have undergone a predetermined process to intentionally cause occlusion. When an object in an image is occluded, the estimation accuracy of the occluded portion tends to deteriorate compared to the non-occluded portion. However, according to the first information processing device 10, since training is performed using images in which occlusion occurs, it is possible to generate a pose estimation model that is highly robust against occlusion. By using the estimation model 200 trained by the first information processing device 10, it is possible to appropriately estimate the pose of an object, even if the object in the image is occluded, for example.
[0035] Second Embodiment A second information processing device 10 will be described with reference to Figures 6 to 12. The second information processing device 10 differs in some configurations and operations from the first information processing device 10 described above, but other parts may be similar to the first information processing device 10. Therefore, the following will describe in detail the parts that differ from the first embodiment, and will omit explanations of other overlapping parts as appropriate.
[0036] (Functional Configuration) First, the functional configuration of the second information processing device 10 will be described with reference to Fig. 6. Fig. 6 is a block diagram showing the functional configuration of the second information processing device. Note that in Fig. 6, the same elements as those described in Fig. 2 are denoted by the same reference numerals.
[0037] 6, the second information processing device 10 is configured to include, as processing blocks for realizing its functions, an image acquisition unit 110, an image processing unit 120, a first learning unit 130, and a second learning unit 140. That is, the second information processing device 10 further includes the second learning unit 140 in addition to the configuration described in the first embodiment (see FIG. 2). Note that the second learning unit 140 may be realized by, for example, the above-mentioned processor 11 (see FIG. 1).
[0038] Second learning unit 140 is configured to be able to perform learning (hereinafter referred to as "second learning") on posture estimation model 200 trained by first learning unit 130. The second learning performed by second learning unit 140 may be performed as fine adjustment (so-called fine tuning) of parameters after learning by first learning unit 130. In other words, the influence of second learning by second learning unit 140 on the parameters of posture estimation model 200 may be smaller than the influence of learning by first learning unit 130 on the parameters of posture estimation model 200.
[0039] The second learning unit 140 may perform the second learning using images different from the images used as learning data by the first learning unit 130. That is, the second learning unit 140 may separately acquire learning data for the second learning and perform the second learning. Note that the learning data used by the second learning unit 140 may not have undergone the predetermined processing by the image processing unit 120. For example, the second learning unit 140 may use the images acquired by the image acquisition unit 110 as they are for the second learning.
[0040] Second learning unit 140 performs second learning by transferring some of the multiple parameters learned by first learning unit 130. Specifically, second learning unit 140 performs second learning of pose estimation model 200 by transferring parameters that are less susceptible to the influence of occlusion, without transferring parameters that are more susceptible to the influence of occlusion. Specific operations performed during the second learning will be described in detail later.
[0041] (Learning Operation) Next, the flow of the learning operation in second information processing device 10 (i.e., a series of operations when first learning unit 130 and second learning unit 140 learn posture estimation model 200) will be described with reference to Fig. 7. Fig. 7 is a flowchart showing the flow of the learning operation by the second information processing device. Note that in Fig. 7, the same processes as those described in Fig. 5 are denoted by the same reference numerals.
[0042] 7, when the second information processing device 10 starts a learning operation, the image acquisition unit 110 first acquires an image including an object (step S101). Then, the image processing unit 120 performs a predetermined process on the image acquired by the image acquisition unit 110 (step S102).
[0043] Next, first learning unit 130 learns posture estimation model 200 using the images that have been subjected to predetermined processing by image processing unit 120 (step S103). First learning unit 130 performs learning by adjusting parameters included in posture estimation model 200. Note that the images used for learning by first learning unit 130 may not have been subjected to predetermined processing by image processing unit 120. The images used for learning may include images in which an object included in the image is occluded. Therefore, step S102 may not be performed, and image acquisition unit 110 may acquire images in which an object is occluded.
[0044] Next, second learning unit 140 transfers the parameters of posture estimation model 200 trained by first learning unit 130 (step S201). At this time, second learning unit 140 transfers only some of the multiple parameters trained by first learning unit 130. Then, second learning unit 140 performs second learning of posture estimation model 200 using the transferred parameters (step S202).
[0045] (Specific learning operation examples) Next, specific examples of learning operations in the second information processing device 10 will be described with reference to Figs. 8 to 12. Fig. 8 is a conceptual diagram showing a specific example of learning operations in the second information processing device. Fig. 9 is a plan view showing a processing example of the first output layer in the second information processing device. Fig. 10 is a plan view (part 1) showing a processing example of the second output layer in the second information processing device. Fig. 11 is a plan view (part 2) showing a processing example of the second output layer in the second information processing device. Fig. 12 is a conceptual diagram showing a specific example of the second learning operation in the second information processing device.
[0046] 8 , in the learning operation in second information processing device 10, first, image acquisition unit 110 acquires an input image from large-scale database 310. Predetermined processing is performed on the input image acquired from large-scale database 310 by image processing unit 120. Then, the image that has undergone the predetermined processing is used as learning data for training pose estimation model 200.
[0047] Training of posture estimation model 200 is performed separately for backbone 201, first output layer 210, and second output layer 220. First output layer 210 is a layer that performs classification (CLS). Second output layer 220 is a layer that performs regression (REG). In this manner, posture estimation model 200 may be configured as a neural network including a classification task and a regression task. The following describes in detail the operations performed by each of first output layer 210 and second output layer 220.
[0048] As shown in FIG. 9 , the first output layer 210 may be a layer that estimates joint points of an object included in an image. Joint points are parts that affect the posture of the object. Specific parts to be estimated as joint points may be set in advance depending on the type of object, etc. For example, if the object is a person as in the example shown in FIG. 9 , the first output layer 210 may estimate positions corresponding to each part of the face, such as the eyes, nose, ears, and mouth, and joints, such as the neck, shoulders, elbows, wrists, waist, hip joints, knees, and ankles, as joint points.
[0049] As shown in FIG. 10 , the second output layer 220 may be a layer that estimates vectors between the multiple joint points estimated in the first output layer 210 as one of the regression tasks. The second output layer 220 may estimate vectors between the multiple joint points based on the connections between the multiple joint points. For example, as shown in the example of FIG. 10 , the second output layer 220 may estimate a vector between a joint point corresponding to a shoulder and a joint point corresponding to an elbow connected to the shoulder. Similarly, the second output layer 220 may estimate a vector between a joint point corresponding to an elbow and a joint point corresponding to a wrist connected to the elbow. Note that information regarding the connections between the joint points (i.e., information indicating which joint points are connected to each other) may be prepared in advance as, for example, a lookup table.
[0050] As shown in FIG. 11 , the second output layer 220 may be a layer that estimates, as one of the regression tasks, the distances between the multiple joint points estimated in the first output layer 210 and a reference point. The reference point is a point set at a reference position on an object. The reference position may be set appropriately depending on the type of object, etc. For example, the reference point may be set as a point at the center of the object. In the example shown in FIG. 11 , the second output layer 220 estimates the distances between each joint point and a reference point set in the stomach area of a person, which is an object.
[0051] Pose estimation model 200 estimates the pose of an object using the estimation results of the above-described first output layer 210 and second output layer 220. In this way, not only the positions of the joint points but also the positional relationships between the joint points are taken into consideration, making it possible to more appropriately estimate the pose of an object.
[0052] Returning to Fig. 8, learning for the first output layer 210 and the second output layer 220 is performed based on the estimation results output from each layer. For example, learning for the first output layer 210 is performed by calculating a loss function such as Cross-entropy from the estimation results of the first output layer 210 (see Fig. 9) and minimizing the loss function. Learning for the second output layer 220 is performed by calculating a loss function such as L1 from the estimation results of the second output layer 220 (see Figs. 10 and 11) and minimizing the loss function.
[0053] When learning by the first learning unit 130 is completed, among the parameters of the posture estimation model 200, the parameters of the first output layer 210 are not transferred, but the parameters of the main body portion 201 and the second output layer 220 are transferred. The first output layer 210 is a layer that estimates the positions of joint points as described in FIG. 9 . The estimation results of this first output layer 210 are easily affected by occlusion occurring in the object. On the other hand, the second output layer 220 is a layer that estimates the regression of vectors between joint points or distances between reference points and joint points as described in FIGS. 10 and 11 . The estimation results of this second output layer 220 are not easily affected by occlusion occurring in the object. In this way, in the second information processing device 10, parameters that are easily affected by occlusion are not transferred, while parameters that are not easily affected by occlusion are transferred.
[0054] 12 , in the second learning operation of the second information processing device 10, the image acquisition unit 110 acquires input images from the small-scale database 320. The small-scale database 320 is a database smaller in scale than the large-scale database 310 (see FIG. 8 ). The images stored in the small-scale database 320 may be narrowed down in anticipation of use after learning. For example, when the posture estimation model 200 is used to estimate the posture of a person, the small-scale database 320 may be configured as a database that stores images including people. The input images acquired from the small-scale database 320 are used as learning data for the second learning of the posture estimation model 200.
[0055] The second learning of the pose estimation model 200 is performed by adding a third output layer 230 in addition to the first output layer 210 and the second output layer 220. The third output layer 220 is a layer that determines the occlusion state of an object. The third output layer 230, for example, determines whether each of the multiple joint points estimated by the first output layer 210 is occluded or not. Hereinafter, the determination made by the third output layer 230 will be referred to as "occlusion determination" as appropriate. Furthermore, the third output layer 230 may determine the visible state of an object in addition to or instead of the above-described occlusion determination. The third output layer 230 may, for example, determine whether each of the multiple joint points estimated by the first output layer 210 is viewed or not. The determination result of the third output layer 230 may be used in the second learning. That is, the second learning of the pose estimation model 200 may be performed using the determination result regarding the occlusion state or visibility state of an object.
[0056] The second learning for the first output layer 210, the second output layer 220, and the third output layer is performed based on the estimation results output from each layer. For example, the second learning for the first output layer 210 is performed by calculating a loss function such as cross-entropy from the estimation results of the first output layer 210 (see FIG. 9 ) and minimizing the loss function. The second learning for the second output layer 220 is performed by calculating a loss function such as L1 from the estimation results of the second output layer 220 (see FIGS. 10 and 11 ) and minimizing the loss function. The second learning for the third output layer 230 is performed by calculating a loss function such as cross-entropy from the estimation results of the third output layer 230 (i.e., the results of the occlusion determination) and minimizing the loss function.
[0057] As already explained, the second learning of posture estimation model 200 starts with the parameters related to main body portion 201 and second output layer 220 transferred. On the other hand, for first output layer 210, to which parameters are not transferred, and newly added third output layer 230, the second learning starts from a state of preset initial values (e.g., random numbers). As a result, parameters that are less susceptible to the influence of occlusion are trained with a higher priority than parameters that are more susceptible to the influence of occlusion.
[0058] (Technical Effects) Next, technical effects obtained by the second information processing device 10 will be described.
[0059] 6 to 12 , in second information processing device 10, after pose estimation model 200 is trained, some parameters are transferred and second learning is performed. Specifically, of the trained parameters, parameters that are susceptible to the influence of occlusion are not transferred, while parameters that are not susceptible to the influence of occlusion are transferred, and second learning is performed in this state. In this way, training is focused on parameters that are not susceptible to the influence of occlusion, making it possible to further improve the robustness of pose estimation model 200 against occlusion.
[0060] Furthermore, the second learning of pose estimation model 200 is performed after adding third output layer 230. That is, the second learning of pose estimation model 200 is performed while taking into consideration the results of occlusion determination by third output layer 230 for each joint point. Whether an object is occluded or not has a significant impact on the pose estimation results. Therefore, if the second learning is performed while taking into consideration the results of occlusion determination by third output layer 230, it is possible to further improve the robustness of pose estimation model 200 against occlusion.
[0061] <Third Embodiment> A third information processing device 10 will be described with reference to Figures 13 to 15. The third information processing device 10 differs in some configurations and operations from the first and second information processing devices 10 described above, but other parts may be similar to the first and second information processing devices 10. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit explanations of other overlapping parts as appropriate.
[0062] (Functional Configuration) First, the functional configuration of the third information processing device 10 will be described with reference to Fig. 13 and Fig. 14. Fig. 13 is a block diagram showing the functional configuration of the third information processing device. Fig. 14 is a block diagram showing the functional configuration of a modified example of the third information processing device. Note that in Fig. 13 and Fig. 14, the same elements as those described in Fig. 2 and Fig. 6 are denoted by the same reference numerals.
[0063] 13, the third information processing device 10 is configured to include, as processing blocks for realizing its functions, an image acquisition unit 110, an image processing unit 120, a first learning unit 130, and a posture estimation unit 150. That is, the third information processing device 10 further includes a posture estimation unit 150 in addition to the configuration described in the first embodiment (see FIG. 2). Note that the posture estimation unit 150 may be realized by, for example, the above-described processor 11 (see FIG. 1).
[0064] 14, the third information processing device 10 may be configured to include, as processing blocks for realizing its functions, an image acquisition unit 110, an image processing unit 120, a first learning unit 130, a second learning unit 140, and a posture estimation unit 150. That is, the third information processing device 10 may further include a posture estimation unit 150 in addition to the configuration described in the second embodiment (see FIG. 6).
[0065] The posture estimation unit 150 is configured to be able to estimate the posture of an object included in an image using the posture estimation model 200 (see FIG. 13 ) trained by the first learning unit 130 or the posture estimation model 200 (see FIG. 14 ) trained by the second learning unit 140. An image acquired by the image acquisition unit 110 is input to the posture estimation unit 150. Note that the image input to the posture estimation unit 150 is not an image used as learning data as described in the first and second embodiments, but an image including a determination target for which posture estimation is required. The posture estimation unit 150 inputs the image including the determination target to the posture estimation model 200 to acquire posture information related to the determination target. The posture estimation unit 150 may have a function of outputting the estimated posture information. For example, the posture estimation unit 150 may be configured to output the posture information to various devices such as a display.
[0066] 13 and 14 show a configuration capable of executing everything from the learning operation of posture estimation model 200 to the operation of estimating the posture after learning, but when only the operation of estimating the posture is executed, the configurations related to learning and second learning may be omitted as appropriate. For example, third information processing device 10 may be configured as a device including image acquisition unit 110 and posture estimation unit 150 including trained posture estimation model 200, without including image processing unit 120, first learning unit 130, and second learning unit 140.
[0067] (Posture Estimation Operation) Next, the flow of the posture estimation operation in the third information processing device 10 (i.e., a series of operations when the posture estimation unit 150 estimates the posture of an object) will be described with reference to Fig. 15. Fig. 15 is a flowchart showing the flow of the posture estimation operation by the third information processing device.
[0068] 15 , when the posture estimation operation by the third information processing device 10 is started, first, the image acquisition unit 110 acquires an image including a target whose posture is to be determined (step S301). The image acquisition unit 110 may be configured to acquire, for example, an image of the target photographed by a camera in real time.
[0069] Next, posture estimation unit 150 inputs an image including the target to posture estimation model 200 and estimates the posture of the target (step S302). For example, if posture estimation model 200 includes first output layer 210, second output layer 220, and third output layer 230 (see FIG. 12 ), posture estimation unit 150 may estimate the posture of the target based on the estimation results of each of first output layer 210, second output layer 220, and third output layer 230.
[0070] Next, posture estimation unit 150 outputs the estimated posture information of the determination target (step S303). Note that when images including the determination target are continuously acquired (for example, when a video including the determination target is acquired), posture estimation unit 150 may estimate the posture for each of the continuously acquired images and continuously output the posture information. In other words, posture estimation unit 150 may output a transition in the posture of the determination target.
[0071] (Technical Effects) Next, technical effects obtained by the third information processing apparatus 10 will be described.
[0072] As described with reference to FIGS. 13 to 15 , in the third information processing device 10, the posture of the object to be determined is estimated using a trained posture estimation model 200. As described in the first embodiment, the posture estimation model 200 is trained to be highly robust against occlusion. Therefore, even if the object to be determined is occluded, it is possible to accurately estimate the posture of the object to be determined. Furthermore, as described in the second embodiment, if the posture estimation model 200 is trained by transferring some parameters, it is possible to estimate the posture of the object to be determined with higher accuracy.
[0073] The third information processing device 10 can be applied to posture analysis tools, human body matching systems, and the like. For example, the third information processing device 10 can determine the anteroposterior relationship of human body parts due to occlusion. For example, the device determines whether occlusion occurs for each of consecutively acquired images. If it is determined that occlusion does not occur, the device estimates the posture using a posture estimation model trained on images without occlusion. If it is determined that occlusion occurs, the device estimates the posture using a posture estimation model 200 trained on images including images with occlusion. By comparing the posture estimation result of a target that is not occluded at a first timing with the posture estimation result of a target that is occluded at a second timing, the movement distance and position of the human body part that was occluded between the first timing and the second timing can be identified. Furthermore, the third information processing device 10 can perform human body matching while ignoring occlusion areas. Furthermore, the third information processing device 10 can also be applied to a system that performs posture tracking.
[0074] The scope of each embodiment also includes a processing method in which a program that operates the configuration of each embodiment to realize the functions of the above-described embodiments is recorded on a recording medium, the program recorded on the recording medium is read as code, and the program is executed on a computer. In other words, a computer-readable recording medium is also included in the scope of each embodiment. Furthermore, each embodiment includes not only a recording medium on which the above-described program is recorded, but also the program itself.
[0075] Examples of recording media that can be used include floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, magnetic tapes, non-volatile memory cards, and ROMs. Furthermore, the scope of each embodiment is not limited to programs that execute processes by themselves, but also includes programs that execute processes by operating on an OS in conjunction with other software or expansion board functions. Furthermore, the program itself may be stored on a server, and part or all of the program may be downloadable from the server to a user terminal. The program may be provided to the user in, for example, a SaaS (Software as a Service) format.
[0076] <Supplementary Notes> The above-described embodiment may be further described as in the following supplementary notes, but is not limited to the following.
[0077] (Supplementary Note 1) The information processing device described in Supplementary Note 1 is an information processing device that includes: an acquisition means that acquires an image including an object; a processing means that executes, as a predetermined process on the image, at least one of a first process of partially masking the image and a second process of copying the object and pasting the object so that the objects overlap; and a first learning means that uses the image that has been subjected to the predetermined process as learning data to learn an estimation model that estimates the posture of the object in the image.
[0078] (Supplementary Note 2) The information processing device according to Supplementary Note 2 is the information processing device according to Supplementary Note 1, in which the estimation model is configured as a neural network and includes a classification task and a regression task.
[0079] (Supplementary Note 3) The information processing device described in Supplementary Note 3 is the information processing device described in Supplementary Note 2, which has a first output layer that estimates articulation points of the object as the classification task, and a second output layer that estimates vectors between the articulation points or distances between a reference point and the articulation points as the regression task.
[0080] (Supplementary Note 4) The information processing device described in Supplementary Note 4 is the information processing device described in Supplementary Note 3, further comprising a second learning means that does not transfer parameters related to the first output layer among the parameters learned by the first learning means, but transfers parameters related to the main part of the neural network and the second output layer, thereby learning the estimation model.
[0081] (Supplementary Note 5) The information processing device according to Supplementary Note 5 is the information processing device according to Supplementary Note 4, wherein the second learning means performs learning by adding a third output layer to the estimation model, the third output layer estimating the hidden joint points.
[0082] (Supplementary Note 6) The information processing device described in Supplementary Note 6 is the information processing device described in any one of Supplements 1 to 3, further comprising an estimation means for inputting an image to the estimation model learned by the first learning means, and estimating the posture of the object in the image based on an output from the estimation model.
[0083] (Appendix 7) The information processing device described in Appendix 7 is the information processing device described in Appendix 4 or 5, further including an estimation means for inputting an image to the estimation model learned by the second learning means, and estimating the posture of the object in the image based on an output from the estimation model.
[0084] (Supplementary Note 8) The information processing method described in Supplementary Note 8 is an information processing method that, by using at least one computer, acquires an image including an object, performs at least one of a first process of partially masking the image and a second process of copying the object and pasting the objects so that they overlap each other as a predetermined process on the image, and uses the image that has been subjected to the predetermined process as training data to learn an estimation model that estimates the posture of the object in the image.
[0085] (Supplementary Note 9) The recording medium described in Supplementary Note 9 is a recording medium having recorded thereon a computer program for causing at least one computer to execute an information processing method, which includes acquiring an image including an object, performing, as predetermined processing on the image, at least one of a first processing of partially masking the image and a second processing of copying the object and pasting the objects so that they overlap, and using the image that has been subjected to the predetermined processing as training data to learn an estimation model that estimates the posture of the object in the image.
[0086] (Supplementary Note 10) The computer program described in Supplementary Note 10 is a computer program that causes at least one computer to execute an information processing method, which includes acquiring an image including an object, performing, as predetermined processing on the image, at least one of a first processing of partially masking the image and a second processing of copying the object and pasting the object so that the objects overlap each other, and using the image that has been subjected to the predetermined processing as training data to learn an estimation model that estimates the posture of the object in the image.
[0087] This disclosure may be modified as appropriate within the scope that does not contradict the gist or idea of the invention that can be read from the claims and the entire specification, and information processing devices, information processing methods, and recording media that involve such modifications are also included in the technical idea of this disclosure.
[0088] REFERENCE SIGNS LIST 10 Information processing device 11 Processor 12 RAM 13 ROM 14 Storage device 15 Input device 16 Output device 110 Image acquisition unit 120 Image processing unit 130 First learning unit 140 Second learning unit 150 Posture estimation unit 200 Posture estimation model 201 Main body part 210 First output layer 220 Second output layer 230 Third output layer 310 Large-scale database 320 Small-scale database
Claims
1. An information processing apparatus comprising: an acquisition unit that acquires an image including an object; a processing unit that executes at least one of a first process of partially masking and a second process of copying the object and pasting the object so as to overlap each other as a predetermined process on the image; and a first learning unit that learns an estimation model for estimating a pose of the object in the image by using the image subjected to the predetermined process as learning data.
2. The information processing apparatus according to claim 1, wherein the estimation model is configured as a neural network and includes a classification task and a regression task.
3. The information processing apparatus according to claim 2, wherein the estimation model has a first output layer that estimates joint points of the object as the classification task, and a second output layer that estimates a vector between the joint points or a distance between a reference point and the joint points as the regression task.
4. The information processing apparatus according to claim 3, further comprising a second learning unit that learns the estimation model by transferring parameters related to the main body portion of the neural network and parameters related to the second output layer without transferring parameters related to the first output layer among the parameters learned by the first learning unit.
5. The information processing apparatus according to claim 4, wherein the second learning unit adds a third output layer for estimating the hidden joint points to the estimation model and learns.
6. The information processing apparatus according to any one of claims 1 to 3, further comprising an estimation unit that inputs an image to the estimation model learned by the first learning unit and estimates a pose of the object in the image based on an output from the estimation model.
7. The information processing apparatus according to claim 4 or 5, further comprising an estimation unit that inputs an image to the estimation model learned by the second learning unit and estimates a pose of the object in the image based on an output from the estimation model.
8. An information processing method, including: acquiring, by at least one computer, an image including an object; executing, as a predetermined process on the image, at least one of a first process of partially masking and a second process of copying the object and pasting the object so as to overlap each other; and learning an estimation model for estimating a pose of the object in the image by using the image subjected to the predetermined process as learning data.
9. A recording medium on which is recorded a computer program for causing a computer to execute an information processing method, the method including: acquiring an image including an object; performing at least one of a first process of partially masking and a second process of copying the object and pasting the object so that the objects overlap each other as a predetermined process on the image; and learning an estimation model for estimating a posture of the object in the image by using the image subjected to the predetermined process as learning data.
Citation Information
Patent Citations
Human body posture detection method and device based on image recognition and storage medium
CN113326778A
Motion control method and device for virtual object
CN115841534A
Predicting subject body poses and subject movement intent using probabilistic generative models
US20210183073A1