Information Processing Apparatus, Information Processing Method, and Program

By using learning data with accurate visibility labels and excluding information from invisible key points, the model improves estimation accuracy for key point detection in images, addressing the issue of decreased accuracy due to partially hidden key points.

JP7683784B2Active Publication Date: 2025-05-27NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024066640
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-04-17
Publication Date
2025-05-27
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

The estimation accuracy decreases when images with partially hidden key points are included in the training data, as the model learns to predict the positions of invisible key points based on incomplete image patterns.

Method used

The proposed solution involves acquiring and using learning data with correct labels indicating the visibility and positions of key points, and learning an estimation model that excludes information from invisible key points during training, thereby improving estimation accuracy.

Benefits of technology

This approach enhances the estimation accuracy by focusing on visible key points and reducing the impact of incomplete or inaccurate data from invisible key points, leading to more reliable key point detection in images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683784000001
    Figure 0007683784000001
  • Figure 0007683784000002
    Figure 0007683784000002
  • Figure 0007683784000003
    Figure 0007683784000003
Patent Text Reader

Abstract

To alleviate a problem that estimation accuracy deteriorates when learning data include an image in which some of key-points are invisible, in a technique for extracting a key-point of a body of a person from an image by using a learned model.SOLUTION: An information processing device comprises: a reception unit which receives designation of a position of a key-point in which a person's body is visible in a teacher image; and an acquisition unit which acquires learning data associating the teacher image, a position of each person, a label indicating whether or not each of the plurality of key points of each person's body is visible in the teacher image, and the position of the received key-point with each other.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Patent Document 1 and Non-Patent Document 1 disclose a technique for extracting key points of a human body from an image using a learned model.

[0003] In the technique of Patent Document 1, when an image in which a part of the body is hidden by another object and not visible is used as learning data, the position information of the key points of the invisible part is also given as correct answer data. By doing so, it is described that key points hidden by other objects and not visible can also be detected.

[0004] In the technique of Non-Patent Document 1, for a map obtained by dividing an image into a grid, a map showing the position of a person (the central position of the person) as a likelihood, a map showing the amount of position correction and the size of the person at the map position showing the position of the person, a map showing the relative position for each type of joint at the map position showing the position of the person, a map showing the joint position for each type of joint as a likelihood, and a map showing the amount of joint position correction at the map position showing the joint position are used to configure a neural network that outputs these maps. Then, in the technique of Non-Patent Document 1, an image is used as the input, and the neural network that outputs each of the above maps is used to estimate the joint positions of a person from the image. Note that the technique of Non-Patent Document 1 will be described in more detail below with reference to the drawings.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Non-Patent Documents

[0006]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] In the case of the prior art, there is a problem that the estimation accuracy decreases if an image in which some of the key points are not visible is included in the training data. The reasons are explained below.

[0008] First, as shown in FIG. 1, the training data is data that associates a teacher image including a person with a correct label indicating the position of each of a plurality of key points of the person's body within the teacher image. In the figure, circles indicate the positions of each of the plurality of key points within the teacher image. Note that the types and numbers of the key points shown in the figure are examples and are not limited thereto.

[0009] When using a teacher image in which some of the key points are not visible as training data, in the prior art, as shown in FIG. 2, not only the position of the visible key points within the teacher image but also the position of the invisible key points within the teacher image are used. A correct label is prepared and learned. In FIG. 2, the feet of the person located in the foreground are hidden by an obstacle and not visible. However, the key points of the feet of this person are specified on the obstacle that hides the feet of this person. For example, an operator predicts the position of the invisible key points within the teacher image based on the visible part of the person's body and creates a correct label as shown in FIG. 2.

[0010] When configured in this way, regarding the key points that are not visible, the position of the key point is learned using an image pattern in which the appearance characteristics of the key point are not shown. Also, since the operator predicts the position of the key point that is actually not visible in the image within the teacher image to create the correct label, there is a risk of deviation from the actual position of the key point. For example, for these reasons, in the case of the prior art, when an image in which a part of the key point is not visible is included in the learning data, there is a problem that the estimation accuracy decreases.

[0011] An object of the present invention is to reduce the problem that the estimation accuracy decreases when an image in which a part of the key point of a person's body is not visible is included in the learning data in a technique for extracting the key points of a person's body from an image using a learned model.

Means for Solving the Problem

[0012] According to the present invention, an acquisition means for acquiring learning data in which a teacher image including a person, a correct label indicating the position of each person, a correct label indicating whether each of a plurality of key points of each person's body is visible in the teacher image, and a correct label indicating the position of the key point visible in the teacher image among the plurality of key points within the teacher image are associated; a learning means for learning an estimation model that estimates information indicating the position of each person, information indicating whether each of the plurality of key points of each person included in the processed image is visible in the processed image, and information related to the position of each key point for calculating the position of the key point visible in the processed image within the processed image based on the learning data; A learning device having the above is provided.

[0013] Also, according to the present invention, a computer An acquisition step of acquiring learning data in which a teacher image including a person, a correct label indicating the position of each person, a correct label indicating whether or not each of a plurality of key points of each person's body is visible in the teacher image, and a correct label indicating the position within the teacher image of the key points that are visible in the teacher image among the plurality of key points are associated; A learning step of learning an estimation model for estimating information indicating the position of each person, information indicating whether or not each of a plurality of the key points of each person included in the processing image is visible in the processing image, and information related to the position of each key point for calculating the position within the processing image of the key points that are visible in the processing image, based on the learning data; A learning method for executing the above is provided.

[0014] Also, according to the present invention, A computer is caused to function as an acquisition means for acquiring learning data in which a teacher image including a person, a correct label indicating the position of each person, a correct label indicating whether or not each of a plurality of key points of each person's body is visible in the teacher image, and a correct label indicating the position within the teacher image of the key points that are visible in the teacher image among the plurality of key points are associated; a learning means for learning an estimation model for estimating information indicating the position of each person, information indicating whether or not each of a plurality of the key points of each person included in the processing image is visible in the processing image, and information related to the position of each key point for calculating the position within the processing image of the key points that are visible in the processing image, based on the learning data; A program is provided.

[0015] Also, according to the present invention, An estimation device having an estimation means for estimating the position within the processing image of each of a plurality of key points of each person included in the processing image, using the estimation model learned by the learning device is provided.

[0016] Also, according to the present invention, A computer There is provided an estimation method in which the computer executes an estimation step of estimating the position of each of a plurality of key points of each person included in a processed image using the estimation model learned by the learning device within the processed image.

[0017] Also, according to the present invention, A program is provided that causes a computer to function as an estimation means for estimating the position of each of a plurality of key points of each person included in a processed image using the estimation model learned by the learning device within the processed image.

Advantages of the Invention

[0018] According to the present invention, in the technology of extracting key points of a person's body from an image using a learned model, it is possible to reduce the problem that the estimation accuracy decreases when an image in which some of the key points are not visible is included in the learning data.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Embodiment for Carrying Out the Invention

[0020] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are denoted by the same reference numerals, and the description will be omitted as appropriate.

[0021] "First Embodiment" <Overview> The learning device 10 of the present embodiment reduces the problem that the estimation accuracy decreases when an image in which some of the keypoints are not visible is included in the learning data by learning while excluding the information of the keypoints that are not visible in the image.

[0022] <Features of the technology of the present embodiment> First, while comparing with the technology described in Non-Patent Document 1, the features of the technology of the present embodiment, specifically, the configuration for realizing "learning excluding the information of the keypoints that are not visible in the image" will be described.

[0023] -Technology described in Non-Patent Document 1- First, the technology described in Non-Patent Document 1 will be described. As shown in FIG. 3, in the case of the technology described in Non-Patent Document 1, when an image is input to the neural network, a plurality of data as shown in the figure are output. In other words, the neural network described in Non-Patent Document 1 is composed of a plurality of layers that output a plurality of data as shown in the figure.

[0024] An example of "likelihood of person position", "correction amount of person position", "size", "relative position of keypoint a", and "relative position of keypoint b" among the plurality of data shown in FIG. 3 is shown in FIG. 4. FIG. 5 shows a figure in which an explanation indicating each concept of the data in FIG. 4 is added to the image that is the source of the data in FIG. 4.

[0025] The data of "likelihood of human position" is data indicating the likelihood of the position within the image of the central position of the human body. For example, based on the feature amounts of the appearance of the human body, the human body is detected within the image, and data indicating the likelihood of the central position of the human body is output based on the detection result. As shown in the drawing, in this data, the likelihood that the central position of the human body is located in each of the plurality of grids obtained by dividing the image is shown. Note that the method of dividing the image into a grid pattern is a matter of design, and the number and size of the grids shown in the drawing are merely examples. According to the data shown in FIG. 4, the "third grid from the left and third grid from the bottom" and the "second grid from the right and third grid from the top" are specified as the grids where the central position of the human body is located. When an image including a plurality of persons is input as shown in FIG. 5, the grids where the central positions of the bodies of the plurality of persons are located are specified.

[0026] The data of "correction amount of human position" is data indicating the amount of movement in the x direction and the amount of movement in the y direction from the center of the grid specified as the location of the central position of the human body to the central position of the human body. As shown in FIG. 5, the central position of the human body exists at a certain position within one grid. By using the likelihood of the human position and the correction amount of the human position, the central position of the human body within the image can be specified.

[0027] The data of "size" is data indicating the vertical and horizontal lengths of the rectangular area including the human body.

[0028] The data of "relative position of key points" is data indicating the positions within the image of each of the plurality of key points. Specifically, it indicates the relative positional relationship between each of the plurality of key points and the center of the grid where the central position of the body is located. Note that in FIGS. 4 and 5, the positions of two key points are shown for each person, but the number of key points can be three or more.

[0029] Next, an example of the "likelihood of the position of key point a", "likelihood of the position of key point b", and "correction amount of the position of the key point" among the plurality of data shown in FIG. 3 is shown in FIG. 6. FIG. 7 shows a drawing in which an explanation indicating the concept of each of the data in FIG. 6 is added to the image that is the source of the data in FIG. 6.

[0030] The data of "likelihood of the position of the key point" is data indicating the likelihood of the position within the image of each of a plurality of key points. For example, each key point is detected within the image based on the feature amounts of the appearances of each of the plurality of key points, and data indicating the likelihood of the position of each key point is output based on the detection result. As shown in the illustration, such data is output for each key point. And in this data, the likelihood of each key point being located in each of the plurality of grids obtained by dividing the image is shown. Note that the number of grids shown in the illustration is merely an example. When an image including a plurality of persons is input as shown in FIG. 7, the likelihood of each key point of each of the plurality of persons being located is shown. According to the data shown in FIG. 6, the "fourth grid from the left and the first grid from the bottom" and the "second grid from the right and the fourth grid from the top" are specified as the grids where the key point a is located. Also, the "fourth grid from the left and the fourth grid from the bottom" and the "second grid from the right and the second grid from the top" are specified as the grids where the key point b is located. Note that although the data of two key points is shown in the figure, the number of key points can be three or more. And for each key point, data as described above is output.

[0031] The data of "amount of correction of the position of the key point" is data indicating the amount of movement in the x direction and the amount of movement in the y direction from the center of the grid specified as the position where each of the plurality of key points is located until moving to the position of each key point. As shown in FIG. 7, each key point exists at a certain position within one grid. By using the likelihood of the position of each key point and the amount of correction of the position of each key point, the position of each key point within the image can be specified.

[0032] In the technique described in Non-Patent Document 1, after outputting a plurality of data as described above from the input image, based on the plurality of data and a given correct label, the value of a predetermined loss function is minimized to calculate (learn) the parameters of the estimation model. Also, at the time of estimation, the position of each keypoint within the image is specified by two methods (the relative position from the center position of the lattice shown in FIG. 4, the likelihood and correction amount shown in FIG. 6). For example, the result of integrating the positions calculated by each of the two methods is used as the position of each of the plurality of keypoints. Examples of the integration method include average, weighted average, selection of either one, etc.

[0033] -The technique of the present embodiment- Next, the technique of the present embodiment will be described while comparing it with the technique described in Non-Patent Document 1. As shown in FIG. 8, also in the technique of the present embodiment, when an image is input to a neural network, a plurality of data as shown in the figure are output. In other words, the neural network of the present embodiment is composed of a plurality of layers that output a plurality of data as shown in the figure.

[0034] As is clear from the comparison between FIG. 3 and FIG. 8, the technique of the present embodiment is different from the technique described in Non-Patent Document 1 in that the data of "hidden information" corresponding to each of the plurality of keypoints is included in the output data.

[0035] An example of "likelihood of person position", "correction amount of person position", "size", "hidden information of keypoint a", "relative position of keypoint a", "hidden information of keypoint b", "relative position of keypoint b" among the plurality of data shown in FIG. 8 is shown in FIG. 9. FIG. 10 shows a figure in which an explanation indicating the concept of each data in FIG. 9 is added to the image that is the source of the data in FIG. 9.

[0036] The data of "likelihood of person position", "correction amount of person position" and "size" are the same concepts as those in the technique described in Non-Patent Document 1.

[0037] The data of "hidden information of key points" is data indicating whether each key point is hidden in the image, that is, whether each key point is visible in the image. The state where a key point is not visible in the image includes the state where the key point is located outside the image and the state where the key point is located inside the image but is hidden by other objects (other people and other objects, etc.).

[0038] As shown in FIG. 9, such data is output for each key point. In the illustrated example, a value of "0" is assigned to the visible key points, and a value of "1" is assigned to the invisible key points. In the case of the example shown in FIG. 10, the key point a of the person 1 located in the front is hidden by other objects and is not visible. Therefore, when using the learned neural network of the present embodiment, data with "1" assigned as the hidden information of the key point a of the person 1 is output as shown in FIG. 9.

[0039] Note that although the data of two key points is shown in the figure, the number of key points can be three or more. And for each key point, data as described above is output.

[0040] The data of "relative position of key points" is data indicating the position of each of a plurality of key points in the image. The data of "relative position of key points" in the present embodiment includes the data of the key points shown to be visible in the data of the hidden information of the key points, and is different from the technology described in Non-Patent Document 1 in that it does not include the data of the key points shown to be invisible in the data of the hidden information of the key points. Otherwise, it is the same concept as the technology described in Non-Patent Document 1.

[0041] In the case of the example shown in FIG. 10, the key point a (the key point at the feet) of the person 1 located in the foreground is hidden by other objects and not visible. Therefore, when using the learned neural network of the present embodiment, as shown in FIG. 9, data on the relative position of the key point a that does not include data on the relative position of the key point a of the person 1 is output. The data on the relative position of the key point a shown in FIG. 9 includes only the data on the relative position of the key point a of the person 2 shown in FIG. 10.

[0042] Next, an example of the "likelihood of the position of key point a", "likelihood of the position of key point b", and "correction amount of the position of the key point" among the plurality of data shown in FIG. 8 is shown in FIG. 11. FIG. 12 shows a figure in which an explanation indicating each concept of the data in FIG. 11 is added to the image that is the source of the data in FIG. 11.

[0043] The data of the "likelihood of the position of the key point" is the same concept as the technology described in Non-Patent Document 1. In the case of the example shown in FIG. 12, the key point a of the person 1 located in the foreground is hidden by other objects and not visible. Therefore, when using the learned neural network of the present embodiment, as shown in FIG. 11, data on the likelihood of the position of the key point a that does not include data on the likelihood of the position of the key point a of the person 1 is output. The data on the likelihood of the position of the key point a shown in FIG. 11 includes only the data on the likelihood of the position of the key point a of the person 2 shown in FIG. 12.

[0044] The data of the "correction amount of the position of the key point" is the same concept as the technology described in Non-Patent Document 1. In the case of the example shown in FIG. 12, the key point a (the key point at the feet) of the person 1 located in the foreground is hidden by other objects and not visible. Therefore, when using the learned neural network of the present embodiment, as shown in FIG. 11, data on the correction amount of the position of the key point a that does not include data on the correction amount of the position of the key point a of the person 1 is output.

[0045] As described above, the technology of the present embodiment is different from the technology described in Non-Patent Document 1 at least in terms of outputting data of hidden information for each of a plurality of key points and not outputting data of the positions of key points that are shown to be invisible in the hidden information. And, by having these features that the technology described in Non-Patent Document 1 does not have, the technology of the present embodiment realizes learning excluding information on key points that are not visible in the image.

[0046] <Functional Configuration> Next, the functional configuration of the learning device of the present embodiment will be described. In FIG. 13, an example of a functional block diagram of the learning device 10 will be described. As shown in the figure, the learning device 10 includes an acquisition unit 11, a learning unit 12, and a storage unit 13. Note that, as shown in the functional block diagram of FIG. 14, the learning device 10 may not include the storage unit 13. In this case, an external device configured to be communicable with the learning device 10 includes the storage unit 13.

[0047] The acquisition unit 11 acquires learning data in which a teacher image and a correct label are associated. The teacher image includes a person. The teacher image may include only one person or a plurality of persons. The correct label indicates at least whether each of a plurality of key points of the person's body is visible in the teacher image and the position of the visible key points within the teacher image. The correct label does not indicate the position of the key points that are not visible in the teacher image within the teacher image. Note that the correct label may include other information such as the position of the person and the size of the person. Also, the correct label may be a new correct label obtained by processing the original correct label. For example, a correct label may be a plurality of data shown in FIG. 8 obtained by processing the position of the key points within the teacher image and the hidden information of the key points.

[0048] For example, an operator who creates the correct label may perform operations such as designating only the key points that are visible within the image within the image. And, the operator does not have to perform troublesome operations such as predicting the position of the key points that are hidden by other objects and not visible within the image and designating them within the image.

[0049] The key points may be at least a part of the joint part, a predetermined part (eyes, nose, mouth, navel, etc.), and the end parts of the body (the tip of the head, the tips of the feet, the tips of the hands, etc.). Further, the key points may be other parts. The number and the definition method of the positions of the key points are various and are not particularly limited.

[0050] For example, a large number of learning data are stored in the storage unit 13. And the acquisition unit 11 can acquire the learning data from the storage unit 13.

[0051] The learning unit 12 learns an estimation model based on the learning data. The storage unit 13 stores the estimation model. The estimation model is configured to include the neural network described with reference to FIG. 8. The estimation model outputs a plurality of data shown in FIG. 8. The plurality of data shown in FIG. 8 includes information indicating the position of each person, information indicating whether each of the plurality of key points included in the processed image is visible in the processed image, and information related to the position of each key point for calculating the position of the key point visible in the processed image within the processed image. The information related to the position of each key point indicates the relative position of each key point, the likelihood of the position of each key point, the correction amount of the position of each key point, and the like.

[0052] Then, various estimation processes can be performed using the plurality of data output by the estimation model. For example, the estimation unit (for example, the estimation unit 21 described in the following embodiments) performs a predetermined arithmetic process based on a part of the plurality of data as described with reference to FIGS. 8 to 12. The estimation unit can estimate the position of a keypoint visible in the processed image within the processed image. For example, the estimation unit integrates the position of each keypoint in the processed image identified based on the likelihood of the position of the person (the central position of the person) shown in FIG. 9 and the position of each keypoint indicated by the relative position from the central position, and the position of each keypoint in the processed image identified based on the likelihood and the correction amount of the position of each keypoint shown in FIG. 11, and calculates the result as the position of each of the plurality of keypoints in the processed image. Examples of the integration method include, but are not limited to, average, weighted average, selection of either one, etc.

[0053] The learning unit 12 learns using only the information of the keypoints shown to be visible in the hidden information of the learning data and the position information of the keypoints of the learning data, that is, without using the information of the keypoints shown to be invisible in the hidden information of the learning data and the position information of the keypoints of the learning data. For example, when learning about the position of a keypoint, the learning unit 12 adjusts the parameters of the estimation model so as to minimize the error between the position information of the keypoint output from the estimation model being learned and the position information of the keypoint of the learning data (correct label) for the position on the grid indicating that the keypoint is visible in the learning data.

[0054] Here, a specific example of the learning method by the learning unit 12 will be described.

[0055] The learning unit 12 learns to minimize the error between the map indicating the likelihood of the human position output from the estimation model during learning and the map indicating the likelihood of the human position in the learning data for the data of the likelihood of the human position (central position). Also, for the data of the correction amount of the human position, the human size, and the hidden information of each keypoint, the learning unit 12 learns to minimize the error between the correction amount of the human position, the human size, and the hidden information of each keypoint output from the estimation model during learning and the correction amount of the human position, the human size, and the hidden information of each keypoint in the learning data only for the positions on the grid indicating the human position in the learning data.

[0056] Also, for the data of the relative position of each keypoint, the learning unit 12 learns to minimize the error between the relative position of each keypoint output from the estimation model during learning and the relative position of each keypoint in the learning data only for the positions on the grid that are not hidden by the hidden information of each keypoint in the learning data among the positions on the grid indicating the human position in the learning data.

[0057] Also, for the data of the likelihood of the position of each keypoint, the learning unit 12 learns to minimize the error between the map indicating the likelihood of the position of each keypoint output from the estimation model during learning and the map indicating the likelihood of the position of each keypoint in the learning data. Also, for the data of the correction amount of the position of each keypoint, the learning unit 12 learns to minimize the error between the correction amount of the position of each keypoint output from the estimation model during learning and the correction amount of the position of each keypoint in the learning data only for the positions on the grid indicating the position of each keypoint in the learning data. Since the likelihood of the position of each keypoint in the learning data and the correction amount of the position of the keypoint in the learning data only indicate the visible keypoints, it will naturally be learned only with the visible keypoints.

[0058] In this way, when learning about the positions of key points, the learning unit 12 adjusts the parameters of the estimation model so as to minimize the error between the position information of the key points output from the estimation model during learning and the position information of the key points in the learning data (correct labels) with respect to the positions on the grid indicating that the key points are visible in the learning data.

[0059] An example of the processing flow of the learning device 10 will be described with reference to FIG. 15.

[0060] In S10, the learning device 10 acquires learning data in which teacher images and correct labels are associated. This process is realized by the acquisition unit 11. The details of the process executed by the acquisition unit 11 are as described above.

[0061] In S11, the learning device 10 learns an estimation model using the learning data acquired in S10. This process is realized by the learning unit 12. The details of the process executed by the learning unit 12 are as described above.

[0062] The learning device 10 repeats the loop of S10 and S11 until the end condition is met. The end condition is defined using, for example, the value of a loss function.

[0063] <Hardware Configuration> Next, an example of the hardware configuration of the learning device 10 will be described. Each functional unit of the learning device 10 is realized by an arbitrary combination of hardware and software centered around a CPU (Central Processing Unit) of an arbitrary computer, a memory, a program loaded into the memory, a storage unit such as a hard disk for storing the program (which can store programs downloaded from storage media such as CDs or servers on the Internet in addition to programs stored in advance at the time of shipping the device), and a network connection interface. And it is understood by those skilled in the art that there are various modifications to the realization method and device.

[0064] FIG. 16 is a block diagram illustrating the hardware configuration of the learning device 10. As shown in FIG. 16, the learning device 10 includes a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The learning device 10 may not have the peripheral circuit 4A. Note that the learning device 10 may be configured by a plurality of physically and / or logically separated devices. In this case, each of the plurality of devices can include the above hardware configuration.

[0065] The bus 5A is a data transmission path for the processor 1A, the memory 2A, the peripheral circuit 4A, and the input / output interface 3A to transmit and receive data to and from each other. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. The input device is, for example, a keyboard, a mouse, a microphone, a physical button, a touch panel, etc. The output device is, for example, a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform operations based on their operation results.

[0066] <Operational effects> The estimation model learned by the learning device 10 of the present embodiment has the feature of outputting data of hidden information indicating whether each of a plurality of key points is visible in the image. And the estimation model further has the feature of not outputting the position information of the key points shown to be invisible in the data of the hidden information. Also, when learning the estimation model, the learning device 10 has the feature that only the position information where the key points are visible in the learning data of the position information of the key points needs to be given. The learning device 10 optimizes the parameters of the estimation model based on the result output from such an estimation model and the correct label (learning data). According to such a learning device 10, it is possible to correctly learn excluding the information of the key points that are not visible in the image. As a result, it is possible to reduce the problem that the estimation accuracy decreases when the learning data includes images in which some of the key points are not visible.

[0067] "Second Embodiment" The estimation device of the present embodiment uses the estimation model learned by the learning device of the first embodiment to estimate the position within the image of each of a plurality of key points of each person included in the image. This will be described in detail below.

[0068] FIG. 17 shows an example of a functional block diagram of the estimation device 20. As shown in the figure, the estimation device 20 includes an estimation unit 21 and a storage unit 22. Note that, as shown in the functional block diagram of FIG. 18, the estimation device 20 may not have the storage unit 22. In this case, an external device configured to be communicable with the estimation device 20 includes the storage unit 22.

[0069] The estimation unit 21 acquires an arbitrary image as a processing image. For example, the estimation unit 21 may acquire an image captured by a surveillance camera as a processing image.

[0070] Then, the estimation unit 21 estimates and outputs the positions of a plurality of key points of each person included in the processed image within the processed image using the estimation model learned by the learning device 10. As described in the first embodiment, when an image is input, the estimation model outputs the data described with reference to FIGS. 8 to 11. The estimation unit 21 further performs an estimation process using the data output by this estimation model, thereby estimating the positions of a plurality of key points of each person included in the processed image within the processed image, and outputs the result as an estimation result. The learned estimation model is stored in the storage unit 22. The output of the estimation result is realized using any means such as a display, a projection device, a printer, and e-mail. Further, the estimation unit 21 may output the data output by the estimation model as the estimation result as it is.

[0071] Note that the estimation unit 21 has a feature of estimating whether or not each of a plurality of key points of each person included in the processed image is visible in the processed image using the estimation model, and estimating the positions of a plurality of key points of each person included in the processed image within the processed image using the result of the estimation. Hereinafter, an example of the process performed by the estimation unit 21 will be described with reference to FIGS. 19 and 20.

[0072] (Step 1): Process the processed image with the estimation model to obtain a plurality of data as shown in FIGS. 8 to 11. (Step 2): Based on the data of the likelihood of the person position, identify the grid (P1 in FIG. 19) in which the central position of the person (P11 in FIG. 19) of each person is located (included). Specifically, identify the grid whose likelihood is equal to or greater than the threshold value. (Step 3): Obtain the correction amount (P10 in FIG. 19) corresponding to the position of the grid identified in (Step 2) from the data of the correction amount of the person position. (Step 4): Based on the position of the grid identified in (Step 2) (including the central position of the grid) and the correction amount obtained in (Step 3), identify the central position of the person in the processed image (P11 in FIG. 19) for each person included in the processed image. Thereby, the central position of the body of each person is identified.

[0073] (Step 5): From the size data, obtain the size of the person corresponding to the position of the grid identified in (Step 2). Thereby, the size of each person is identified. (Step 6): From the data of the hidden information of each keypoint, obtain the data corresponding to the position of the grid identified in (Step 2). Thereby, the information that a certain keypoint of each person is not visible and the information that it is visible are identified. (Step 7): From the data of the relative positions of each keypoint, obtain only the data corresponding to the position of the grid where the keypoint is identified as visible in (Step 6) (P12 in Fig. 19). Thereby, only the relative position of each visible keypoint of each person is obtained. (Step 8): Using the center of the grid identified in (Step 2) and the data obtained in (Step 7), identify the position in the processed image of each visible keypoint (P2 in Fig. 19). Thereby, the position in the processed image of each visible keypoint of each person is identified.

[0074] (Step 9): Based on the data of the likelihood of the position of the keypoint, identify the grid (P4 in Fig. 20) where each keypoint (P5 in Fig. 20) is located (included). Specifically, identify the grid whose likelihood is greater than or equal to the threshold. (Step 10): From the data of the correction amount of the position of the keypoint, obtain the correction amount (P6 in Fig. 20) corresponding to the position of the grid identified in (Step 9). (Step 11): Based on the position of the grid identified in (Step 9) (including the center position of the grid) and the correction amount obtained in (Step 10), identify the position in the processed image of each keypoint included in the processed image (P5 in Fig. 20). (Step 12): For the positions of the keypoints in the processed images of each person obtained in (Step 8) and the positions of the keypoints in the processed image obtained in (Step 11), those with a close distance among the same type of keypoints (e.g., those with a distance less than or equal to the threshold) are associated. By integrating the associated positions, the positions of the keypoints in the processed images of each person obtained in (Step 8) are corrected, thereby calculating the positions of each of the multiple visible keypoints of each person in the processed image. Examples of the integration method include average, weighted average, selection of either one, etc.

[0075] Since the positions of each of the keypoints calculated in (Step 12) in the processed image and the positions of the grids indicating the positions of the people are associated in (Step 8), it will be known which person each of the calculated positions of the keypoints in the processed image corresponds to. Also, in (Step 7), only the data corresponding to the positions of the grids where the keypoints were identified as visible in (Step 6) was acquired, but it is also possible to acquire data including the positions of the grids identified as not visible.

[0076] Note that the estimation unit 21 may or may not estimate the positions of each of the multiple non-visible keypoints of each person in the processed image. If not estimated, since the types of non-visible keypoints are known for each person, it is also possible to output that information (the types of non-visible keypoints) for each person. Furthermore, as shown in P40 of FIG. 24, it is also possible to represent the types of non-visible keypoints for each person as an object imitating a person and display them for each person.

[0077] When making a presumption, as the presumption process, for example, the following can be considered. The presumption unit 21 identifies visible keypoints that are directly connected to invisible keypoints based on the connection relationships of a plurality of predefined keypoints for a person. Then, the presumption unit 21 presumes the position of the invisible keypoint within the processed image based on the position of the visible keypoint that is directly connected to the invisible keypoint within the processed image. The details are various and can be realized using any technology.

[0078] Also, the position of the presumed invisible keypoint within the processed image can be displayed as a range of a circle centered on that position. Since the position of the presumed invisible keypoint within the processed image is actually an approximate position, this is a display method that can represent it. The range of the circle may be calculated based on the spread of the positions of the keypoints corresponding to the person to whom the keypoint belongs, or may be fixed. Incidentally, since the position of the presumed visible keypoint within the processed image is accurate, it may be displayed with an object (point, figure, etc.) that can indicate that position at a single point.

[0079] Next, an example of the processing flow of the presumption device 20 will be described using the flowchart of FIG. 21.

[0080] In S20, the presumption device 20 acquires a processed image. For example, an operator inputs the processed image into the presumption device 20. Then, the presumption device 20 acquires the input processed image.

[0081] In S21, the presumption device 20 presumes the position of each of the plurality of keypoints of each person included in the processed image using the presumption model learned by the learning device 10. This process is realized by the presumption unit 21. The details of the process executed by the presumption unit 21 are as described above.

[0082] In S22, the presumption device 20 outputs the presumption result of S21. The presumption device 20 can utilize any means such as a display, a projection device, a printer, and e-mail.

[0083] Next, an example of the hardware configuration of the estimation device 20 will be described. Each functional unit of the estimation device 20 is realized by an arbitrary combination of hardware and software centered around a CPU, a memory, a program loaded into the memory, a storage unit such as a hard disk for storing the program (which can store not only programs stored in advance at the time of shipping the device but also programs downloaded from a storage medium such as a CD or a server on the Internet), and a network connection interface. And it is understood by those skilled in the art that there are various modifications to the realization method and device.

[0084] FIG. 16 is a block diagram illustrating the hardware configuration of the estimation device 20. As shown in FIG. 16, the estimation device 20 includes a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The estimation device 20 may not have the peripheral circuit 4A. Note that the estimation device 20 may be composed of a plurality of physically and / or logically divided devices. In this case, each of the plurality of devices can have the above hardware configuration.

[0085] According to the estimation device 20 of the present embodiment described above, the position of each of the plurality of keypoints of each person included in the processed image can be estimated using the estimation model correctly learned except for the information of the keypoints not visible in the image. According to such an estimation device 20, the accuracy of the estimation is improved.

[0086] "Modification Example" Hereinafter, several modification examples will be described. The above embodiment can also be configured to adopt one or more of the following plurality of modification examples.

[0087] -First Modification Example- Based on at least one of the number of keypoints estimated to be visible in the processed image and the number of keypoints estimated to be invisible in the processed image for each estimated person, the estimation unit 21 may calculate and output information indicating at least one of the degree to which the body of the person is visible in the processed image and the degree to which the body of the person is hidden in the processed image for each estimated person.

[0088] For example, the estimation unit 21 may calculate, as information indicating the degree to which the body of the person is visible in the processed image for each estimated person, the ratio of the number of keypoints estimated to be visible in the processed image to the total number of keypoints for each estimated person.

[0089] Alternatively, the estimation unit 21 may calculate, as information indicating the degree to which the body of the person is hidden in the processed image for each estimated person, the ratio of the number of keypoints estimated to be invisible in the processed image to the total number of keypoints for each estimated person.

[0090] The information (or ratio) indicating the degree to which the body of each person is visible / invisible calculated as described above may be displayed for each person based on the center position of each person or the position of a specified keypoint, as shown at P30 in FIG. 22. Also, based on a specified threshold value, the information (or ratio) may be converted into information indicating whether each person is visible or hidden, and the converted information may be displayed in the same manner as described above (P31 in FIG. 23). Further, a color / pattern may be assigned to the information indicating whether each person is visible or hidden, and the keypoints for each person may be displayed in that color, as shown at P32 in FIG. 23.

[0091] -Second Modification Example- The estimation model of the above-described embodiment learns and estimates whether each of a plurality of key points of each person is visible in the processed image. As a modification, the estimation model may further learn and estimate the state of how each key point that is not visible in the processed image is hidden, instead of or in addition to the hidden information described above. In this modification, in the correct label of the learning data, the state of how each key point that is not visible in the teacher image is hidden is further shown. The state of how a non-visible key point is hidden can include, for example, a state of being located outside the image, a state of being located inside the image but hidden by another object, and a state of being located inside the image but hidden by its own part.

[0092] As an example of realizing this modification, an example of adding this information to the hidden information can be considered. For example, in the above-described embodiment, in the hidden information, a value of "0" is assigned to a visible key point, and a value of "1" is assigned to a non-visible key point. In the modification, in the hidden information, for example, a value of "0" is assigned to a visible key point, a value of "1" is assigned to a non-visible key point due to being located outside the image, a value of "2" is assigned to a non-visible key point due to being located inside the image but hidden by another object, and a value of "3" is assigned to a non-visible key point due to being located inside the image but hidden by its own part. One or more of the hidden information indicates a non-visible key point.

[0093] -Third Modification- The estimation model of the above-described embodiment learns and estimates whether each of a plurality of key points of each person is visible in the processed image. As a modification, the estimation model may further learn and estimate the state of how each key point that is not visible in the processed image overlaps, as the number of objects hiding the key point, instead of or in addition to the hidden information described above. In this modification, in the correct label of the learning data, the state of how each key point that is not visible in the teacher image overlaps, as the number of objects hiding the key point, is further shown.

[0094] As an example of implementing this modification, an example of adding this information to the hidden information can be considered. For example, in the above embodiment, in the hidden information, a value of "0" is given to the visible keypoints, and a value of "1" is given to the invisible keypoints. In the modification, in the hidden information, for example, a value of "0" is given to the visible keypoints, and a value corresponding to the number M of objects hiding the keypoint, for example, a value of "M", is given to the invisible keypoints. One or more of the hidden information indicates an invisible keypoint.

[0095] Regarding the number of objects hiding each keypoint for each person shown above, calculate the maximum value for each person, and calculate the calculated maximum value as the overlapping state for each person. The calculated overlapping state (or the maximum value) for each person may be displayed for each person based on the central position of each person or the position of the specified keypoint, as shown in P35 of FIG. 25. Also, assign a color / pattern to the overlapping state for each person, and display the keypoints for each person in that color, as shown in P36 of FIG. 25.

[0096] Since the number of objects hiding each keypoint for each person, or the overlapping state (or the maximum value) for each person, shown above is known, it is also possible to construct depth information for each person or each keypoint based on that information. The depth information shown here indicates the order of the distances from the camera.

[0097] Note that the third modification can also be combined with the second modification.

[0098] Although the embodiments of the present invention have been described above with reference to the drawings, these are examples of the present invention, and various configurations other than the above can also be adopted.

[0099] In addition, in this specification, "acquisition" means, based on user input or based on program instructions, "the self-device goes to obtain data stored in other devices or storage media (active acquisition)", for example, requesting or inquiring of other devices and receiving, accessing other devices or storage media and reading out, etc., and, based on user input or based on program instructions, "inputting data output from other devices to the self-device (passive acquisition)", for example, receiving data distributed (or transmitted, push-notified, etc.), and also selecting and acquiring from the received data or information, and "generating new data by editing (textualizing, rearranging data, extracting some data, changing file format, etc.) the data and acquiring the new data", including at least any one of them.

[0100] Some or all of the above embodiments may be described as follows in the appended claims, but are not limited thereto. 1. Acquisition means for acquiring learning data in which a teacher image including a person, a correct label indicating the position of each person, a correct label indicating whether each of a plurality of key points of each person's body is visible in the teacher image, and a correct label indicating the position within the teacher image of the key point that is visible in the teacher image among the plurality of key points are associated; Learning means for learning an estimation model that estimates information indicating the position of each person, information indicating whether each of a plurality of the key points of each person included in the processed image is visible in the processed image, and information related to the position of each key point for calculating the position within the processed image of the key point that is visible in the processed image, based on the learning data; A learning device having the above. 2. The learning device according to 1, wherein in the correct label, the position within the teacher image of the key point that is not visible in the teacher image is not shown. 3. The learning means is Based on the estimation model being learned, estimate information indicating the position of each person, information indicating whether each of the plurality of keypoints of each person included in the processed image is visible in the processed image, and information regarding the position of each keypoint for calculating the position of each of the plurality of keypoints in the teacher image, Adjust the parameters of the estimation model so as to minimize the difference between the estimation result of the information indicating the position of each person and the information indicating the position of each person indicated by the correct label, Adjust the parameters of the estimation model so as to minimize the difference between the estimation result of the information indicating whether each of the plurality of keypoints of each person included in the processed image is visible in the processed image and the information indicating whether each of the plurality of keypoints of the body of each person indicated by the correct label is visible in the teacher image, The learning device according to 1 or 2, wherein the parameters of the estimation model are adjusted so as to minimize the difference between the estimation result of the information regarding the position of each keypoint for calculating the position of each of the plurality of keypoints in the teacher image and the information regarding the position of each keypoint obtained from the position of the keypoint visible in the teacher image among the plurality of keypoints indicated by the correct label, only for the keypoints visible in the teacher image indicated by the correct label. 4. The correct label further indicates the state of each keypoint that is not visible for each person in the teacher image, The learning device according to any one of 1 to 3, wherein the estimation model further estimates the state of each keypoint that is not visible for each person in the processed image. 5. The learning device according to 4, wherein the state includes a state of being located outside the image, a state of being located inside the image but hidden by another object, and a state of being located inside the image but hidden by one's own part. 6. The learning device according to 4, wherein the state indicates the number of objects hiding the keypoint that is not visible in the teacher image or the processed image. 7. An estimation device having an estimation means for estimating the position of each of a plurality of keypoints of each person included in a processed image using an estimation model learned by the learning device according to any one of 1 to 6. 8. The estimation means estimates, using the estimation model, whether or not each of the plurality of keypoints of each person included in the processed image is visible in the processed image, and estimates the position of each of the plurality of keypoints of each person included in the processed image using the result of the estimation. The estimation device according to 7. 9. The estimation means outputs, for each person, the type of keypoint that is not visible, or represents the type of the non-visible keypoint as an object imitating a person and displays it for each person, using the estimated information as to whether or not each of the plurality of keypoints of each person included in the processed image is visible in the processed image. The estimation device according to 8. 10. The estimation means identifies non-visible keypoints using the estimated information as to whether or not each of the plurality of keypoints of each person included in the processed image is visible in the processed image, and based on a connection relationship of a plurality of keypoints for a person defined in advance, identifies visible keypoints directly connected to the identified non-visible keypoints, and estimates the position of the identified non-visible keypoints in the processed image based on the position of the identified visible keypoints in the processed image. The estimation device according to 8 or 9. 11. The estimation means calculates information indicating at least one of the degree to which a person's body is visible and the degree to which a person's body is hidden in the processed image, for each estimated person, based on at least one of the number of keypoints estimated to be visible and the number of keypoints estimated to be non-visible in the processed image for each estimated person. The estimation device according to any one of 7 to 10. 12. The estimation means according to claim 11, wherein information indicating at least one of the degree to which the calculated human body is visible and the degree to which the human body is hidden is displayed for each person based on the central position of each person or the designated key point position. 13. The estimation means according to claim 11, wherein information indicating at least one of the degree to which the calculated human body is visible and the degree to which the human body is hidden is converted into information indicating visible / hidden for each person based on a designated threshold value, and the converted information is displayed for each person based on the central position of each person or the designated key point position. 14. The estimation means according to claim 7, wherein for the number of the objects hiding each key point for each person, a maximum value is calculated for each person, the calculated maximum value is calculated as the overlapping state for each person, and the calculated overlapping state for each person is displayed for each person based on the central position of each person or the position of the designated key point, or a color / pattern is assigned to the overlapping state for each person and the key points in units of persons are displayed in the assigned colors. 15. A computer performs an acquisition step of acquiring training data in which a teacher image including a person is associated with a correct label indicating the position of each person, a correct label indicating whether each of a plurality of key points of each person's body is visible in the teacher image, and a correct label indicating the position within the teacher image of the key point that is visible in the teacher image among the plurality of key points; a learning step of learning an estimation model for estimating information indicating the position of each person, information indicating whether each of the plurality of key points of each person included in the processed image is visible in the processed image, and information related to the position of each key point for calculating the position within the processed image of the key point that is visible in the processed image, based on the training data; A learning method for performing the above. 16. A computer An acquisition means for acquiring learning data in which a teacher image including a person, a correct label indicating the position of each person, a correct label indicating whether or not each of a plurality of key points of the body of each person is visible in the teacher image, and a correct label indicating the position in the teacher image of the key points that are visible in the teacher image among the plurality of key points are associated with each other. A learning means for learning an estimation model that estimates information indicating the position of each person, information indicating whether or not each of a plurality of the key points of each person included in the processing image is visible in the processing image, and information related to the position of each key point for calculating the position in the processing image of the key points that are visible in the processing image, based on the learning data. A program that functions as. 17. A computer An estimation method in which a computer executes an estimation step of estimating the position in the processing image of each of a plurality of key points of each person included in the processing image, using the estimation model learned by the learning device according to any one of 1 to 6. 18. A computer A program that causes a computer to function as an estimation means for estimating the position in the processing image of each of a plurality of key points of each person included in the processing image, using the estimation model learned by the learning device according to any one of 1 to 6.

Explanation of Signs

[0101] 10 Learning device 11 Acquisition unit 12 Learning unit 13 Storage unit 20 Estimation device 21 Estimation unit 22 Storage unit 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus

Claims

1. A receiving means for receiving designation of positions of key points at which a person's body is visible in a teacher image; an acquisition means for acquiring learning data that associates the teacher image, the position of each person, a label indicating whether each of a plurality of key points on the body of each person is visible in the teacher image, and the position of the accepted key point; having An information processing device, wherein the learning data further includes a state of each of the key points that are not visible for each person in the teacher image.

2. The information processing apparatus according to claim 1 , wherein the training data does not indicate positions in the teacher image of key points that are not visible in the teacher image.

3. The information processing apparatus according to claim 1 , wherein the states include a state in which the object is located outside the image, a state in which the object is located within the image but hidden by another object, and a state in which the object is located within the image but hidden by a part of the object itself.

4. The information processing apparatus according to claim 1 , wherein the state indicates a number of objects that are hiding the key points that are not visible in the teacher image or the processed image.

5. an estimation means for estimating, based on an estimation model that has learned the learning data, information indicating the position of each person, information indicating whether each of a plurality of key points of each person included in a processed image is visible in the processed image, and information related to the position of each key point for calculating the position in the processed image of the key point that is visible in the processed image; 5. The information processing device according to claim 1, further comprising:

6. The information processing device according to any one of claims 1 to 4, further comprising an estimation means for estimating a position within the processed image of each of a plurality of key points of each person contained in the processed image based on an estimation model trained with the learning data.

7. The computer a receiving step of receiving designation of positions of key points at which a person's body is visible in a teacher image; an acquisition step of acquiring learning data in which the teacher image, the position of each person, a label indicating whether each of a plurality of key points on the body of each person is visible in the teacher image, and the position of the accepted key point are associated with each other; Run An information processing method in which the learning data further includes a state of each of the key points that are not visible for each person in the teacher image.

8. The computer, The information processing method according to claim 7 , further comprising an estimation step of estimating a position within the processed image of each of a plurality of key points of each person included in the processed image based on an estimation model trained with the training data.

9. Computer, A receiving means for receiving designation of positions of key points at which a person's body is visible in a teacher image; an acquisition means for acquiring learning data in which the teacher image, the position of each person, a label indicating whether each of a plurality of key points on the body of each person is visible in the teacher image, and the position of the accepted key point are associated with each other; Function as a The program, wherein the learning data further includes a state of each of the key points that are not visible for each person in the training image.

10. The computer, 10. The program according to claim 9, further functioning as an estimation means for estimating a position within the processed image of each of a plurality of key points of each person included in the processed image, based on an estimation model trained with the learning data.

Citation Information

Patent Citations

  • Document management device and document management program

    JP2004295436A

  • Image retrieving apparatus, image retrieving method, and setting screen used therefor

    JP2019091138A

  • Learning device, learning method, learning program, and object recognition device

    JP2020123105A

  • Method, device, medium and apparatus for determining a bounding box of an object

    JP2020525959A

  • Image processing device that recognizes state of subject and method for same

    WO2020217812A1