Image processing device, image processing method, and program

JP7899954B2Active Publication Date: 2026-08-04NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2022-12-05
Publication Date
2026-08-04

AI Technical Summary

Benefits of technology

【0009】 本開示によれば、ーポイント対応付けの新規な技術が提供される。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007899954000005
    Figure 0007899954000005
  • Figure 0007899954000006
    Figure 0007899954000006
  • Figure 0007899954000007
    Figure 0007899954000007
Patent Text Reader

Abstract

The keypoint association device (2000) acquires a target image (10) in which one or more people (80) are captured, detects keypoints (20) from the target image (10), and generates a spatial feature map (30) for each pair of body parts. The spatial feature map (30) includes a first direction region (32) for each keypoint (20) representing a first body part of the pair and a second direction region (32) for each keypoint (20) representing a second body part of the pair. The first direction region (32) and the second direction region (32) belonging to the same person (80) represent directions from the keypoint (20) corresponding to the first direction region (32) to the keypoint (20) corresponding to the second direction region (32). The keypoint association device (2000) generates a keypoint group (40) in which the keypoints (20) belong to the same person (80) for each person (80) captured in the target image (10).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , , , , , , , ,

[0003] ,

[0001] The present disclosure relates generally to a keypoint association device, a keypoint association method, and a non - transient computer - readable storage medium.

Background Art

[0002] There are various types of analyses performed on images in which one or more persons are imaged. Some of these analyses, such as pose estimation, use the keypoints of a person. Specifically, keypoints are detected from an image, and the keypoints are divided into groups such that each group contains keypoints belonging to the same person. This process of dividing the keypoints into groups is called "keypoint association".

[0003] Non - Patent Document 1 discloses one algorithm for keypoint association. The system of Non - Patent Document 1 generates a feature map including a region called a PAF (Part Affinity Field) corresponding to each pair of predetermined body parts for each person from an input image. The PAF corresponding to the pair of body parts connects two keypoints representing the pair of body parts, and the two keypoints belong to the same person. The PAF is filled with pixel values representing the direction between the two keypoints.

Prior Art Documents

Non - Patent Documents

[0004]

Non - Patent Document 1

[0005] Since a PAF needs to connect two corresponding keypoints, it may include regions that are far from either of those two keypoints, for example, regions near the midpoint between them. The object of this disclosure is to provide a novel technique for keypoint mapping. [Means for solving the problem]

[0006] The keypoint mapping device provided in this disclosure comprises at least one memory configured to store instructions and at least one processor configured to execute the instructions. The at least one processor, by executing the instruction, performs the following actions: acquire a target image in which one or more people are captured; detect keypoints from the target image for each of the body parts of the people; generate a spatial feature map using the target image for each predetermined pair of the body parts; the spatial feature map relating to the pair of body parts includes a first directional region corresponding to each keypoint representing the first body part of the pair and a second directional region corresponding to each keypoint representing the second body part of the pair, wherein the first and second directional regions belonging to the same person represent the direction from the keypoint in the first directional region to the keypoint in the second directional region; and generate a keypoint group for each of the people captured in the target image, including the keypoints belonging to the same person.

[0007] The keypoint mapping method provided in this disclosure is performed by computer. The keypoint mapping method includes acquiring a target image in which one or more people are captured; detecting keypoints from the target image for each of the body parts of the people; generating a spatial feature map using the target image for each predetermined pair of the body parts; the spatial feature map relating to the pair of body parts includes a first directional region corresponding to each keypoint representing the first body part of the pair, and a second directional region corresponding to each keypoint representing the second body part of the pair, wherein the first and second directional regions belonging to the same person represent the direction from the keypoint in the first directional region to the keypoint in the second directional region; and generating a keypoint group for each of the people captured in the target image, which includes keypoints belonging to the same person.

[0008] The non-temporary computer-readable storage medium provided in this disclosure stores programs. The program causes the computer to perform the following actions: acquire a target image in which one or more people are captured; detect keypoints from the target image for each of the body parts of the people; generate a spatial feature map using the target image for each predetermined pair of the body parts; the spatial feature map for the pair of body parts includes a first directional region corresponding to each keypoint representing the first body part of the pair and a second directional region corresponding to each keypoint representing the second body part of the pair, wherein the first and second directional regions belonging to the same person represent the direction from the keypoint in the first directional region to the keypoint in the second directional region; and generate a keypoint group for each of the people captured in the target image, which includes the keypoints belonging to the same person. [Effects of the Invention]

[0009] According to this disclosure, a novel technology for one-point mapping is provided. [Brief explanation of the drawing]

[0010] [Figure 1] Figure 1 shows an overview of the keypoint mapping device. [Figure 2] Figure 2 shows an example of a spatial feature map. [Figure 3] Figure 3 is a block diagram showing an example of the functional configuration of a keypoint mapping device. [Figure 4] Figure 4 is a block diagram showing an example of the hardware configuration of a keypoint mapping device. [Figure 5] Figure 5 is a flowchart showing an example of the process performed by a keypoint mapping device. [Figure 6] Figure 6 shows an example of the configuration of the feature generation unit. [Figure 7] Figure 7 shows an example of a pair of horizontal and vertical space feature maps, where the directions between keypoints are represented in three-dimensional space. [Figure 8]Figure 8 shows an example of the configuration of the feature generation unit when the keypoint positions are represented by three-dimensional coordinates. [Figure 9] Figure 9 shows an example of a keypoint mapping method. [Modes for carrying out the invention]

[0011] Examples of embodiments relating to this disclosure will be described below with reference to the drawings. Throughout the drawings, the same elements are denoted by the same reference numerals, and redundant descriptions are omitted where necessary. Unless otherwise stated, predetermined information (for example, predetermined values ​​or predetermined thresholds) is pre-stored in a storage device accessible to the computer using the information.

[0012] <Overview> Figure 1 is a diagram showing an overview of a keypoint mapping device 2000 according to one embodiment. Please note that the overview shown in Figure 1 is an example of the operation of the keypoint mapping device 2000 to facilitate understanding of the device, and does not limit or narrow the range of possible operations of the keypoint mapping device 2000.

[0013] The keypoint mapping device 2000 acquires a target image 10 in which one or more people are captured, detects keypoints 20 from the target image 10, and performs keypoint mapping on the detected keypoints 20. The target image 10 may be any type of image data, such as an RGB image or a grayscale image in which people can be visibly captured.

[0014] Keypoint 20 can indicate the position of a body part of a person imaged in the target image 10. The position of the body part may be represented by two-dimensional (2D) coordinates on the image plane of the target image 10 or three-dimensional (3D) coordinates in a predetermined three-dimensional (3D) space. The keypoint association device 2000 is configured to detect one or more keypoints 20 from the target image 10 for each of the predetermined body parts. The predetermined body parts may include the neck, right and left eyes, right and left ears, right and left shoulders, right and left elbows, right and left wrists, waist, right and left knees, and right and left feet.

[0015] Keypoint association is a process of generating a group called "keypoint group 40" for each person included in the target image 10. The keypoint group 40 of a specific person includes only the keypoints 20 belonging to the specific person.

[0016] To generate a keypoint group 40 for each person, the keypoint association device 2000 generates a spatial feature map 30 for each pair of predetermined body parts based on the target image 10. The pairs of predetermined body parts may include pairs of adjacent body parts such as the pair of the right eye and the neck, the pair of the neck and the right shoulder, the pair of the right shoulder and the right elbow, and the pair of the right elbow and the right wrist. However, the body parts of a specific pair do not necessarily have to be adjacent to each other.

[0017] The spatial feature map 30 related to a specific pair of body parts may be image data of the same size as the target image 10 and may include a region called a "direction region" for each keypoint 20 indicating one of the body parts of the specific pair. The direction regions belonging to the same person indicate the direction between the keypoints 20 (the direction from one of the keypoints 20 to the other). In one embodiment, different colors (in other words, pixel values) may be assigned to different directions. In this case, the direction region is filled with the color corresponding to the direction that the direction region should represent. The region not included in any direction region may be filled with a color not assigned to any direction.

[0018] FIG. 2 is a diagram showing an example of the spatial feature map 30. The target image 10 shown in FIG. 2 includes two persons 80. The spatial feature map 30 shown in FIG. 2 is generated for the pair of the left elbow and the left wrist. Therefore, the spatial feature map 30 includes four direction regions 32-1 to 32-4 representing the left elbow of the person 80-1, the left wrist of the person 80-1, the left elbow of the person 80-2, and the left wrist of the person 80-2, respectively.

[0019] In FIG. 2, the direction region 32 represents the direction from the left elbow to the left wrist of the corresponding person 80. For example, the direction regions 32-1 and 32-2 corresponding to the person 80-1 represent the direction from the left elbow to the left wrist of the person 80-1. Since the left elbow and the left wrist of the person 80-1 are represented by the key points 20-1 and 20-2, respectively, the direction regions 32-1 and 32-2 represent the direction from the key point 20-1 to the key point 20-2.

[0020] After generating the spatial feature map 30, the key point association device 2000 divides the key points 20 into key point groups 40 based on the spatial feature map 30. A specific method for generating the key point group 40 will be described later.

[0021] <Example of the operation and effect> According to the key point association device 2000, the key points 20 detected from the target image 10 are classified into key point groups 40 such that each key point group 40 includes only the key points 20 belonging to the same person as each other. For this purpose, the key point association device 2000 generates the spatial feature map 30 for each pair of predetermined body parts. Thus, the key point association device 2000 provides a novel technique for key point association.

[0022] Furthermore, the keypoint mapping device 2000 has advantages in the following respects. As mentioned above, Non-Patent Literature 1 generates a feature map for each person, including a PAF that connects the two keypoints corresponding to each pair of body parts. This feature map is generated using a Convolutional Neural Network (CNN). Since the PAF may include regions that are far from either of the corresponding keypoints (for example, intermediate regions between the keypoints), a problem may arise in the training of the CNN where the convergence of such regions in the PAF is slow.

[0023] In this regard, the spatial feature map 30 relating to a pair of body parts includes a separate directional region 32 for each of the two keypoints of the pair. Therefore, regions far from the keypoints, such as the intermediate region between the keypoints, are not included in the directional region 32. Consequently, when the spatial feature map 30 is generated by a machine learning-based model, it is possible to prevent adverse effects on model training caused by delays in convergence in regions far from the keypoints.

[0024] The following provides a detailed explanation of the keypoint mapping device 2000.

[0025] <Example of functional configuration> Figure 3 is a block diagram showing an example of the functional configuration of a keypoint mapping device 2000 according to one embodiment. The keypoint mapping device 2000 includes an acquisition unit 2020, a keypoint detection unit 2040, a feature map generation unit 2060, and a keypoint mapping unit 2080. The acquisition unit 2020 acquires a target image 10. The keypoint detection unit 2040 detects keypoints 20 from the target image 10. The feature map generation unit 2060 generates a spatial feature map 30 for each of a predetermined pair of body parts using the target image 10. The keypoint mapping unit 2080 generates a keypoint group 40 based on the spatial feature map 30.

[0026] <Example of hardware configuration> The keypoint mapping device 2000 can be implemented using one or more computers. Each of these computers may be a dedicated computer manufactured specifically for implementing the keypoint mapping device 2000, or it may be a general-purpose computer such as a personal computer (PC), server, or mobile device.

[0027] The keypoint mapping device 2000 can be implemented by installing an application on a computer. This application is implemented by a program that makes the computer function as the keypoint mapping device 2000. In other words, this program implements the functional components of the keypoint mapping device 2000.

[0028] Figure 4 is a block diagram showing an example of the hardware configuration of a computer 1000 that implements a keypoint mapping device 2000 according to one embodiment. In Figure 4, the computer 1000 includes a bus 1020, a processor 1040, memory 1060, storage device 1080, input / output (I / O) interface 1100, and network interface 1120.

[0029] Bus 1020 is a data transmission path for sending and receiving data between the processor 1040, memory 1060, storage device 1080, input / output interface 1100, and network interface 1120. The processor 1040 is a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), DSP (Digital Signal Processor), or FPGA (Field-Programmable Gate Array). Memory 1060 is main memory such as RAM (Random Access Memory) or ROM (Read Only Memory). Storage device 1080 is auxiliary storage such as a hard disk, SSD (Solid State Drive), or memory card. The input / output interface 1100 is an interface between the computer 1000 and peripheral devices such as a keyboard, mouse, and display device. The network interface 1120 is an interface between the computer 1000 and a network. This network may be a LAN (Local Area Network) or a WAN (Wide Area Network).

[0030] The processor 1040 is configured to load the aforementioned program instructions from the storage device 1080 into memory 1060 and execute those instructions, thereby causing the computer 1000 to operate as a keypoint mapping device 2000.

[0031] The hardware configuration of computer 1000 is not limited to that shown in Figure 4. For example, as described above, the keypoint mapping device 2000 may be implemented as a combination of multiple computers. In this case, these computers may be connected to each other via a network.

[0032] <Processing flow> Figure 5 is a flowchart illustrating an example of the process performed by a keypoint mapping device 2000 according to one embodiment. The acquisition unit 2020 acquires the target image 10 (S102). The keypoint detection unit 2040 detects keypoints from the target image 10 (S104). The feature map generation unit 2060 generates a spatial feature map 30 for each of a predetermined pair of body parts (S106). The keypoint mapping unit 2080 generates a keypoint group 40 for each person (S108).

[0033] <Acquisition of target image 10: S102> The acquisition unit 2020 acquires the target image 10 (S102). There are various ways to acquire the target image 10. In one embodiment, the target image 10 is pre-stored in a storage device in a manner that allows it to be acquired from the keypoint mapping device 2000. In this case, the acquisition unit 2020 can access the storage device to acquire the target image. In another embodiment, the target image 10 may be transmitted from another computer, such as a camera, that generates the target image 10. In this case, the acquisition unit 2020 can acquire the target image 10 by receiving the target image 10.

[0034] In one embodiment, the target image 10 may be one of a time-series image, such as a time-series video frame that constitutes a video. In this case, the keypoint mapping device 2000 can acquire all or part of the time-series image as the target image 10 and perform keypoint detection and keypoint mapping for each target image 10.

[0035] <Keypoint detection: S104> The keypoint detection unit 2040 detects keypoints 20 from the target image 10 (S104). There are various methods for detecting the location of one or more predetermined parts of a person's body as keypoints from an image, and the keypoint detection unit 2040 can detect keypoints 20 from the target image 10 using one of these methods.

[0036] In one embodiment, the keypoint detection unit 2040 may include a machine learning-based model (e.g., a neural network) that receives an image as input and, in response to the input image, is pre-trained to detect one or more keypoints 20 for each predetermined part of the input image. Hereinafter, this model will be referred to as the "keypoint detection model".

[0037] A keypoint detection model may receive a target image 10 as input, extract features from the target image 10, detect one or more locations of predetermined body parts based on the extracted features, and output pairs of locations and labels as keypoints. The label of a keypoint indicates which body part the keypoint represents. In this case, the keypoint detection model may include a first model pre-trained to extract features from the target image 10, and a second model pre-trained to detect one or more locations of predetermined body parts based on the features extracted by the first model. Each of the first and second models may be configured as a machine learning-based model such as a neural network. There are various types of machine learning models that can detect keypoints from an input image, and the keypoint detection model may be configured as one such model.

[0038] <Generating spatial feature maps: S106> The feature map generation unit 2060 generates a spatial feature map 30 for each of a predetermined pair of body parts (S106). To generate the spatial feature map 30, the feature map generation unit 2060 may include a machine learning-based model called a "feature map generation model" for each of the predetermined pairs of body parts. Figure 6 shows an example of the configuration of the feature map generation unit 2060. In Figure 6, it is assumed that there are N predetermined pairs of body parts. Therefore, the feature map generation unit 2060 includes a feature map generation model 70 for each of the predetermined N pairs of body parts.

[0039] The feature map generation model 70 for a specific pair of body parts is configured to receive image data and information on keypoints 20 detected from the image data that represent one of the body parts in the pair as input. The feature map generation model 70 is pre-trained to generate a spatial feature map 30 for the corresponding pair of body parts in response to the input data.

[0040] When the positions of keypoints 20 are represented in two-dimensional coordinates, as shown in Figure 6, the direction between two keypoints 20 can be represented by a single angle, for example, the angle between the X-axis and the line connecting the two keypoints 20. Therefore, the feature map generation unit 2060 can generate one spatial feature map 30 for each of a given pair of body parts. On the other hand, when the positions of keypoints 20 are represented in three-dimensional coordinates, the direction between two keypoints 20 can be represented by a pair of angles. Therefore, the feature map generation unit 2060 can generate two spatial feature maps 30 for each of a given pair of body parts. The case where the positions of keypoints 20 are represented in three-dimensional coordinates will be explained in more detail below.

[0041] If the positions of keypoints 20 are represented in three-dimensional coordinates, the direction between two keypoints 20 can be represented as a pair of horizontal and vertical directions. To represent the direction between keypoints 20 as a pair of horizontal and vertical directions in three-dimensional space, the feature map generation unit 2060 can generate a pair of spatial feature maps 30 representing the horizontal direction between keypoints 20 and a spatial feature map 30 representing the vertical direction between keypoints 20. Hereafter, the spatial feature map 30 representing the horizontal direction between the 20 key points will be referred to as the "horizontal spatial feature map," and the spatial feature map 30 representing the vertical direction between the 20 key points will be referred to as the "vertical spatial feature map."

[0042] Figure 7 shows an example of representing the direction between keypoints 20 in three-dimensional space using a pair of horizontal and vertical spatial feature maps. In Figure 7, it is assumed that the spatial feature map 30 was generated for the pair of the left elbow and left wrist. It is also assumed that keypoint 20-1 and keypoint 20-2 represent the positions of the person's left elbow and left wrist, respectively.

[0043] The positions of keypoint 20-1 and keypoint 20-2 in three-dimensional space are represented by points Q1 and Q2, respectively. Therefore, the direction from keypoint 20-1 to keypoint 20-2 in three-dimensional space is represented by a vector V whose starting and ending points are Q1 and Q2, respectively.

[0044] The horizontal direction of vector V can be represented by the angle between the vector and the X-axis when vector V is projected onto the XY plane. This angle is shown as θ in Figure 7. Therefore, the horizontal spatial feature map 50 is generated to include directional regions 32-1 and 32-2, where the angle θ is represented by pixel values.

[0045] The vertical direction of vector V can be represented by the angle between the XY plane and vector V. This angle is shown as φ in Figure 7. Therefore, the vertical spatial feature map 60 is generated to include directional regions 32-3 and 32-4, where the angle φ is represented by pixel values.

[0046] When the position of keypoint 20 is represented by three-dimensional coordinates, for each of a predetermined pair of body parts, the feature map generation model 70 may include a first model that generates a horizontal spatial feature map 50 for the pair of body parts, and a second model that generates a vertical spatial feature map 60 for the pair of body parts. By using these feature map generation models, the feature map generation unit 2060 can generate pairs of horizontal spatial feature maps 50 and vertical spatial feature maps 60 for each of a predetermined pair of body parts, based on the target image 10 and the keypoints 20 detected from the target image 10.

[0047] Figure 8 shows an example of the configuration of the feature map generation unit 2060 when the location of the key point 20 is represented by three-dimensional coordinates. Each feature map generation model 70 includes a pair consisting of a first model 72 that generates a horizontal spatial feature map 50 and a second model 74 that generates a vertical spatial feature map 60.

[0048] <Keypoint mapping: S108> The keypoint mapping unit 2080 performs keypoint mapping by generating keypoint groups 40 based on the spatial feature map 30 (S108). As described above, the keypoint groups 40 are generated to contain only keypoints 20 belonging to the same person. Assume that the number of people captured in the target image 10 is N. In this case, the keypoint mapping unit 2080 can generate a keypoint group 40 for each of the N people. Therefore, N keypoint groups 40 can be generated.

[0049] The following describes a specific method for keypoint mapping using the spatial feature map 30. For simplicity, we first assume that the location of keypoint 20 is represented in two-dimensional coordinates. The method for keypoint mapping when the location of keypoint 20 is represented in three-dimensional coordinates will be described later.

[0050] For each of a predetermined pair of body parts, the keypoint mapping unit 2080 divides the keypoints 20 into keypoint groups 40 using the spatial feature map 30 related to that pair. For example, the keypoint mapping unit 2080 uses the spatial feature map 30 related to the pair of left elbow and left wrist to generate a keypoint group 40 that includes a pair of keypoints 20 for the left elbow and keypoints 20 for the left wrist, each belonging to the same person.

[0051] Theoretically, a pair of directional regions 32 in the spatial feature map 30 corresponds to a pair of keypoints 20 belonging to the same person if the two directional regions 32 point in the same direction to each other. Therefore, the keypoint matching unit 2080 can identify a pair of keypoints 20 belonging to the same person by identifying a pair of keypoints 20 whose directional regions 32 point in the same direction to each other.

[0052] However, in reality, there may be some differences in the indicated directions between direction regions 32 that belong to the same person. Therefore, in one embodiment, the keypoint matching unit 2080 identifies pairs of keypoints 20 whose direction regions 32 indicate directions that are substantially close to each other, and generates a keypoint group 40 that includes the identified pair of keypoints 20.

[0053] Figure 9 shows an example of a keypoint mapping method. In this example, a spatial feature map relating to the left elbow and left wrist pair is used. Therefore, for each person captured in the target image 10, a keypoint group 40 is generated that includes a pair of keypoints 20 representing the left elbow and 20 representing the left wrist.

[0054] By referring to the detection results of keypoints 20 performed by the keypoint detection unit 2040, the keypoint mapping unit 2080 identifies keypoints 20 of the left elbow (keypoints 20-2 and 20-3) and keypoints 20 of the left wrist (keypoints 20-1 and 20-4) on the spatial feature map 30. Next, the keypoint mapping unit 2080 identifies the directional regions 32 corresponding to each identified keypoint 20. Specifically, there are four directional regions 32-1 to 32-4, corresponding to keypoints 20-1 to 20-4, respectively.

[0055] As will be described later, the feature map generation model 70 may be trained to generate a spatial feature map 30 in which the directional region 32 has a predetermined shape and size, and the position of the directional region 32 is defined based on the position of the corresponding key point 20. Therefore, the key point mapping unit 2080 can identify the directional region 32 based on the predetermined shape and size and the position of the corresponding key point 20.

[0056] In the example shown in Figure 9, it is assumed that the shape of the direction region 32 is defined as a circle, and the size of the direction region 32 is defined by its radius R. It is also assumed that the center of the direction region 32 is located at the corresponding key point 20. Therefore, for each key point 20, the key point mapping unit 2080 identifies the region that is circular in shape, has a radius R, and whose center is located at the key point 20 as the direction region 32 corresponding to that key point 20.

[0057] If two or more directional regions 32 overlap, the keypoint mapping unit 2080 may adjust the size of the directional regions 32 so that they do not overlap. There are various ways to adjust the size of the directional regions. For example, the keypoint mapping unit 2080 may reduce the size of the directional regions 32 by repeatedly multiplying the size of the directional regions 32 by an adjustment coefficient that is a real number greater than 0 and less than 1, until they no longer overlap. As another example, multiple options for the size of the directional regions 32 may be defined in advance. In this case, the keypoint mapping unit 2080 may select the option with the largest size that does not overlap.

[0058] Furthermore, as will be described later, the adjustment of the size of the directional region 32 is a characteristic. map This can also be performed when generating the training dataset used to train the generative model. Therefore, it is desirable for the keypoint mapping unit 2080 to adjust the size of the direction region 32 in the same way as the method used to adjust the size of the direction region 32 when generating the training dataset.

[0059] After identifying the directional region 32 for each keypoint 20, the keypoint mapping unit 2080 identifies pairs of keypoints 20 to generate a keypoint group 40. To facilitate explanation of the operation of the keypoint mapping unit 2080, pairs of body parts corresponding to the spatial feature map 30 are referred to as the first body part and the second body part, respectively. For example, in the example shown in Figure 9, the left elbow is referred to as the first body part and the left wrist is referred to as the second body part.

[0060] The keypoint matching unit 2080 selects one of the keypoints 20 of the first body part. Next, the keypoint matching unit 2080 evaluates the keypoints 20 of the second body part in relation to the selected keypoint 20 of the first body part, and identifies which of the keypoints 20 of the second body part should be paired with the keypoint 20 of the first body part.

[0061] For example, in the example shown in Figure 9, the keypoint matching unit 2080 may select keypoint 20-2 as one of the keypoints 20 of the left elbow. Next, the keypoint matching unit 2080 evaluates each of the keypoints 20 of the left wrist (i.e., keypoints 20-1 and 20-4) and determines which of them should be paired with keypoint 20-2.

[0062] Keypoint 20 can be evaluated using an index value called "coefficient distance." The coefficient distance between two keypoints 20 represents how different the directions indicated by their respective directional regions are. For example, the coefficient distance between keypoint 20-2 and keypoint 20-1 represents the degree of difference between the direction indicated by directional region 32-2 and the direction indicated by directional region 32-1.

[0063] After selecting one of the keypoints 20 of the first body part, the keypoint matching unit 2080 calculates a coefficient distance between each keypoint 20 of the second body part and the selected keypoint 20 of the first body part. Subsequently, the keypoint matching unit 2080 generates pairs of the selected keypoint 20 of the first body part and the keypoint 20 of the second body part that have the smallest coefficient distance.

[0064] In one embodiment, a threshold for the coefficient distance may be predetermined. In this case, only when the coefficient distance is less than that threshold, the key point 20 of the second body part having the smallest coefficient distance is paired with the selected key point 20 of the first body part.

[0065] To calculate the coefficient distance between the keypoints 20, the keypoint matching unit 2080 identifies a value representing the direction (hereinafter referred to as the "direction value") for each of their direction regions 32. As described above, a direction region can represent the direction by the pixel values ​​within it. Therefore, the keypoint matching unit 2080 can calculate the statistical value of the pixel values ​​within a direction region 32 as the direction value of that direction region 32.

[0066] The coefficient distance between the 20 key points can be expressed as the absolute difference between the directional values ​​of the corresponding directional regions 32. This can be expressed by the following formula:

number

[0067] In one embodiment, the coefficient distance between key points 20 may be calculated by considering the Euclidean distance between those key points 20. This is because the longer the Euclidean distance between key points 20, the less likely it is that those key points 20 belong to the same person. When considering the Euclidean distance between key points 20, the coefficient distance between key points 20 can be expressed by the following formula.

number

[0068] After generating keypoint groups 40 for each of a predetermined pair of body parts, the keypoint mapping unit 2080 may combine keypoint groups 40 that correspond to the same person. Specifically, the keypoint mapping unit 2080 may repeatedly detect two keypoint groups 40 that each contain at least one identical keypoint 20, and combine the two detected keypoint groups 40 into a single keypoint group 40, until no two keypoint groups 40 contain the same keypoint 20 as any other keypoint group 40.

[0069] <<Regarding cases where the keypoint position is represented by 3D coordinates>> When the keypoint locations are represented by 3D coordinates, two types of spatial feature maps 30, namely a horizontal spatial feature map 50 and a vertical spatial feature map 60, are generated for each predetermined pair of body parts. Therefore, for each predetermined pair of body parts, the keypoint mapping unit 2080 generates a keypoint group 40 using the horizontal spatial feature map 50 and the vertical spatial feature map 60 related to that pair of body parts.

[0070] When the positions of keypoints 20 are represented in 3D coordinates, keypoint mapping differs from mapping when the positions of keypoints 20 are represented in 2D coordinates in that the coefficient distance is calculated based on the horizontal and vertical directions between the keypoints 20. For this reason, the keypoint mapping unit 2080 calculates the direction value of the direction region 32 in the horizontal spatial feature map 50 and the direction value of the direction region 32 in the vertical spatial feature map 60 for each keypoint 20. The coefficient distance between keypoints 20 whose positions are represented in 3D coordinates can be calculated as follows.

number

[0071] Furthermore, when calculating the coefficient distance between the 20 key points, if the Euclidean distance between those 20 key points is taken into consideration, the coefficient distance between the 20 key points can be expressed by the following formula.

number

[0072] <Output from keypoint mapping device 2000> The keypoint mapping device 2000 may be configured to output information (referred to as output information) indicating the result of keypoint mapping. For example, the output information may include an identifier of the target image 10 (e.g., frame number) and keypoint group information. For each keypoint group 40, the keypoint group information includes the identifier of the keypoint group 40 and information on each keypoint 20 included in the keypoint group 40. The information on the keypoint 20 may include the identifier of the keypoint 20, the position indicated by the keypoint 20, and the identifier of the body part indicated by the keypoint 20.

[0073] There are various methods for outputting output information. In one embodiment, the output information may be stored in a storage device, displayed on a display device, or transmitted to another computer such as a PC or smartphone of the user of the keypoint mapping device 2000.

[0074] <About training the feature map generation model 70> The feature map generation model 70 is trained using multiple training datasets, which include training input images, ground truth keypoint information, and ground truth spatial feature maps. The training input images are image data in which one or more people are captured, such as the target image 10. The ground truth keypoint information indicates the location and body part represented by each keypoint 20 to be detected from the target image 10. The ground truth spatial feature map is the ideal spatial feature map 30 that should be output by the trained feature map generation model 70 when the corresponding training input image is input. The training datasets include ground truth spatial feature maps for each of a given pair of body parts.

[0075] In the following, the device used to train the feature map generation model 70 will be referred to as the "training device." The training device may be the same device as the keypoint mapping device 2000, or it may be a different device from the keypoint mapping device 2000. In the former case, it means that the keypoint mapping device 2000 also has the function of training the feature map generation model 70.

[0076] For each of a given pair of body parts, the training device can train the feature map generation model 70 for that pair as follows: The training device provides input data extracted from the training dataset to the feature map generation model 70 and obtains the spatial feature map 30 output from the feature map generation model 70. The training device calculates a loss based on the obtained spatial feature map 30 and the ground truth spatial feature map, and updates the trainable parameters of the feature map generation model 70. The above process can be repeatedly performed for each of multiple training datasets.

[0077] In one embodiment, the ground truth spatial feature map may be pre-generated by an administrator of the keypoint mapping device 2000, for example. For instance, the administrator operates a computer called a "dataset generation device" to display the training input images on a display device. The dataset generation device may be the same device as the keypoint mapping device 2000, the same device as the training device, or a different device from both the keypoint mapping device 2000 and the training device. The first case means that the keypoint mapping device 2000 is configured to also function as a dataset generation device.

[0078] Administrators and other personnel operate a dataset generation device to generate training datasets. For example, administrators receive training input images from the dataset generation device. Next, for each pair of predetermined body parts, administrators and other personnel specify keypoints for each person included in the provided training input images. The dataset generation device generates a ground truth spatial feature map based on the training input images and the specified keypoints.

[0079] Assume that the training input images contain individuals P1 and P2. Also assume that the ground truth spatial feature map is generated for a pair of left elbows and left wrists. In this case, an administrator or similar can specify the keypoints of person P1's left elbow and left wrist. Hereafter, the keypoints of person P1's left elbow and left wrist will be denoted as E1 and H1, respectively.

[0080] Depending on the specifications of E1 and H1, the dataset generator automatically generates directional regions R1 and R2 for each. The directional regions can be generated as areas having a predetermined shape and size (e.g., a circle with a predetermined radius, a square with a predetermined side length, etc.). The directional region for a particular keypoint is positioned based on the location of that keypoint. For example, the center of the directional region is positioned at the corresponding keypoint. That is, the center of the directional region for keypoint E1 is positioned at keypoint E1.

[0081] To generate directional regions R1 and R2, the dataset generator calculates the direction from E1 to H1 and identifies the pixel value corresponding to that direction. The identified pixel value is then set for all pixels within directional regions R1 and R2.

[0082] Similarly, administrators specify the keypoints of person P2's left elbow and left wrist. These are denoted as E2 and H2, respectively. In response to the specification of E2 and H2, the dataset generator generates direction regions R3 and R4 for E2 and H2, respectively. Specifically, the dataset generator calculates the direction from E2 to H2, identifies the pixel value corresponding to the calculated direction, and generates direction regions R3 and R4 with a predetermined shape and size, filled with that pixel value.

[0083] Here, the dataset generator may dynamically adjust the size of the directional regions in the ground truth spatial feature map so that they do not overlap. Assume that the predetermined shape and size of the directional region is a circle with radius R. In this case, if the distance between directional regions R1 and R2 in the ground truth spatial feature map is less than 2*R, then directional regions R1 and R2 will overlap. Therefore, the dataset generator reduces the size of directional regions R1 and R2 so that they do not overlap. An example of reducing the size of directional regions has already been described.

[0084] Furthermore, if the keypoint locations are represented by 3D coordinates, the dataset generator will generate a horizontal spatial feature map and a vertical spatial feature map according to the specified keypoints.

[0085] <Using Keypoint Groups> The results of keypoint mapping (i.e., keypoint groups 40) can be used in various ways. For example, keypoint groups 40 can be used for posture estimation. As a result of posture estimation, the type of posture held by the person corresponding to each keypoint group 40 can be estimated.

[0086] Furthermore, by performing pose estimation on each target image 10 of time-series data (e.g., video frames in a video), a time-series of the pose of each person captured in the target image 10 can be obtained. This time-series of the person's pose can be used to identify the action that the person is performing or the time-series of an action.

[0087] Programs can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs, CD-Rs, CD-R / Ws, and semiconductor memory (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, RAMs). Programs may also be provided to a computer using various types of transient computer-readable media. Examples of transient computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable media can be supplied to a computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels.

[0088] While this disclosure has been described with reference to embodiments, it is not limited to the embodiments described above. Various modifications that will be understood by those skilled in the art can be made to the configuration and details within the scope of the invention.

[0089] Some or all of the above embodiments may also be described as follows, but are not limited to the following: <Note> (Note 1) At least one memory configured to store instructions, It includes at least one processor configured to execute the instruction, and by executing the instruction, The process involves acquiring target images in which one or more people are captured, For each of the body parts of the person, key points are detected from the target image. For each of the predetermined pairs of body parts, a spatial feature map is generated using the target image, and the spatial feature map relating to the pair of body parts includes a first directional region corresponding to each key point representing the first body part of the pair, and a second directional region corresponding to each key point representing the second body part of the pair, wherein the first and second directional regions belonging to the same person represent the direction from the key point in the first directional region to the key point in the second directional region. A keypoint matching device that performs the following actions: for each person captured in the target image, generates a keypoint group that includes keypoints belonging to the same person. (Note 2) The generation of the key point group is performed for each of the predetermined pairs of body parts, For each of the key points of the first body part, the first directional region of that key point is detected from the pair of spatial feature maps based on the position of the key point and the predetermined shape and size of the first directional region. For each of the key points of the second body part, the second directional region of that key point is detected from the pair of spatial feature maps based on the position of the key point and the predetermined shape and size of the second directional region. For each of the aforementioned key points of the first body part, For each of the key points of the second body part, calculate the coefficient distance between that key point of the first body part and that key point of the second body part. This includes including the key point of the first body part and the key point of the second body part having the smallest coefficient distance in the same key point group, The keypoint matching device as described in Appendix 1, wherein the coefficient distance between the keypoint of the first body part and the keypoint of the second body part represents the degree of difference between the direction indicated by the first directional region of the keypoint of the first body part and the direction indicated by the second directional region of the keypoint of the second body part. (Note 3) The calculation of the coefficient distance between the key point of the first body part and the key point of the second body part is as follows: The statistical value of the pixel values ​​within the first directional region of the key point of the first body part is calculated as the direction indicated by the first directional region, The statistical value of the pixel values ​​within the second directional region of the key point of the second body part is calculated as the direction indicated by that second directional region, The keypoint matching device according to Appendix 2, which includes calculating the absolute value of the difference between the direction indicated by the first direction region and the direction indicated by the second direction region. (Note 4) The keypoint matching device according to Appendix 3, wherein the calculation of the coefficient distance between the keypoint of the first body part and the keypoint of the second body part further includes a process of correcting the absolute value of the difference by the Euclidean distance between the keypoints. (Note 5) The position of the aforementioned key point is represented by 3D coordinates, For each of the predetermined pairs of body parts, a horizontal spatial feature map and a vertical spatial feature map are generated as spatial feature maps for that pair. In the aforementioned horizontal spatial feature map, the first and second directional regions belonging to the same person represent the horizontal direction from the key point in the first directional region to the key point in the second directional region. In the aforementioned vertical spatial feature map, the first and second directional regions belonging to the same person represent the vertical direction from the key point in the first directional region to the key point in the second directional region, as described in Appendix 1, and the key point mapping device. (Note 6) The generation of the key point group is performed for each of the predetermined pairs of body parts, For each of the keypoints of the first body part, the first directional region of that keypoint is detected from the pair of the horizontal spatial feature map and the vertical spatial feature map, based on the position of the keypoint and the predetermined shape and size of the first directional region. For each of the keypoints of the second body part, the second directional region of that keypoint is detected from the pair of the horizontal spatial feature map and the vertical spatial feature map, based on the position of the keypoint and the predetermined shape and size of the second directional region. For each of the aforementioned key points of the first body part, For each of the key points of the second body part, calculate the coefficient distance between that key point of the first body part and that key point of the second body part. This includes including the key point of the first body part and the key point of the second body part having the smallest coefficient distance in the same key point group, The calculation of the coefficient distance between the key point of the first body part and the key point of the second body part is as follows: For each of the horizontal and vertical feature maps, the statistical value of the pixel values ​​within the first direction region of the key point of the first body part is calculated as the direction indicated by that first direction region. For each of the horizontal and vertical feature maps, the statistical value of the pixel values ​​within the second directional region of the key point of the second body part is calculated as the direction indicated by that second directional region. For each of the horizontal and vertical feature maps, the absolute value of the difference between the direction indicated by the first directional region and the direction indicated by the second directional region is calculated. A keypoint matching device as described in Appendix 5, which includes calculating the sum of the absolute values ​​of the aforementioned differences. (Note 7) The keypoint matching device according to Appendix 6, wherein the calculation of the coefficient distance between the keypoint of the first body part and the keypoint of the second body part further includes a process of correcting the sum of the absolute values ​​of the differences by the Euclidean distance between the keypoint of the first body part and the keypoint of the second body part. (Note 8) The process involves acquiring target images in which one or more people are captured, For each of the body parts of the person, key points are detected from the target image. For each of the predetermined pairs of body parts, a spatial feature map is generated using the target image, and the spatial feature map relating to the pair of body parts includes a first directional region corresponding to each key point representing the first body part of the pair, and a second directional region corresponding to each key point representing the second body part of the pair, wherein the first and second directional regions belonging to the same person represent the direction from the key point in the first directional region to the key point in the second directional region. A computer-based keypoint mapping method, which includes generating a keypoint group for each person captured in the target image, including keypoints that belong to the same person. (Note 9) The generation of the key point group is performed for each of the predetermined pairs of body parts, For each of the key points of the first body part, the first directional region of that key point is detected from the pair of spatial feature maps based on the position of the key point and the predetermined shape and size of the first directional region. For each of the key points of the second body part, the second directional region of that key point is detected from the pair of spatial feature maps based on the position of the key point and the predetermined shape and size of the second directional region. For each of the aforementioned key points of the first body part, For each of the key points of the second body part, calculate the coefficient distance between that key point of the first body part and that key point of the second body part. This includes including the key point of the first body part and the key point of the second body part having the smallest coefficient distance in the same key point group, The keypoint correspondence method described in Appendix 8, wherein the coefficient distance between the keypoint of the first body part and the keypoint of the second body part represents the degree of difference between the direction indicated by the first directional region of the keypoint of the first body part and the direction indicated by the second directional region of the keypoint of the second body part. (Note 10) The calculation of the coefficient distance between the key point of the first body part and the key point of the second body part is as follows: The statistical value of the pixel values ​​within the first directional region of the key point of the first body part is calculated as the direction indicated by the first directional region, The statistical value of the pixel values ​​within the second directional region of the key point of the second body part is calculated as the direction indicated by that second directional region, The key point correspondence method described in Appendix 9, which includes calculating the absolute value of the difference between the direction indicated by the first direction region and the direction indicated by the second direction region. (Note 11) The keypoint correspondence method according to Appendix 10, wherein the calculation of the coefficient distance between the keypoint of the first body part and the keypoint of the second body part further includes a process of correcting the absolute value of the difference by the Euclidean distance between the keypoints. (Note 12) The position of the aforementioned key point is represented by 3D coordinates, For each of the predetermined pairs of body parts, a horizontal spatial feature map and a vertical spatial feature map are generated as spatial feature maps for that pair. In the aforementioned horizontal spatial feature map, the first and second directional regions belonging to the same person represent the horizontal direction from the key point in the first directional region to the key point in the second directional region. In the aforementioned vertical spatial feature map, the first and second directional regions belonging to the same person represent the vertical direction from the key point in the first directional region to the key point in the second directional region, according to the key point mapping method described in Appendix 8. (Note 13) The generation of the key point group is performed for each of the predetermined pairs of body parts, For each of the keypoints of the first body part, the first directional region of that keypoint is detected from the pair of the horizontal spatial feature map and the vertical spatial feature map, based on the position of the keypoint and the predetermined shape and size of the first directional region. For each of the keypoints of the second body part, the second directional region of that keypoint is detected from the pair of the horizontal spatial feature map and the vertical spatial feature map, based on the position of the keypoint and the predetermined shape and size of the second directional region. For each of the aforementioned key points of the first body part, For each of the key points of the second body part, calculate the coefficient distance between that key point of the first body part and that key point of the second body part. This includes including the key point of the first body part and the key point of the second body part having the smallest coefficient distance in the same key point group, The calculation of the coefficient distance between the key point of the first body part and the key point of the second body part is as follows: For each of the horizontal and vertical feature maps, the statistical value of the pixel values ​​within the first direction region of the key point of the first body part is calculated as the direction indicated by that first direction region. For each of the horizontal and vertical feature maps, the statistical value of the pixel values ​​within the second directional region of the key point of the second body part is calculated as the direction indicated by that second directional region. For each of the horizontal and vertical feature maps, the absolute value of the difference between the direction indicated by the first directional region and the direction indicated by the second directional region is calculated. A keypoint correspondence method as described in Appendix 12, which includes calculating the sum of the absolute values ​​of the aforementioned differences. (Note 14) The keypoint correspondence method according to Appendix 13, wherein the calculation of the coefficient distance between the keypoint of the first body part and the keypoint of the second body part further includes a process of correcting the sum of the absolute values ​​of the differences by the Euclidean distance between the keypoint of the first body part and the keypoint of the second body part. (Note 15) The process involves acquiring target images in which one or more people are captured, For each of the body parts of the person, key points are detected from the target image. For each of the predetermined pairs of body parts, a spatial feature map is generated using the target image, and the spatial feature map relating to the pair of body parts includes a first directional region corresponding to each key point representing the first body part of the pair, and a second directional region corresponding to each key point representing the second body part of the pair, wherein the first and second directional regions belonging to the same person represent the direction from the key point in the first directional region to the key point in the second directional region. A non-temporary computer-readable storage medium containing a program that causes a computer to perform the following actions: generate a keypoint group for each person captured in the aforementioned target image, including keypoints that belong to the same person. (Note 16) The generation of the key point group is performed for each of the predetermined pairs of body parts, For each of the key points of the first body part, the first directional region of that key point is detected from the pair of spatial feature maps based on the position of the key point and the predetermined shape and size of the first directional region. For each of the key points of the second body part, the second directional region of that key point is detected from the pair of spatial feature maps based on the position of the key point and the predetermined shape and size of the second directional region. For each of the aforementioned key points of the first body part, For each of the key points of the second body part, calculate the coefficient distance between that key point of the first body part and that key point of the second body part. This includes including the key point of the first body part and the key point of the second body part having the smallest coefficient distance in the same key point group, The coefficient distance between the key point of the first body part and the key point of the second body part represents the degree of difference between the direction indicated by the first directional region of the key point of the first body part and the direction indicated by the second directional region of the key point of the second body part, as described in Appendix 15. (Note 17) The calculation of the coefficient distance between the key point of the first body part and the key point of the second body part is as follows: The statistical value of the pixel values ​​within the first directional region of the key point of the first body part is calculated as the direction indicated by the first directional region, The statistical value of the pixel values ​​within the second directional region of the key point of the second body part is calculated as the direction indicated by that second directional region, The storage medium according to Appendix 16, which includes calculating the absolute value of the difference between the direction indicated by the first direction region and the direction indicated by the second direction region. (Note 18) The storage medium according to Appendix 17, wherein the calculation of the coefficient distance between the key point of the first body part and the key point of the second body part further includes a process of correcting the absolute value of the difference by the Euclidean distance between the key points. (Note 19) The position of the aforementioned key point is represented by 3D coordinates, For each of the predetermined pairs of body parts, a horizontal spatial feature map and a vertical spatial feature map are generated as spatial feature maps for that pair. In the aforementioned horizontal spatial feature map, the first and second directional regions belonging to the same person represent the horizontal direction from the key point in the first directional region to the key point in the second directional region. In the aforementioned vertical spatial feature map, the first and second directional regions belonging to the same person represent the vertical direction from the key point in the first directional region to the key point in the second directional region, as described in Appendix 15. (Note 20) The generation of the key point group is performed for each of the predetermined pairs of body parts, For each of the keypoints of the first body part, the first directional region of that keypoint is detected from the pair of the horizontal spatial feature map and the vertical spatial feature map, based on the position of the keypoint and the predetermined shape and size of the first directional region. For each of the keypoints of the second body part, the second directional region of that keypoint is detected from the pair of the horizontal spatial feature map and the vertical spatial feature map, based on the position of the keypoint and the predetermined shape and size of the second directional region. For each of the aforementioned key points of the first body part, For each of the key points of the second body part, calculate the coefficient distance between that key point of the first body part and that key point of the second body part. This includes including the key point of the first body part and the key point of the second body part having the smallest coefficient distance in the same key point group, The calculation of the coefficient distance between the key point of the first body part and the key point of the second body part is as follows: For each of the horizontal and vertical feature maps, the statistical value of the pixel values ​​within the first direction region of the key point of the first body part is calculated as the direction indicated by that first direction region. For each of the horizontal and vertical feature maps, the statistical value of the pixel values ​​within the second directional region of the key point of the second body part is calculated as the direction indicated by that second directional region. For each of the horizontal and vertical feature maps, the absolute value of the difference between the direction indicated by the first directional region and the direction indicated by the second directional region is calculated. A storage medium as described in Appendix 19, which includes calculating the sum of the absolute values ​​of the aforementioned differences. (Note 21) The storage medium according to Appendix 20, wherein the calculation of the coefficient distance between the key point of the first body part and the key point of the second body part further includes a process of correcting the sum of the absolute values ​​of the differences by the Euclidean distance between the key point of the first body part and the key point of the second body part. [Explanation of Symbols]

[0090] 10 Target Images 20 Key Points 30 Spatial Feature Maps 32 direction area 40 Key Point Groups 50 Horizontal spatial feature maps 60 Vertical Spatial Feature Maps 70 Feature Extraction Models 72 First Model 74 Second Model 80 people 1000 computers 1020 Bus 1040 processor 1060 memory 1080 storage device 1100 Input / Output Interface 1120 Network Interface 2000 Keypoint Alignment Device 2020 Acquisition Department 2040 Keypoint Detection Unit 2060 Feature Map Generation Unit 2080 Keypoint Alignment Section

Claims

1. A keypoint detection unit that detects keypoints indicating the location of a person's body parts from an image, A feature map generation unit generates a spatial feature map including a first directional region corresponding to a first key point included in a predetermined pair of body parts, and a second directional region corresponding to a second key point included in the predetermined pair. It has a keypoint matching unit that generates a keypoint group including the aforementioned keypoints belonging to the same person, The first and second directional regions belonging to the same person represent the direction from the first key point to the second key point. The first directional region is a region having a predetermined shape and a predetermined size, The position of the first directional region is determined based on the position of the first key point. The second directional region is a region having the predetermined shape and predetermined size, The position of the second directional region is determined based on the position of the second key point in an image processing apparatus.

2. The keypoint matching unit performs the following for each of the predetermined pairs: For each of the first keypoints, the first directional region of the first keypoint is detected from the pair of spatial feature maps based on the position of the first keypoint and the predetermined shape and size of the first directional region. For each of the second keypoints, the second directional region of the second keypoint is detected from the pair of spatial feature maps based on the position of the second keypoint and the predetermined shape and size of the second directional region. Regarding each of the above first key points, For each of the aforementioned second key points, calculate the coefficient distance between its first key point and the second key point. The first key point and the second key point having the smallest coefficient distance are included in the same key point group. The image processing apparatus according to claim 1, wherein the coefficient distance represents the degree of difference between the direction indicated by the first direction region of the first key point and the direction indicated by the second direction region of the second key point.

3. The calculation of the coefficient distance is as follows: The statistical value of the pixel values ​​within the first direction region of the first key point is calculated as the direction indicated by that first direction region, The statistical value of the pixel values ​​within the second direction region of the second key point is calculated as the direction indicated by that second direction region, The image processing apparatus according to claim 2, comprising calculating the absolute value of the difference between the direction indicated by the first direction region and the direction indicated by the second direction region.

4. The image processing apparatus according to claim 3, wherein the calculation of the coefficient distance between the first key point and the second key point further includes a process of correcting the absolute value of the difference by the Euclidean distance between the first key point and the second key point.

5. The position of the aforementioned key point is represented by three-dimensional coordinates. The feature map generation unit generates a horizontal spatial feature map and a vertical spatial feature map for each of the predetermined pairs as the spatial feature maps for that pair. In the aforementioned horizontal spatial feature map, the first and second directional regions belonging to the same person represent the horizontal direction from the first key point in the first directional region to the second key point in the second directional region. The image processing apparatus according to claim 1, wherein in the vertical spatial feature map, the first directional region and the second directional region belonging to the same person represent the vertical direction from the first key point in the first directional region to the second key point in the second directional region.

6. The keypoint matching unit performs the following for each of the predetermined pairs: For each of the aforementioned first keypoints, the first directional region of the first keypoint is detected from the pair of the aforementioned horizontal spatial feature map and the aforementioned vertical spatial feature map, based on the position of the first keypoint and the predetermined shape and size of the first directional region. For each of the aforementioned second keypoints, the second directional region of the second keypoint is detected from the pair of the aforementioned horizontal spatial feature map and the aforementioned vertical spatial feature map, based on the position of the second keypoint and the predetermined shape and size of the second directional region. Regarding each of the above first key points, For each of the aforementioned second key points, calculate the coefficient distance between its first key point and the second key point. The first key point and the second key point having the smallest coefficient distance are included in the same key point group. The calculation of the coefficient distance between the first key point and the second key point is as follows: For each of the horizontal spatial feature map and the vertical spatial feature map, the statistical value of the pixel value within the first direction region of the first key point is calculated as the direction indicated by the first direction region. For each of the horizontal spatial feature map and the vertical spatial feature map, the statistical value of the pixel value within the second direction region of the second key point is calculated as the direction indicated by the second direction region. For each of the horizontal spatial feature map and the vertical spatial feature map, the absolute value of the difference between the direction indicated by the first directional region and the direction indicated by the second directional region is calculated. The image processing apparatus according to claim 5, comprising calculating the sum of the absolute values ​​of the differences.

7. The image processing apparatus according to claim 6, wherein the calculation of the coefficient distance further includes a process of correcting the sum of the absolute values ​​of the differences by the Euclidean distance between the first key point and the second key point.

8. Detecting keypoints that indicate the location of a person's body parts from an image, To generate a spatial feature map that includes a first directional region corresponding to a first key point included in a predetermined pair of body parts, and a second directional region corresponding to a second key point included in the predetermined pair, This includes generating a keypoint group that includes the aforementioned keypoints belonging to the same person, The first and second directional regions belonging to the same person represent the direction from the first key point to the second key point. The first directional region is a region having a predetermined shape and a predetermined size, The position of the first directional region is determined based on the position of the first key point. The second directional region is a region having the predetermined shape and predetermined size, A computer-based image processing method wherein the position of the second directional region is determined based on the position of the second keypoint.

9. The generation of the aforementioned keypoint group is performed for each of the predetermined pairs, For each of the first keypoints, the first directional region of the first keypoint is detected from the pair of spatial feature maps based on the position of the first keypoint and the predetermined shape and size of the first directional region. For each of the second keypoints, the second directional region of the second keypoint is detected from the pair of spatial feature maps based on the position of the second keypoint and the predetermined shape and size of the second directional region. Regarding each of the above first key points, For each of the aforementioned second key points, calculate the coefficient distance between its first key point and the second key point. The first key point and the second key point having the smallest coefficient distance are included in the same key point group. The image processing method according to claim 8, wherein the coefficient distance between the first key point and the second key point represents the degree of difference between the direction indicated by the first directional region of the first key point and the direction indicated by the second directional region of the second key point.

10. Detecting keypoints that indicate the location of a person's body parts from an image, To generate a spatial feature map that includes a first directional region corresponding to a first key point included in a predetermined pair of body parts, and a second directional region corresponding to a second key point included in the predetermined pair, The computer is made to generate a keypoint group that includes the aforementioned keypoints belonging to the same person, The first and second directional regions belonging to the same person represent the direction from the first key point to the second key point. The first directional region is a region having a predetermined shape and a predetermined size, The position of the first directional region is determined based on the position of the first key point. The second directional region is a region having the predetermined shape and predetermined size, A program in which the position of the second directional region is determined based on the position of the second key point.