Recognition device, recognition method, and program
The recognition device enhances posture recognition accuracy by determining keypoint positions and regions in time-series images, addressing inaccuracies in existing methods through region-based comparisons and confidence-level adjustments.
Patent Information
- Application Number
- PCT/JP2024/022129
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2025-12-26
AI Technical Summary
Existing techniques for recognizing changes in the posture of an object, such as human gestures, suffer from inaccuracies due to variations in keypoint extraction and posture, making it difficult to accurately recognize pose changes.
A recognition device that determines keypoint positions and regions based on time-series images, using confidence levels to adjust region sizes, and compares these regions with pre-defined patterns to recognize posture changes, reducing calculation load.
Accurately recognizes posture changes with high precision and reduced computational burden by utilizing region-based comparisons and confidence-level adjustments.
Smart Images

Figure JP2024022129_26122025_PF_FP_ABST
Abstract
Description
Recognition device, recognition method, and program
[0001] The present invention relates to a recognition device, a recognition method, and a program.
[0002] Various techniques have been proposed for recognizing changes in the posture of an object (such as gestures). Patent Literature 1 describes a technique for extracting human key points from images captured by an in-vehicle camera and recognizing human gestures based on the positions of the key points.
[0003] Special Publication No. 2020-533662
[0004] The accuracy of the keypoint extraction results may vary depending on the type of keypoint, the posture of the object, etc. Therefore, it may not be possible to accurately recognize a change in the posture of an object based solely on the positions of the keypoints. Some aspects of the present invention provide a technique for accurately recognizing a change in the posture of an object.
[0005] According to some embodiments, there is provided a recognition device for recognizing pose changes of an object, the recognition device comprising: a position determination means for determining positions of a plurality of key points for an object included in time-series images; a region determination means for determining a plurality of regions based on the determined positions of the plurality of key points; and a recognition means for recognizing pose changes of the object in the time-series images based on the plurality of regions determined for the time-series images.
[0006] According to some embodiments, changes in the posture of an object can be recognized with high accuracy.
[0007] Other features and advantages of the present invention will become apparent from the following description taken in conjunction with the accompanying drawings, in which the same or similar elements are designated by the same reference numerals.
[0008] The accompanying drawings, which are incorporated into and constitute a part of the specification, illustrate embodiments of the present invention and, together with the description, serve to explain the principles of the present invention. A schematic diagram illustrating an example hardware configuration of a computer in some embodiments. A flow diagram illustrating an example method for recognizing a pose change of an object in some embodiments. A schematic diagram illustrating an example method for determining a key region of an object in some embodiments. A schematic diagram illustrating an example pattern set in some embodiments. A flow diagram illustrating an example method for recognizing a pose change of an object in some embodiments. A schematic diagram illustrating a method for calculating an error between key regions in some embodiments. A schematic diagram illustrating a method for calculating an error between key regions in some embodiments.
[0009] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention as claimed, and not all combinations of features described in the embodiments are necessarily essential to the invention. Two or more of the features described in the embodiments may be combined in any desired manner. Furthermore, the same reference numerals are used to designate identical or similar components, and redundant descriptions will be omitted.
[0010] Referring to FIG. 1 , an example hardware configuration of a computer 100 according to some embodiments will be described. As described in detail below, the computer 100 is used to recognize a change in the posture of an object. Therefore, the computer 100 may be referred to as a recognition device. In the following embodiments, the object may be a living thing such as a human or an animal, or an inanimate object such as a robot or a power shovel. The change in posture of the object recognized by the computer 100 may be a change in posture of the entire object (e.g., the entire human body) or a part of the object (e.g., a human hand). A posture change intentionally made by a human may be referred to as a gesture.
[0011] The computer 100 may be, for example, a server computer or a personal computer (for example, a desktop or laptop computer). The computer 100 may also be a computer resource located in a cloud environment.
[0012] The computer 100 may include the hardware devices shown in Fig. 1. The processor 101 controls the overall operation of the computer 100. The processor 101 may be configured by, for example, a central processing unit (CPU). The processor 101 may be a single processor or a set of multiple processors connected to each other so that they can communicate with each other.
[0013] The memory 102 stores programs and data used in the processing of the computer 100. The memory 102 may be configured, for example, by a combination of random access memory (RAM) and read-only memory (ROM).
[0014] The input device 103 is a device for obtaining instructions from a user of the computer 100. The input device 103 may be configured, for example, by a combination of one or more of a keyboard, buttons, a touchpad, and a microphone. The display device 104 is a device for visually presenting information to a user of the computer 100. The display device 104 may be, for example, a dot-matrix display such as a liquid crystal display. The computer 100 may have a device (e.g., a touch screen) in which the input device 103 and the display device 104 are integrated. The input device 103 and the display device 104 may be external to the computer. In this case, the computer 100 may have an interface for communicating with the external input device 103 and display device 104.
[0015] The communication device 105 is a device for communicating with devices external to the computer 100. When the computer 100 performs wired communication, the communication device 105 may be a network interface card (NIC) having a connector for connecting a cable. When the computer 100 performs wireless communication, the communication device 105 may be a wireless communication module including an antenna and a baseband processing circuit.
[0016] The secondary storage device 106 is a device for non-volatilely storing programs and data used in the processing of the computer 100. The secondary storage device 106 is configured by, for example, a hard disk drive (HDD) or a solid state drive (SSD).
[0017] An example of a recognition method for recognizing a change in the posture of an object will be described with reference to Fig. 2. Each step of the method in Fig. 2 may be realized by the processor 101 executing a program loaded into the memory 102. Alternatively, the computer 100 may include a dedicated circuit such as an application specific integrated circuit (ASIC), and at least some of the steps of the method in Fig. 2 may be executed by this dedicated circuit.
[0018] In S201, the processor 101 acquires time-series images including an object for which pose change is to be recognized. In the following description, the time-series images including an object for which pose change is to be recognized are simply referred to as time-series images. The processor 101 may read time-series images stored in advance in the secondary storage device 106, or may read time-series images from an external device (e.g., a database). The time-series images include multiple images generated at different times. The time-series images may be a moving image obtained by photographing an object. The moving image is composed of multiple frames (images) obtained by photographing an object at different times.
[0019] An example of time-series images 300 will be described with reference to Fig. 3. In the example of Fig. 3, the time-series images 300 are composed of three images 301a to 301c. Alternatively, the time-series images 300 may be composed of any other number of images. Of the images 301a to 301c, it is assumed that the image 301a was generated the oldest and the image 301c was generated the newest. Each of the images 301a to 301c includes an object 302.
[0020] In S202, the processor 101 determines the positions of multiple key points related to an object for each image in the time-series image. A key point is a part of an object used to recognize the pose of the object. The key points may vary depending on the type of object for which pose changes are to be recognized. For example, if the object is a human body, the key points may be both shoulders, both elbows, both wrists, both hip joints, both knees, both ankles, and the head. If the object is a human hand, the key points may be the knuckles, fingertips, and wrist. If the object is a power shovel, the key points may be the arm joints. Furthermore, even if the type of object is the same, the key points may vary depending on the algorithm or its settings used to determine the key point positions.
[0021] The processor 101 further determines a confidence level of the keypoint position. The confidence level of the keypoint position refers to the likelihood that the keypoint actually exists at the determined position. For example, the confidence level may have a value between 0 and 1. When the confidence level has a value in this range, the confidence level can be interpreted as the probability that the keypoint actually exists at the determined position. The higher the confidence level of the keypoint position, the greater the likelihood that the keypoint actually exists at the determined position. The lower the confidence level, the less likely it is that the keypoint actually exists at the determined position. In other words, the lower the confidence level, the more likely it is that the keypoint actually exists at a position displaced from the determined position. The confidence level of the keypoint has a probability distribution within the image.
[0022] In the following description, data determined for multiple keypoints included in one image will be referred to as keypoint data, and the time series of keypoint data will be referred to as keypoint time-series data. The keypoint time-series data includes keypoint data for each of multiple images included in the time-series images. The keypoint data includes, for each of the multiple keypoints, a keypoint identifier, the keypoint position, and the reliability of the position. The keypoint identifier is information that identifies the type of keypoint (e.g., right shoulder). The position of the keypoint is represented, for example, by the coordinates of a point in the image.
[0023] The processor 101 may calculate the positions of the keypoints and their reliability. Alternatively, the processor 101 may transmit the time-series images to an external server and receive the positions of the keypoints and their reliability in response. The positions of the keypoints and their reliability may be calculated using existing techniques. Known examples of such techniques include YOLO pose and Blaze pose.
[0024] An example of keypoint time-series data 310 will be described with reference to Fig. 3. In the example of Fig. 3, the keypoint time-series data 310 includes keypoint data 311a to 311c for three images 301a to 301c included in the time-series image 300. Below, the keypoint data 311a will be described as a representative. The same description applies to the other keypoint data 311b to 311c.
[0025] The keypoint data 311a includes the identifier, position, and reliability of each of multiple keypoints of the object 302 included in the image 301a. The keypoint data 311a may be represented by an array. In FIG. 3, the positions of the keypoints included in the keypoint data 311a are indicated by points in the image 301a. In FIG. 3, only one of the points in the image 301a (specifically, the point representing the position of the right shoulder) is assigned a reference symbol 312. In FIG. 3, thirteen body parts are used as keypoints: both shoulders, both elbows, both wrists, both hips, both knees, both ankles, and the head. For the sake of explanation, the outer edge of the object 302 is indicated by a dashed line in FIG. 3, but the keypoint data 311a does not necessarily include information regarding the outer edge of the object 302. As shown in the keypoint data 311a to 311c in FIG. 3, the positions of the keypoints may vary depending on the posture of the object 302. Also, the confidence of the keypoint location may vary depending on the type of keypoint in an image, and may even vary between images for the same type of keypoint. The location of point 312 may be offset from the actual location of the keypoint.
[0026] In S203, the processor 101 determines a plurality of regions for each image in the time-series images based on the positions of the plurality of key points determined in S202. In the following description, a region determined based on the positions of the key points will be referred to as a key region. The shape of the key region may be any predetermined shape, such as a rectangle or a circle. For example, the processor 101 may determine the position of the key region so that the position of the key point is at the center of the key region. In the following description, a case where the key region is rectangular will be described. In the following description, a key region determined based on the position of a specific key point will sometimes be simply referred to as a key region of a specific key point.
[0027] The processor 101 may determine the size of the key region for each keypoint based on the confidence level of the position of each keypoint. The size of the key region may refer to the area of the key region or the diameter of the key region. The processor 101 may make the key region smaller as the confidence level of the position of each keypoint increases. The size of the key region may change linearly with the confidence level or may change according to another function.
[0028] In the following description, data related to multiple key regions included in one image will be referred to as key region data, and the time series of key region data will be referred to as key region time series data. The key region time series data includes key region data for each of multiple images included in the time series images. The key region data includes, for each of the multiple key regions, an identifier of a key point on which the key region is based, and information for specifying the position and size of the key region. If the key region is rectangular, the information for specifying the position and size of the key region may be the coordinates of the top left point of the key region and the coordinates of the bottom right point of the key region. If the key region is circular, the information for specifying the position and size of the key region may be the coordinates of the center of the key region and the radius of the key region.
[0029] An example of key region time-series data 320 will be described with reference to Fig. 3. In the example of Fig. 3, the key region time-series data 320 includes key region data 321a to 321c for three images 301a to 301c. The following description will be given using key region data 321a as a representative. The same description applies to the other key region data 321b to 321c.
[0030] The key region data 321a represents the multiple key regions 322 determined in S203. The key region data 321a may be represented by an array. For the sake of explanation, Fig. 3 illustrates the key regions 322 included in the key region data 321a using images.
[0031] In FIG. 3 , a key region 322 is determined for each of 13 body parts: both shoulders, both elbows, both wrists, both hip joints, both knees, both ankles, and the head. In FIG. 3 , only one of the 13 key regions 322 is labeled with a reference symbol. For the sake of explanation, points 312 representing the positions of key points are shown in FIG. 3 , but the key region data 321 a does not necessarily include the positions of the key points and their reliability. As shown by the key region data 321 a to 321 c in FIG. 3 , the key region 322 may vary depending on the posture of the object 302. Furthermore, the size of the key region 322 in one image may vary depending on the type of key point. For example, the size of the key region 322 determined based on the position (point 312) of the key point representing the right shoulder is determined based on the reliability of the position of the key point representing the right shoulder. The same applies to other types of key points.
[0032] In S204, the processor 101 recognizes a posture change of the object in the time-series images based on a plurality of key regions determined for the time-series images. For example, the processor 101 may determine whether a posture change of the object is a specific posture change based on a comparison between a plurality of key regions determined for the time-series images and a plurality of key regions of a region change pattern representing the specific posture change. A region change pattern is a pattern of changes in key regions. In the following description, a region change pattern will simply be referred to as a pattern.
[0033] An example of a pattern representing a specific posture change will be described with reference to FIG. 4 . A set of patterns representing a specific posture change is referred to as a pattern set 400. The pattern set 400 may be pre-stored in the secondary storage device 106 or may be read from an external device (e.g., a database). In the example of FIG. 4 , the pattern set 400 is composed of two patterns 401a and 401b. Alternatively, the pattern set 400 may be composed of any other number of patterns. Each pattern included in the pattern set 400 represents the change over time in key regions when an object assumes a specific posture. For example, pattern 401a represents the change over time in key regions of the left shoulder, left elbow, and left wrist when a person raises their left arm. Pattern 401b represents the change over time in key regions of the left hip joint, left knee, and left ankle when a person abducts their left leg. The following description will be given using pattern 401a as a representative. The same description applies to the other patterns 401b.
[0034] Pattern 401a is represented by time-series key region data 402a to 402c. Key region data 402a to 402c represent key region data at different times. The following explanation will be given using key region data 402a as a representative. The same explanation applies to the other key region data 402b to 402c.
[0035] The key region data 402a represents key regions 403 of multiple key points used to recognize a posture change in the pattern 401a. The key region data 402a may be represented by an array. For the sake of explanation, FIG. 4 illustrates the key regions 403 included in the key region data 402a using images. The multiple key points used to recognize a posture change in the pattern 401a may be all or some of the key points detectable for the object. For example, the multiple key points for recognizing the raising of the left arm are the left shoulder, left elbow, and left wrist. The key region data 402a includes, for each of the multiple key regions 403, an identifier of the key point associated with the key region 403 and information for specifying the position and size of the key region 403. The data structure of the key region data 402a may be the same as the data structure of the key region data 321a.
[0036] The pattern 401a may be generated by previously executing the above-described steps S201 to S203. For example, an apparatus for generating the pattern 401a (hereinafter referred to as a pattern generating apparatus) determines key regions related to specific key points (e.g., left shoulder, left elbow, and left wrist) by executing the above-described steps S201 to S203 on time-series images of multiple trials in which an object performs a specific posture change (e.g., raising the left arm). The multiple trials may be performed by multiple objects (e.g., multiple people), or may be performed multiple times by a single object.
[0037] The pattern generation device generates the pattern 401a by statistically processing each key region 322 of each image obtained through multiple trials. Specifically, the pattern generation device may determine a representative value (e.g., average value) of the positions of each key region 322 of each image obtained through multiple trials as the position of the key region 403 of the pattern 401a. The pattern generation device may determine a representative value (e.g., average value, first quartile, median, or third quartile) of the size of each key region 322 of each image obtained through multiple trials as the size of the key region 403 of the pattern 401a. Furthermore, the pattern generation device may generate multiple patterns for a specific posture change. For example, the pattern generation device may generate a pattern for each different representative value of the size of each key region 322 of each image obtained through multiple trials (e.g., for each of the first quartile, median, and third quartile). The multiple key regions 403 of these patterns are different for each pattern.
[0038] Furthermore, the pattern generation device may normalize the multiple key regions 403 included in the key region data 402a. For example, the pattern generation device may translate and scale the multiple key regions 403 so that the center of the positions of the multiple key regions 403 becomes the origin of a two-dimensional coordinate system and the variance of the distances from the center to the positions of the multiple key regions 403 becomes a predetermined value (for example, 1).
[0039] The process of S204 in Fig. 2 will be described in detail with reference to Fig. 5. In the method of Fig. 5, a posture change of an object is recognized by determining whether an object included in time-series images has undergone a specific posture change represented by a pattern included in the pattern set 400. In Fig. 5, it is determined whether an object included in time-series images has undergone a specific posture change represented by one pattern. The processor 101 may execute the method of Fig. 5 for all patterns included in the pattern set 400. The following describes pattern 401a. When the pattern set 400 includes multiple patterns for a specific posture change, it may be determined whether the object has undergone a specific posture change for each of the multiple patterns.
[0040] In S501, the processor 101 selects one image included in the time-series images acquired in S201. In the method of Fig. 5, it is determined whether or not the object has undergone a specific pose change in the portion of the time-series images starting from the selected image. The processor 101 may select various images included in the time-series images acquired in S201 to execute the method of Fig. 5. In the following, it is assumed that image 301a is selected.
[0041] In S502, key regions 322 of key points used in pattern 401a are extracted from the multiple key regions 322 in key region data 321a generated by executing S202 and S203 on image 301a, and the extracted multiple key regions 322 are normalized. In pattern 401a, the left shoulder, left elbow, and left wrist are used as key points. Therefore, from the multiple key regions 322 of key points throughout the body in key region data 321a, the key regions 322 of the left shoulder, left elbow, and left wrist are extracted. The processor 101 may translate and scale the multiple key regions 322 so that the center of the positions of the multiple key regions 322 becomes the origin of a two-dimensional coordinate system and the variance of the distances from the center to the positions of the multiple key regions 322 is 1. As described above, the multiple key regions 403 included in pattern 401a are normalized. Therefore, the multiple key regions 322 are also normalized for comparison with these multiple key regions 403.
[0042] In S503, the processor 101 selects the first key region data from among the unselected key region data of the pattern 401a. As will be described later, S503 is executed repeatedly. When S503 is executed for the first time, the first key region data 402a from among the key region data 402a to 402c of the pattern 401a is selected. When S503 is executed next time, the key region data 402b is selected.
[0043] In S504, the processor 101 selects one of the unselected key regions 403 of the key region data selected in S503 (for example, the key region data 402a). As will be described later, S504 is executed repeatedly. When S504 is executed for the first time, one of the three key regions 403 (the key regions 403 of the left shoulder, left elbow, and left wrist) of the key region data 402a is selected. These key regions 403 may be selected in any order.
[0044] In S505, the processor 101 calculates the error between the key region 403 selected in S504 and the key region 322 of the selected image, which is the key point related to this key region 403. For example, if the left shoulder key region 403 is selected in S504, the error between this left shoulder key region 403 and the left shoulder key region 322 of the selected image is calculated.
[0045] 6A and 6B, a method for calculating the error between key region 322 and key region 403 will be described. As will be described below, the error between key region 322 and key region 403 may be based on the distance between a point on the outer edge of key region 322 and a point on the outer edge of key region 403.
[0046] 6A shows a case where the key region 322 and the key region 403 are rectangular. Each side of the key region 322 and the key region 403 is parallel to the x-axis or y-axis. The error between the key region 322 and the key region 403 may be the sum of the distance between the top-left point 322a of the key region 322 and the top-left point 403a of the key region 403, and the distance between the bottom-right point 322b of the key region 322 and the bottom-right point 403b of the key region 403. These distances may be Euclidean distance, Manhattan distance (i.e., the sum of the difference between the x component and the y component), or other distances. By using the Manhattan distance, the error can be calculated simply by calculating the difference between the x component and the y component, thereby reducing the burden of calculation processing.
[0047] 6B shows a case where key region 322 and key region 403 are circular. The distance may be the sum of the maximum difference between the outer edge of key region 322 and the outer edge of key region 403 (i.e., the distance between point 322c and point 403d) and the minimum difference between the outer edge of key region 322 and the outer edge of key region 403 (i.e., the distance between point 322d and point 403c). These distances may be Euclidean distance, Manhattan distance (i.e., the sum of the difference in the x component and the difference in the y component), or other distances.
[0048] In S506, the processor 101 determines whether or not there is a key region 403 that was not selected in S504 in the selected key region data. If it is determined that there is an unselected key region 403 ("YES" in S506), the processor 101 transitions the process to S504, and otherwise ("NO" in S506), the processor 101 transitions the process to S507. As a result, errors are calculated in S505 for all key regions 403 included in the selected key region data.
[0049] In S507, the processor 101 determines whether or not there is any unselected key region data in the pattern 401a in S503. If it is determined that there is any unselected key region data (YES in S507), the processor 101 transitions the process to S508, and otherwise (NO in S507), the processor 101 transitions the process to S509.
[0050] In S508, the processor 101 selects the image next (i.e., later in time) to the currently selected image in the time-series images, selects the image next (i.e., later in time) to the currently selected key region data in pattern 401a, and returns the process to S502. This determines the posture of the object in the temporally later image. For example, if the currently selected image is image 301a, image 301b is selected next, and if the currently selected key region data is key region data 402a, key region data 402b is selected next. This executes S502 to S506 for all key region data included in pattern 401a.
[0051] In S509, the processor 101 determines whether the object included in the time-series images 300 has undergone the specific posture change represented by the pattern 401 a, based on the average value of the multiple errors calculated in S505, which are repeatedly executed in the above-mentioned step. For example, the processor 101 may determine that the object included in the time-series images 300 has undergone the specific posture change represented by the pattern 401 a, if the average value of the errors is smaller than a preset threshold. In this way, the posture change of the object included in the time-series images 300 is recognized.
[0052] In the above-described method, the multiple key region data 321a to 321c of the time-series image 300 are compared with the key region data 402a to 402c of the pattern 401a in order from oldest to newest. Alternatively, the multiple key region data 321a to 321c of the time-series image 300 may be compared with the key region data 402a to 402c of the pattern 401a in order from newest to newest, or in another order. In the above-described method, a pose change of an object is recognized by comparing the multiple key region data 321a to 321c of the time-series image 300 with the key region data 402a to 402c of the pattern 401a. Alternatively, the processor 101 may recognize a pose change of an object by inputting the key region time-series data 320 of the time-series image 300 into a model for estimating a pose change of an object. This model may be generated in advance by machine learning.
[0053] As described above, the above-described method uses regions based on the reliability of the keypoint positions, allowing for accurate recognition of changes in the pose of an object. Furthermore, since changes in the pose of an object can be recognized simply by calculating the distance between regions in step S505, the calculation load can be reduced.
[0054] Summary of Embodiments (Item 1) A recognition device (100) for recognizing a pose change of an object (302), comprising: a position determination means for determining positions (312) of a plurality of key points related to the object included in time-series images (300); a region determination means for determining a plurality of regions (322) based on the determined positions of the plurality of key points; and a recognition means for recognizing a pose change of the object in the time-series images based on the plurality of regions determined for the time-series images. According to this item, a pose change of the object can be recognized with high accuracy. (Item 2) The recognition device according to item 1, wherein the recognition means determines whether the object has undergone a specific pose change based on an error between the plurality of regions determined for the time-series images and a plurality of regions (403) of a region change pattern (401 a, 401 b) representing a specific pose change. According to this item, it is possible to accurately recognize whether an object has undergone a specific pose change. (Item 3) The recognition device according to Item 2, wherein the error is based on the distance between points (322a-322d) on the outer edges of the plurality of regions determined for the time-series images and points (403a-403d) on the outer edges of the plurality of regions in the region change pattern. According to this item, it is possible to recognize whether an object has undergone a specific pose change with a low processing load. (Item 4) The recognition device according to Item 2 or 3, wherein the recognition means determines whether the object has undergone the specific pose change based on the error between the plurality of regions determined for the time-series images and each of a plurality of regions in a plurality of region change patterns representing the specific pose change, and the sizes of the plurality of regions in the plurality of region change patterns vary for each region change pattern. According to this item, it is possible to accurately recognize whether an object has undergone a specific pose change. (Item 5) The recognition device according to any one of Items 1 to 4, further comprising a reliability determination means for determining reliability of the positions of the plurality of key points, and the region determination means for determining the size of each of the plurality of regions based on the reliability. According to this item, changes in the posture of an object can be recognized with high accuracy.(Item 6) The recognition device according to Item 5, wherein the region determination means reduces the size of a region among the plurality of regions that is located at a higher reliability. According to this item, it is possible to accurately recognize a change in the posture of an object. (Item 7) A program for causing a computer (100) to function as each means of the recognition device according to any one of Items 1 to 6. According to this item, it is possible to accurately recognize a change in the posture of an object. (Item 8) A method for recognizing a change in the posture of an object (302), the method comprising: determining (S202) positions (312) of a plurality of key points related to the object included in time-series images (300); determining (S203) a plurality of regions (322) based on the determined positions of the plurality of key points; and recognizing (S204) a change in the posture of the object in the time-series images based on the plurality of regions determined for the time-series images. According to this item, it is possible to accurately recognize a change in the posture of an object.
[0055] The invention is not limited to the above-described embodiment, and various modifications and variations are possible within the scope of the gist of the invention.
Claims
1. A recognition device for recognizing changes in the posture of an object, comprising: a position determination means for determining the positions of a plurality of key points relating to an object included in time-series images; a region determination means for determining a plurality of regions based on the determined positions of the plurality of key points; and a recognition means for recognizing changes in the posture of the object in the time-series images based on the plurality of regions determined for the time-series images.
2. The recognition device according to claim 1, wherein the recognition means determines whether the object has undergone the specific pose change based on the error between the multiple regions determined for the time-series images and multiple regions of a region change pattern representing the specific pose change.
3. The recognition device according to claim 2, wherein the error is based on the distance between a point on the outer edge of the plurality of regions determined for the time-series images and a point on the outer edge of the plurality of regions in the region change pattern.
4. The recognition device described in claim 2 or 3, wherein the recognition means determines whether the object has undergone the specific posture change based on the error between the multiple regions determined for the time-series images and the multiple regions of each of the multiple region change patterns representing the specific posture change, and the sizes of the multiple regions of the multiple region change patterns differ for each region change pattern.
5. The recognition device according to any one of claims 1 to 4, further comprising a reliability determination means for determining the reliability of the positions of the plurality of key points, and the region determination means for determining the size of each of the plurality of regions based on the reliability.
6. The recognition device according to claim 5, wherein said region determining means reduces the size of a region among said plurality of regions as the region is closer to said position with higher reliability.
7. A program for causing a computer to function as each means of the recognition device according to any one of claims 1 to 6.
8. A method for recognizing pose changes of an object, comprising: determining the positions of a plurality of key points for an object included in time-series images; determining a plurality of regions based on the determined positions of the plurality of key points; and recognizing pose changes of the object in the time-series images based on the plurality of regions determined for the time-series images.
Citation Information
Patent Citations
Position and orientation estimation method, device, electronic device, and storage medium
JP2021517649A
Image processing method and apparatus, electronic device, and recording medium
JP2023511243A