Information processing device, information processing method, and information processing program
The information processing device accurately determines the subject's position by recognizing the skeleton and estimating the ground contact point, addressing the issue of partial obscuration and enhancing positional accuracy.
Patent Information
- Application Number
- JP2021203587
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-10-30
- Estimated Expiration
- 2041-12-15
AI Technical Summary
Existing systems struggle to accurately determine the position of a subject when parts of the subject are obscured by objects within the imaging unit's field of view, leading to incomplete recognition and positioning errors.
An information processing device that recognizes the subject's skeleton, estimates the contact position with the ground based on the skeleton, and identifies the subject's position using image information, even when the lower body is hidden.
Enables accurate identification of the subject's position by estimating the contact point with the ground, even when the lower body is obscured, improving positional accuracy and enabling tracking and situation estimation.
Smart Images

Figure 0007762554000001 
Figure 0007762554000002 
Figure 0007762554000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, an information processing method, and an information processing program. [Background technology]
[0002] 2. Description of the Related Art Conventionally, the position of a subject has been acquired based on image information generated by capturing an image of the subject. The device described in Patent Document 1 creates a difference image from two images captured by an imaging unit (a reference image serving as a reference and an image showing the subject), recognizes the subject from the difference image, and obtains the position of the recognized subject based on the angle of view, installation angle, and installation position of the imaging unit. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-181786 Summary of the Invention [Problem to be solved by the invention]
[0004] The device described in Patent Document 1 recognizes a subject and acquires its position. However, depending on the location where the imaging unit is installed and the surrounding conditions after the imaging unit is installed, an object may be placed within the angle of view of the imaging unit, causing part of the subject to be hidden by the object. In such cases, it may be difficult to accurately acquire the position of the subject based on the recognition of the subject, compared to when the entire subject is recognized.
[0005] The present disclosure provides an information processing device, an information processing method, and an information processing program for acquiring the position of a subject. [Means for solving the problem]
[0006] An information processing device of one embodiment includes a recognition unit that recognizes the skeleton of a subject based on image information in which the subject is recorded, a first estimation unit that estimates the contact position of the subject with the ground based on the skeleton of the subject recognized by the recognition unit, and an identification unit that identifies the position of the subject based on the contact position of the subject estimated by the first estimation unit. [Effects of the Invention]
[0007] According to one aspect, the skeleton of the subject can be recognized based on image information in which the subject is recorded, the contact position between the subject and the ground can be estimated based on the skeleton of the subject, and the position of the subject can be identified based on the estimated contact position of the subject. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an overview of an information processing device according to an embodiment. [Figure 2] FIG. 1 is a block diagram illustrating an information processing device according to an embodiment. [Figure 3] 1 is a diagram illustrating the skeleton of a person as an example of a subject, in which (A) shows an example of a person's skeleton, (B) shows an example of how feet are estimated based on the skeleton, (C) shows an example of how feet are estimated when the lower half of a person's body is hidden by an object, and (D) shows an example of how feet are estimated when the person's lower body is hidden by an object. [Figure 4] FIG. 1 is a diagram illustrating the skeleton of a sitting person as an example of a subject. [Figure 5] 1A and 1B are diagrams illustrating the skeletal structure of a person as an example of a subject when the person is crouching and kneeling, where (A) shows the person when crouching and (B) shows the person when kneeling. [Figure 6] 1A and 1B are diagrams illustrating the skeleton of a person as an example of a subject when the person is crouching and when the person is lying down. FIG. 1A shows the case where the person is crouching, and FIG. 1B shows the case where the person is lying down. [Figure 7] FIG. 10 is a diagram for explaining a skeleton when estimating a person's state (posture and behavior). [Figure 8] FIG. 2 is a block diagram illustrating an example of an acquisition unit. [Figure 9] FIG. 10 is a diagram for explaining a case where image position information is acquired. [Figure 10] 3A and 3B are diagrams for explaining an example of an image captured by an imaging unit; [Figure 11] 10 is a graph showing the relationship between the Y coordinate and the number of pixels. [Figure 12] 10 is a graph for explaining an example of image position information. [Figure 13] 10 is a graph showing the relationship between the Y coordinate and the ratio of the image width. [Figure 14] 1 is a flowchart illustrating an information processing method according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] An embodiment will be described below.
[0010] [Overview of information processing device 1] First, an overview of an information processing device 1 according to an embodiment will be described. FIG. 1 is a diagram illustrating an overview of an information processing device 1 according to an embodiment.
[0011] The information processing device 1 may be configured, for example, as a position identification device that identifies the position of the subject 110. The information processing device 1 may be configured, for example, as a situation estimation device that estimates the situation of the subject 110. The information processing device 1 may be configured, for example, as a tracking device that tracks (e.g., tracks) the subject 110 whose position has been identified.
[0012] The information processing device 1 may be, for example, a computer such as a server, a desktop, a laptop, or a tablet. The information processing device 1 receives image information in which a subject 110 is recorded. The information processing device 1 may receive image information generated by the imaging unit 11, for example, or may receive image information from an external camera, a server, or the like external to the information processing device 1. For example, based on image information, the information processing device 1 recognizes the skeleton of the subject 110. The skeleton of the subject 110 may be represented by, for example, feature points of the subject 110 and straight lines connecting adjacent feature points according to the shape of the body of the subject 110. The information processing device 1 estimates the contact position between the subject 110 and the ground based on, for example, the skeleton of the subject 110. That is, the information processing device 1 estimates the contact position between the skeleton of the subject 110 and the ground. In this case, as an example, when the lower half of the body of the subject 110 is hidden by an object 200 or the like and cannot be seen, the skeleton of the lower half of the body is estimated using the skeleton of the upper half of the body of the subject 110, and the contact position between the subject 110 and the ground is estimated. Here, the ground on which the subject 110 touches the ground may be a concept that includes the contact surface at the position where the subject 110 is located, such as, for example, an indoor floor. The information processing device 1 identifies the position of the subject 110 based on, for example, the contact position between the subject 110 and the ground estimated as described above. In this case, the information processing device 1 may identify, for example, the position of the subject 110 relative to the imaging unit 11 that generates image information. Furthermore, when the imaging unit 11 is disposed above the ground, the information processing device 1 may identify the distance between the position on the ground of the pole or the like on which the imaging unit 11 is installed and the contact position of the subject 110.
[0013] [Details of information processing device 1] Next, the information processing device 1 according to an embodiment will be described in detail. FIG. 2 is a block diagram illustrating an information processing device 1 according to an embodiment. Fig. 3 is a diagram for explaining the skeleton of a person as an example of a subject. Fig. 3(A) shows an example of a person's skeleton, Fig. 3(B) shows an example of estimating feet based on the skeleton, Fig. 3(C) shows an example of estimating feet when the lower half of a person's body is hidden by an object, and Fig. 3(D) shows an example of estimating feet when the person's lower body is hidden by an object. FIG. 4 is a diagram for explaining the skeleton of a sitting person as an example of a subject. 5A and 5B are diagrams for explaining the skeleton of a person as an example of a subject when the person is crouching and when the person is kneeling. Fig. 5A shows the person when crouching, and Fig. 5B shows the person when kneeling. 6A and 6B are diagrams illustrating the skeleton of a person as an example of a subject when the person is crouching and when the person is lying down. Fig. 6A shows the case where the person is crouching, and Fig. 6B shows the case where the person is lying down. FIG. 7 is a diagram for explaining the skeleton when estimating the state (posture and behavior) of a person.
[0014] As illustrated in FIG. 2 , the information processing device 1 includes, for example, an imaging unit 11, a communication unit 12, a storage unit 13, a display unit 14, and a control unit 20. The communication unit 12, the storage unit 13, and the display unit 14 may be an embodiment of an output unit. The control unit 20 includes, for example, a reception unit 21, a recognition unit 22, a first estimation unit 23, a second estimation unit 24, a third estimation unit 25, an acquisition unit 26, a specification unit 27, and an output control unit 28. The acquisition unit 26 may include, for example, a first acquisition unit 261, a second acquisition unit 262, a third acquisition unit 263, a fourth acquisition unit 264, a fifth acquisition unit 265, a sixth acquisition unit 266, a seventh acquisition unit 267, an eighth acquisition unit 268, and a ninth acquisition unit 269 (see FIG. 8 ). The control unit 20 may be configured by, for example, an arithmetic processing unit of the information processing device 1. The control unit 20 (e.g., a calculation processing device, etc.) may realize the functions of each unit (e.g., the reception unit 21, the recognition unit 22, the first estimation unit 23, the second estimation unit 24, the third estimation unit 25, the acquisition unit 26 (the first to ninth acquisition units 261 to 269), the identification unit 27, the output control unit 28, etc.) by, for example, appropriately reading and executing various programs, etc. stored in the memory unit 13, etc.
[0015] The imaging unit 11 captures an image of a subject and generates image information. The imaging unit 11 may be, for example, a camera capable of capturing still or moving images of the subject. In this case, when capturing still images, the imaging unit 11 may capture the still images continuously. That is, as an example, the imaging unit 11 may capture still images at predetermined time intervals.
[0016] As an example, the imaging unit 11 may be installed on a pole or the like fixed to the ground, or on a movable tripod or the like. That is, the imaging unit 11 may be, for example, a fixed-point imaging unit or the like that is capable of capturing images of a predetermined location constantly or for a certain period of time. When the imaging unit 11 is installed on, for example, a movable tripod, it can be moved together with the tripod. This allows the imaging unit 11 to be easily placed in a location desired by the user, and to be placed in the desired location only for the period of time required by the user. If there is distortion or the like in the image captured by the imaging unit 11, this may affect the positional accuracy when acquiring the position of the subject. For this reason, it is preferable that the imaging unit 11 uses an imaging lens or the like that can obtain an image with relatively little distortion.
[0017] The communication unit 12 is capable of transmitting and receiving various types of information to and from devices external to the information processing device 1 (external devices), for example.
[0018] The storage unit 13 may store, for example, various information and programs. Examples of the storage unit 13 may be a memory, a solid state drive, a hard disk drive, etc. Note that the storage unit 13 may be, for example, a storage area or a server on a cloud.
[0019] The display unit 14 is capable of displaying, for example, various characters, symbols, images, and the like.
[0020] The reception unit 21 receives image information. For example, the reception unit 21 may receive image information from the imaging unit 11. Furthermore, the reception unit 21 may receive image information from an external imaging unit (external camera) (not shown) outside the information processing device 1 via the communication unit 12, or may receive image information from a server (not shown) or the like outside the information processing device 1.
[0021] The recognition unit 22 recognizes the skeleton of the subject (see FIG. 3(A)) based on image information in which the subject is recorded. The subject may be, for example, a person, an animal, or a moving object such as a bicycle, a motorcycle, or an automobile. The recognition unit 22 may recognize, for example, a skeleton represented by feature points of the subject and lines connecting adjacent feature points. For example, if the subject is a person, the recognition unit 22 may recognize the feature points as various distinctive body parts such as the head, nose, shoulders, neck, elbows, wrists, groin, knees, and ankles. For example, the recognition unit 22 may recognize the skeleton of an animal, similar to a person, represented by feature points of body parts and lines connecting adjacent feature points. Furthermore, for a moving object with wheels, the recognition unit 22 may recognize the skeleton represented by feature points such as the center points of the wheels, points at the top and bottom of the wheels, and points at the edges or corners of the vehicle outline, and lines connecting adjacent feature points. As an example, the recognition unit 22 may recognize the skeleton of the subject using publicly known technology. In the present disclosure, a person will be exemplified as the subject.
[0022] Furthermore, when recognizing the skeleton of the subject (for example, a person) described above, the recognition unit 22 may recognize the waist as a feature point. As a specific example, the recognition unit 22 may estimate the center position of a line segment connecting the feature points of the groin of both feet as the feature point of the waist. The recognition unit 22 is not limited to the example of setting the feature point of the waist as described above, and may set the feature point of the waist using various methods. In this case, the recognition unit 22 may newly recognize the feature points of the waist and feet of the AI model using, for example, "human-pose-estimation-0001" or the like.
[0023] The recognition unit 22 may recognize whether the subjects are the same subject based on the skeleton of the subject. That is, the recognition unit 22 may estimate the identity of the subjects recorded in the image information based on, for example, whether the positions of the feature points of the skeleton and the lengths of adjacent feature points (lengths of straight lines) are the same (approximately the same). As an example, the recognition unit 22 may estimate the subjects as the same subject when the skeletons (positions of feature points and lengths of straight lines, etc.) of the subjects recorded in the image information generated at different times are the same (approximately the same), and may estimate the subjects as different subjects when the skeletons (at least one of the positions of feature points and lengths of straight lines, etc.) are different. Note that the recognition unit 22 may recognize whether the subjects are the same or not using various methods, not limited to the above example.
[0024] The first estimation unit 23 estimates the contact position between the subject and the ground based on the skeleton of the subject recognized by the recognition unit 22. That is, the first estimation unit 23 estimates the feature points of the subject's feet (contact positions between the subject and the ground) based on the skeleton of the subject. 3(B), the first estimation unit 23 may set the feature point of the feet at a position on an imaginary line extending downward from the feature point of the waist. Alternatively, the first estimation unit 23 may set the feature point of the feet on an imaginary line extending downward from a line segment connecting the feature points of the neck and waist.
[0025] In this case, as an example, the first estimation unit 23 may estimate the feature points of the feet from the positions of the two ankles or the knees and the position of the waist downward (toward the ground). That is, when the entire skeleton of the subject can be recognized, for example, the first estimation unit 23 may estimate the positions of the feature points of the feet using the positions of the feature points of the ankles (or knees) and the feature points of the waist. In this case, the first estimation unit 23 may newly estimate the feature points of the feet of the AI model using, for example, "human-pose-estimation-0001" or the like.
[0026] As a specific example, the first estimation unit 23 may set the foot feature point at a position on an imaginary line extending downward from the waist feature point toward the ground, corresponding to the height of the lower (or higher) ankle of the two ankle feature points. Alternatively, as another example, the first estimation unit 23 may set the foot feature point at a position on an imaginary line corresponding to the height of the middle (for example, the center or approximately the center) between the two ankles. The first estimation unit 23 is not limited to the example of setting the waist and foot feature points as described above, and may estimate the waist and foot feature points using various methods.
[0027] The first estimation unit 23 may also estimate the contact position between the subject and the ground based on the distance between at least two characteristic points of the subject's skeleton. 3(C), when the subject's skeleton below the waist is hidden by an object 200 or the like (when the subject's skeleton up to the waist is visible), the first estimation unit 23 may estimate a foot feature point at a position vertically downward from the waist based on the length between the neck feature point and the waist feature point. That is, the first estimation unit 23 may estimate that the foot feature point is located at a position below the waist feature point, a length obtained by multiplying the length between the neck feature point and the waist feature point by a predetermined coefficient. The value of the coefficient may be, for example, any value in the range of 1.2 to 1.8, or may be any value in a range outside this range.
[0028] 3(D), when the subject's skeleton below the neck is hidden by an object 200 or the like (when the subject's skeleton up to the neck is visible), the first estimation unit 23 may estimate a foot feature point at a position vertically downward from the neck based on the length between the nose feature point and the neck feature point. That is, the first estimation unit 23 may estimate that the foot feature point is located at a position below the neck feature point, at a length obtained by multiplying the length between the nose feature point and the neck feature point by a predetermined coefficient. The value of the coefficient may be, for example, any value in the range of 10 to 20, or may be any value in a range outside this range.
[0029] In other words, the first estimation unit 23 may estimate, as the ground contact position of the subject, the position where an imaginary line connecting at least two characteristic points of the skeleton of the upper body of the subject touches the ground. As an example, the first estimation unit 23 may estimate the ground contact position of the subject using a central part of both shoulders (for example, a characteristic point of the neck) and a part of the waist (a characteristic point of the waist) as the at least two characteristic points. In this case, when the state of the subject estimated by the second estimation unit 24 described later is estimated to be at least one of a case where the subject is standing, a case where the subject is sitting, and a case where the subject is crouching, the first estimation unit 23 may estimate the ground contact position of the subject using the central part of both shoulders and the part of the waist.
[0030] 3(B), for example, when the subject is standing, the first estimation unit 23 may set the X and Y coordinates of the feature points at the feet based on the Y coordinate of the lower of the two ankles of the person's skeleton in the image (reference Y coordinate) and the X coordinate of a perpendicular line drawn from the waist to the ground. As an example, the origin (0,0) of the X and Y coordinates may be set at the upper left corner of the image, with the horizontal direction being the X axis direction and the vertical direction being the Y axis direction.
[0031] As illustrated in FIG. 4, for example, when the subject is sitting, the first estimation unit 23 may set the X and Y coordinates of the feature point at the feet based on the Y coordinate (reference Y coordinate) of the lower of the two ankles of the person's skeleton in the image and the X coordinate of the perpendicular line from the waist.
[0032] Furthermore, as illustrated in FIG. 5(A), for example, when the subject is crouching, the first estimation unit 23 may set the X and Y coordinates of the feature point at the feet based on the Y coordinate (reference Y coordinate) of the lower of the two ankles of the person's skeleton in the image and the X coordinate of the perpendicular line drawn from the waist to the ground.
[0033] Furthermore, first estimation unit 23 may estimate, as the ground contact position of the subject, the position where an imaginary line set downward from the waist, which is a characteristic part of the subject's skeleton, touches the ground. In this case, when the state of the subject estimated by second estimation unit 24 is at least one of a case where the subject is kneeling, a case where the subject is crouching, or a case where the subject has fallen, first estimation unit 23 may estimate the ground contact position of the subject using the waist.
[0034] As illustrated in FIG. 5(B), for example, when the subject is kneeling, the first estimation unit 23 may set the X and Y coordinates of the feature point at the feet based on the Y coordinate of the lower of the two knees of the person's skeleton in the image (reference Y coordinate) and the X coordinate of the perpendicular line drawn from the waist to the ground.
[0035] As illustrated in FIG. 6, for example, when the subject is crouching or when the subject has fallen, the first estimation unit 23 may set the X and Y coordinates of the feature points at the feet based on the Y coordinate (reference Y coordinate) of the lowest of the two feature points of the ankle and knee of the skeleton of the person in the image and the X coordinate of the perpendicular line drawn from the waist to the ground.
[0036] The second estimation unit 24 may estimate the state of the subject based on the skeletal shape of the subject recognized by the recognition unit 22. For example, the second estimation unit 24 may estimate the state of the subject to be any one of standing, sitting, crouching, kneeling, crouching, and fallen.
[0037] Here, the second estimation unit 24 may estimate that the person is standing if, for example, the length between the waist feature point and the ankle feature point (length between the waist and the ankle) is longer than the length between the neck feature point and the waist feature point (length of the spine) and the skeleton as a whole is close to vertical. Here, the meaning of "length between the waist and the ankle" being longer than "length of the spine" may be a concept that includes a certain range, such as "almost longer" or "to some extent longer." Furthermore, "the skeleton as a whole is close to vertical" may include a concept that, for example, the neck feature point, waist feature point, knee feature point, and ankle feature point are close to forming an approximately straight line, and this straight line extends in the Y-axis direction (or a direction close to the Y-axis direction).
[0038] Furthermore, the second estimation unit 24 may estimate that the person is sitting if, for example, the "length between the waist and ankles" is longer than the "length of the spine" and the length between the waist characteristic point and the knee characteristic point (thigh height) is relatively short.
[0039] In addition, the second estimation unit 24 may estimate that the person is squatting, for example, when the "length between the waist and the ankles" is shorter than the "length of the spine" and the position of the knee characteristic point is higher than the position of the ankle characteristic point.
[0040] In addition, the second estimation unit 24 may estimate that the person is kneeling, for example, when the "length between the waist and the ankles" is shorter than the "length of the spine" and the position of the knee feature point is lower than the position of the ankle feature point.
[0041] Furthermore, the second estimation unit 24 may estimate that the person is crouching, for example, when the height between the neck feature point and the waist feature point (length of the spine in the Y-axis direction (height of the spine)) and the height between the knee feature point and the ankle feature point (height of the shin in the Y-axis direction) are relatively small, and the height between the waist feature point and the knee feature point (height of the thigh in the Y-axis direction) is relatively large.
[0042] Furthermore, the second estimation unit 24 may estimate that the person has fallen if, for example, the "spine height," "shin height," and "thigh height" are all relatively small.
[0043] As shown in FIG. 7, the second estimation unit 24 may estimate at least one of the posture and behavior of the subject, including posture, motion, hand movement, and facial movement, based on the skeletal shape of the subject. The second estimation unit 24 may estimate, for example, the above-described posture as the posture of the subject.
[0044] The second estimation unit 24 may also estimate at least one of the following groups of motions: when the subject is standing still, when the subject is walking, and when the subject is running, based on the shape and motion of the subject's skeleton.
[0045] Furthermore, the second estimation unit 24 may estimate at least one hand movement from a group of cases where the subject is doing nothing, where the hands are around the stomach, where the hands are around the face, and where the hands are raised, based on the shape and movement of the subject's skeleton. In this case, the second estimation unit 24 may use, for example, feature points of the neck, nose, elbow, neck, and hand (wrist) of the subject's skeleton, and estimate the hand positions from the positions of the feature points, the intervals between the feature points, etc. For example, the second estimation unit 24 may estimate that the subject is raising their hands if, in the subject's skeleton, the feature points of the hands are located above the feature points of the nose and the feature points of the elbows are located above the feature points of the neck. Furthermore, the second estimation unit 24 may estimate that the subject has raised their hand, for example, if the shorter of the length between the nose feature point and the hand feature point (nose-hand length) and the length between the neck feature point and the hand feature point (neck-hand length) is longer than the length between the neck feature point and the nose feature point (neck-nose length). In addition, the second estimation unit 24 may estimate that the hand is near the face if, for example, the longer of the "nose-hand length" and the "neck-hand length" is shorter than the "neck-nose length." In addition, the second estimation unit 24 may estimate that the hand is near the abdomen if, for example, the longer of the "nose-hand length" and the "neck-hand length" is longer than the "neck-nose length."
[0046] Furthermore, the second estimation unit 24 may estimate the facial movement of at least one of the following groups based on the shape and movement of the subject's skeleton: when the subject is doing nothing, when looking to the side, when looking up, when looking down, and when looking around. Here, for example, if the subject is looking around restlessly, the second estimation unit 24 may make an estimation based on the movement of the subject within a predetermined time period.
[0047] The third estimation unit 25 may recognize the face of the subject based on the image information and estimate the gender and age of the subject. The third estimation unit 25 can recognize the face of the subject and estimate the gender and age of the subject, for example, by using various known techniques. As a specific example, the third estimation unit 25 can recognize the face and estimate the gender and age by using "face-detection-retail-0004" and "age-gender-recognition-retail-0013", etc.
[0048] Here, the third estimation unit 25 may, for example, associate the recognized face with the position of the face in the skeletal shape recognized by the recognition unit 22. As an example, if the third estimation unit 25 can recognize a nose in the recognized face, it may associate the recognized nose with a feature point of the nose in the skeletal shape of the subject. As another example, if the third estimation unit 25 cannot recognize the nose but can recognize a neck, it may associate the recognized neck with a feature point of the neck in the skeletal shape of the subject.
[0049] For example, the acquisition unit 26 may generate the image position information in advance. Details of the acquisition unit 26 will be described later.
[0050] The identification unit 27 identifies the position of the subject based on the ground contact position of the subject estimated by the first estimation unit 23. In this case, the identification unit 27 may identify the position of the subject based on the image position information acquired by the acquisition unit 26, with the ground contact position corresponding to the imaging unit 11 as a reference. That is, for example, when a subject is imaged by the imaging unit 11, the identification unit 27 may identify the position of the subject based on image information in which the subject is present and image position information acquired by the acquisition unit 26. For example, when a subject is imaged by the imaging unit 11, the identification unit 27 may identify the position of the subject based on the positions (X, Y coordinates) of feature points at the feet of the subject in the image and the image position information. Furthermore, for example, the identification unit 27 may identify the position of the subject by identifying the distance from the ground position where the imaging unit 11 is located to the ground contact position of the subject estimated by the first estimation unit 23. As a specific example, when a subject is recorded in an image captured by the imaging unit 11, the identifying unit 27 may acquire the position (pixel position (X, Y coordinates)) of the subject in the image. The identifying unit 27 may compare the acquired pixel position of the subject with image position information to identify the distance from the imaging unit 11 to the subject.
[0051] The identification unit 27 may track the position of a subject recognized as the same subject by the recognition unit 22. Furthermore, in this case, the identification unit 27 may track the position of the subject by using, for example, at least one selected from the face of the subject and the gender and age group of the subject recognized by the third estimation unit 25. That is, the identification unit 27 can identify a subject (person) based on the person's skeletal shape, face, gender, age, etc. In this case, the identification unit 27 may assign, for example, one identification information (individual ID) to the same subject.
[0052] Here, the identification unit 27 may, for example, identify the same identification information (ID) based on the history of skeleton recognition, calculate the position of the subject's feet, and correct the position of the feet based on the position history of the feature points of the neck. That is, for example, when the skeleton of the subject is recognized by the recognition unit 22, the identification unit 27 may predict the position (predicted position) of the neck feature point at the current time from the history of the positions of the neck feature points of the skeleton (history of multiple points), and may assign identification information (ID) to the positions of the neck feature points that are close to the predicted positions actually detected by the recognition unit 22. Furthermore, the identification unit 27 may adjust past positions by correcting, for example, the history of the position of the neck feature point of the subject's skeleton recognized by the recognition unit 22. The identification unit 27 may calculate the position of the feet based on, for example, the positions of the ankle feature point and the waist feature point of the subject's skeleton. The identification unit 27 may also adjust past positions by correcting, for example, the calculated position of the feet based on the history.
[0053] Furthermore, for example, when the third estimation unit 25 recognizes the face of a subject and estimates the gender and age of the subject, the identification unit 27 may assign the same identification information (ID) if the face of the subject is recognizable. Thereafter, for example, when the face of the subject and the skeleton of the subject are associated with each other, if the face of the subject cannot be recognized, the identification unit 27 may estimate the identity of the subject and the gender and age of the subject by using past information from which the face of the subject can be recognized and the skeleton of the subject. Thereafter, when the face of the subject becomes recognizable, the identification unit 27 may estimate the identity of the subject and the gender and age of the subject based on a score indicating the degree of correspondence between the recognized face and a face recognized in the past (for example, when the score is equal to or greater than a threshold).
[0054] Furthermore, the specifying unit 27 may specify whether a subject (e.g., a person) for each identification information (ID) (individual ID) has entered a predetermined area set in an image. For example, when a subject has entered the predetermined area, the specifying unit 27 may record at least one selected from a group of the time when the subject for each individual ID entered the predetermined area, the time when the subject left the predetermined area, and the duration of stay in the predetermined area. Furthermore, the second estimation unit 24 etc. described above may estimate the state of the subject in the predetermined area (for example, events such as the subject's posture, actions, hand movements, and facial movements). The second estimation unit 24 etc. may acquire the occurrence time and duration of an event for each individual ID. The second estimation unit 24 may set a "warning" flag if a predetermined event continues for a predetermined time or longer. The predetermined event may be, for example, when the subject performs an event that is different from normal events. The second estimation unit 24 may monitor for each individual ID when the subject leaves the predetermined area and when the monitoring time has elapsed, and may terminate monitoring when at least one of these two cases occurs.
[0055] The output control unit 28 may control the output unit to output at least one of the position of the subject identified by the identification unit 27 and the warning corresponding to the state of the subject estimated by the second estimation unit 24. Here, the output control unit 28 may, for example, control the output unit to output a warning when the subject is in a state different from normal. The state different from normal may be, for example, an emergency state, an abnormal state, or the like. In other words, the output control unit 28 may, for example, control the output unit to output a warning when the "warning" flag is set as described above. The output unit may be, for example, a communication unit 12, a storage unit 13, a display unit 14, and the like. That is, the output control unit 28 may control the communication unit 12 to transmit, for example, information regarding the position of the subject identified by the identification unit 27 and at least one selected from a group of warnings according to the state of the subject estimated by the second estimation unit 24 to an external device (not shown). The external device may be, for example, various terminals, servers, etc. Here, as an example of a warning, the output control unit 28 may transmit information regarding the warning to the terminal via the communication unit 12 so that the terminal outputs a warning by voice, text, image, etc. In addition, the output control unit 28 may control the memory unit 13 to store, for example, information regarding the position of the subject identified by the identification unit 27 and at least one selected from a group of warnings according to the state of the subject estimated by the second estimation unit 24. In addition, the output control unit 28 may control the display unit 14 to display, for example, at least one selected from a group of warnings according to the position of the subject identified by the identification unit 27 and the state of the subject estimated by the second estimation unit 24.
[0056] As a specific example, the output control unit 28 may output at least one selected from the group consisting of the distance of the subject from the imaging unit 11 (or longitude, latitude, etc.), the age and gender of the subject, and the state of the subject (for example, events such as the subject's posture, movement, hand movement, and facial movement), store the file in the memory unit 13, and display it on a plan view (map) of the area imaged by the imaging unit 11.
[0057] [Details of Acquisition Section 26] Next, the acquisition unit 26 will be described. FIG. 8 is a block diagram illustrating an example of the acquisition unit 26. As shown in FIG.
[0058] The acquisition unit 26 may, for example, generate image position information to be used in the identification unit 27 described above in advance. The acquisition unit 26 acquires image position information within an image based on the height of the subject recorded in image information generated by the imaging unit 11 and the angle of view of the imaging unit 11. That is, the acquisition unit 26 acquires information relating to the position within the image (image position information) based on the image captured by the imaging unit 11. In detail, when the recognition unit 22 recognizes the subject (skeleton of the subject) based on the image information captured by the imaging unit 11, the acquisition unit 26 acquires position information within the image captured by the imaging unit 11 (image position information (information relating to the distance from the imaging unit 11 to the subject)) based on the height of the subject and the angle of view of the imaging unit 11.
[0059] FIG. 9 is a diagram used to explain the case where image position information is acquired. FIG. 10 is a diagram for explaining an example of an image captured by the imaging unit 11. As shown in FIG. Fig. 11 is a graph showing the relationship between the Y coordinate and the number of pixels. The horizontal axis of Fig. 11 represents the Y coordinate, and the vertical axis represents the number of pixels per meter. Here, the depth direction within the image captured by the imaging unit 11 is the Y coordinate, and the width direction of the image is the X coordinate. Fig. 12 is a graph for explaining an example of image position information, in which the horizontal axis represents the Y coordinate and the vertical axis represents the distance from the position on the ground where the imaging unit 11 is placed to an arbitrary Y coordinate position. Fig. 13 is a graph showing the relationship between the Y coordinate and the ratio of the image width, where the horizontal axis represents the Y coordinate and the vertical axis represents the ratio of the image width.
[0060] Here, in FIG. 9, the Y-axis direction (Y coordinate direction) is the depth direction. Also, in FIG. 9, when a subject 110 at a predetermined position in the Y-axis direction (Y coordinate direction) is imaged by the imaging unit 11, the image at that predetermined position is indicated by symbol 100a. Symbol w denotes the width (size in the X-axis direction) of the image 100a. If the center position of the image 100a is designated by symbol Pc, the distance from the imaging unit 11 to the center position Pc is designated as a first distance L1, and the distance from the center position Pc to the ground is designated as a second distance L2. Symbol θ denotes the angle between the ground and a line extending from the imaging unit 11 along the Y-axis direction to a position on the image 100a corresponding to the ground. Furthermore, the height of the imaging unit 11 from the ground is designated as H, the ground position at which the imaging unit 11 is disposed is designated as P0, and the distance from the ground position P0 along the Y-axis direction to the image 100a is designated as a third distance L3. The acquisition unit 26 acquires the third distance L3 and associates a position in the image captured by the imaging unit 11 with the third distance L3, thereby acquiring image position information.
[0061] Specifically, as illustrated in FIG. 8, the acquisition unit 26 includes a first acquisition unit 261, a second acquisition unit 262, a third acquisition unit 263, a fourth acquisition unit 264, a fifth acquisition unit 265, a sixth acquisition unit 266, a seventh acquisition unit 267, an eighth acquisition unit 268, and a ninth acquisition unit 269.
[0062] The first acquisition unit 261 acquires the height of the subject, the angle of view of the imaging unit 11, and the width of the image captured by the imaging unit 11 in advance. To acquire image position information, first, the subject 110 is moved within the imaging range of the imaging unit 11. The subject 110 is, for example, a person. When the subject 110 is a person, the height of the person (the height of the subject) is input using, for example, an input device (not shown). The first acquisition unit 261 acquires the input height of the person. Furthermore, the first acquisition unit 261 acquires the angle of view θ2 captured by the imaging unit 11. The angle of view θ2 is input using, for example, an input device (not shown). The first acquisition unit 261 acquires the input angle of view θ2. The angle of view θ2 is determined by the planar size of an imaging element (not shown) of the imaging unit 11, the focal length of an imaging lens (not shown) of the imaging unit 11, etc. Furthermore, the first acquisition unit 261 acquires the width (for example, the number of pixels) of the image obtained by the imaging unit 11. The width of the image (image width) is input using, for example, an input device (not shown).
[0063] The second acquisition unit 262 acquires the number of pixels constituting the image data per 1 meter based on the height of the subject acquired by the first acquisition unit 261. The second acquisition unit 262 acquires the number of pixels per 1 meter in the image captured by the imaging unit 11 at any position in the image by having the subject (person) 110 walk within the imaging range. That is, the second acquisition unit 262 acquires the size (length) of 1 meter in the image 100a based on the height of the person acquired by the first acquisition unit 261, and further acquires the number of pixels in the image per acquired size of 1 meter. 10, the size of the imaging element used in imaging unit 11 (the image obtained by imaging unit 11) is 1280 pixels x 960 pixels. When subject 110 (not shown in FIG. 10) is recorded in the image illustrated in FIG. 10, second acquisition unit 262 acquires the number of pixels per meter based on the height of subject 110 and the number of pixels of the imaging element (image).
[0064] The third acquisition unit 263 acquires the relationship between the Y coordinate, which is the coordinate in the depth direction of the image based on the image information, and the number of pixels per 1 m of the subject acquired by the second acquisition unit 262. In the example shown in Figure 10, the vertical direction of the image corresponds to the depth direction in real space, and the vertical coordinate of the image is the Y coordinate. Also, in the example shown in Figure 10, the horizontal direction of the image corresponds to the width direction in real space, and the horizontal coordinate of the image is the X coordinate. If the upper left pixel in Figure 10 is the origin coordinate (0,0), the coordinate of the upper right pixel is (1280,0). The coordinate of the pixel at the center of Figure 10 is (640,480). The coordinate of the lower left pixel in Figure 10 is (0,960), and the coordinate of the lower right pixel is (1280,960).
[0065] For example, by having a subject (person) walk vertically and horizontally within the imaging range of the imaging unit 11 (the range surrounded by coordinates (0,0), coordinates (1280,0), coordinates (1280,960), and coordinates (0,960)), the third acquisition unit 263 acquires the number of pixels per meter at multiple Y coordinate positions.
[0066] As shown in an example in Fig. 11, the relationship between the Y coordinate acquired by the third acquisition unit 263 and the number of pixels per meter is linear. Here, as shown in Fig. 9, the value of the Y coordinate increases as the subject approaches the imaging unit 11. In other words, the number of pixels per meter increases as the Y coordinate increases (as the subject approaches the imaging unit 11). In other words, the number of pixels per meter decreases as the Y coordinate decreases (as the subject moves farther away from the imaging unit 11).
[0067] The fourth acquisition unit 264 acquires the actual width w of the image 100a in which the subject 110 is detected at a predetermined position based on the width of the image acquired by the first acquisition unit 261 and the number of pixels per meter acquired by the second acquisition unit 262. Consider image 100a at the Y coordinate position (a predetermined position in the Y coordinate direction) where subject 110 is located, as exemplified in Fig. 9. Since the number of pixels in the width direction of the image captured by imaging unit 11 (1280 pixels in the example shown in Fig. 10) is acquired by first acquisition unit 261, fourth acquisition unit 264 acquires the actual length of width w of image 100a at the predetermined Y coordinate position based on the number of pixels per meter.
[0068] The fifth acquisition unit 265 acquires a first distance, which is the distance from the imaging unit 11 to the center position Pc in the image 100a, based on the width w of the actual image acquired by the fourth acquisition unit 264 and the angle of view θ2 acquired by the first acquisition unit 261, and also acquires a second distance, which is the distance from the ground to the center position Pc. The fifth acquisition unit 265 acquires the actual distance (first distance L1) from the imaging unit 11 to the center position Pc of the image 100a, based on the actual length of the width w of the image 100a at the predetermined position on the Y coordinate and the angle of view θ2 of the imaging unit 11. The acquisition unit 26 also acquires the actual distance (second distance L2) from the center position Pc of the image 100a to the ground, based on the number of pixels in the image, the number of pixels per meter, the actual length of the width w of the image 100a at the predetermined position on the Y coordinate, and the angle of view of the imaging unit 11.
[0069] The sixth acquisition unit 266 calculates the angle θ between the ground and the center line in the imaging direction of the imaging unit 11, based on the first distance L1 and the second distance L2 acquired by the fifth acquisition unit 265. That is, the sixth acquisition unit 266 calculates the angle θ between the ground and the imaging unit 11, based on the first distance L1 and the second distance L2 acquired by the fifth acquisition unit 265, by using trigonometric functions.
[0070] The seventh acquisition unit 267 acquires the height H from the ground to the imaging unit 11 based on the first distance L1 and the second distance L2 acquired by the fifth acquisition unit 265 and the formed angle θ acquired by the sixth acquisition unit 266. That is, the seventh acquisition unit 267 acquires the height H from the ground to the imaging unit 11 based on the first distance L1, the second distance L2, and the formed angle θ by using trigonometric functions.
[0071] The eighth acquisition unit 268 acquires the position from the imaging position where the imaging unit 11 is located to the subject based on the height H from the ground to the imaging unit 11 acquired by the seventh acquisition unit 267 and the angle θ acquired by the sixth acquisition unit 266, and further acquires image position information based on acquiring multiple positions of the subject. That is, the eighth acquisition unit 268 acquires a distance L3 from the ground position P0 where the imaging unit 11 is located to a predetermined position on the Y coordinate (image 100a) based on the height H from the ground to the imaging unit 11 and the angle θ between the imaging unit 11 and the ground. The eighth acquisition unit 268 acquires information about the position within the image (image position information) by associating the distance L3 with the Y coordinate of the predetermined position within the image. Furthermore, the eighth acquisition unit 268 acquires image position information by moving the subject 110 within the imaging range of the imaging unit 11 and acquiring the distance from the ground position P0 where the imaging unit 11 is located to the subject (person) 110 at any position within the imaging range. The eighth acquisition unit 268 stores the image position information in the storage unit 13.
[0072] The image position information is information that indicates the distance in meters from the imaging unit 11 to the position (coordinates) of any pixel in an image captured by the imaging unit 11. As a specific example, the image position information is information that indicates that the pixel coordinates (640,480) shown in Fig. 10 are 10 meters away from the imaging unit 11, the pixel coordinates (640,720) are 15 meters away from the imaging unit 11, and the pixel coordinates (640,840) are 20 meters away from the imaging unit 11.
[0073] 12, the image position information indicates that as the Y coordinate increases, the distance from ground position P0 where the imaging unit 11 is disposed decreases. In other words, the image position information indicates that as the Y coordinate decreases, the distance from ground position P0 where the imaging unit 11 is disposed increases.
[0074] Here, the ninth acquisition unit 269 acquires the width ratio to the Y coordinate based on the width at a reference position in the Y coordinate direction and the width at a predetermined position in the Y coordinate direction, out of the actual width w of each of the multiple images acquired by the fourth acquisition unit 264. That is, the fourth acquisition unit 264 acquires the actual image width w. At positions where the subject 110 is not present (Y coordinates corresponding to positions where no person is walking), the fourth acquisition unit 264 cannot acquire the actual image width w. Therefore, when the fourth acquisition unit 264 acquires image widths w at multiple Y coordinate positions, the ninth acquisition unit 269 calculates the image width at an arbitrary Y coordinate position based on the image widths w at those multiple positions. For example, the ninth acquisition unit 269 sets one of the multiple image widths acquired by the fourth acquisition unit 264 as the image width at the reference Y coordinate position (reference width), and calculates the ratio (width ratio) of the remaining image widths of the multiple images excluding the reference width to the reference width. The ninth acquisition unit 269 calculates the ratio (width ratio) to the reference width based on each of the multiple images acquired by the fourth acquisition unit 264. This width ratio can be plotted as shown in FIG. 13.
[0075] 13, the width ratio of the image decreases as the Y coordinate increases (as the distance from the imaging unit 11 decreases). In other words, the width ratio of the image increases as the Y coordinate decreases (as the distance from the imaging unit 11 increases).
[0076] The above-mentioned fifth acquisition unit 265 can acquire the width of the image at any Y coordinate by utilizing the width ratio of the image acquired by the ninth acquisition unit 269, even at a position where the image width w has not been acquired by the fourth acquisition unit 264, and therefore can acquire the first distance L1 and the second distance L2. This completes the acquisition of image position information.
[0077] [Specific example] The information processing device 1 sets the position of the imaging unit 11, a plan view (map) of the area that the imaging unit 11 captures, and the like. The information processing device 1 acquires image information (still image information, video information, etc.) generated by capturing images using various imaging units 11, including, for example, a USB camera, an IP camera, a 360° camera, a 4K / 8K camera, etc. The information processing device 1 performs data processing using, for example, the control unit 20 (e.g., an image recognition engine (AI / image processing)). The data processing may include, for example, subject skeletal recognition, subject facial recognition, object distance calculation, object tracking, global coordinate transformation, and map mapping. In the data processing, for example, a trained model may be generated by performing supervised learning using images (still images and videos) and data science such as data analysis, behavior analysis, and algorithm generation. In the data processing, for example, deep learning may be used to perform AI object detection (various objects including people, cars, motorcycles, etc.) and AI behavior detection (e.g., face, skeletal structure, age, gender, etc.). The information processing device 1 may output, for example, GUIs such as object detection time, object distance, object position, object identification information (ID), object information, and a floor plan (map) as output from the control unit 20. That is, the information processing device 1 can input, for example, an image captured by the imaging unit 11 to an image recognition engine, and perform image processing, deep learning, and the like to output the above-mentioned object information.
[0078] As an application example 1 of such an information processing device 1, it is possible to monitor the number of people waiting in line and the time they are waiting in line at a counter in a facility (for example, an ATM, etc.). As an application example 2, it is possible to monitor a victim (suspicious person) of a "bank transfer fraud" who is making a phone call at an ATM, etc. As an application example 3, it is possible to monitor lifesaving efforts, such as when a person has collapsed and someone nearby raises their hand and asks for help. As an application example 4, it is possible to take COVID-19 countermeasures at the entrance of a facility, such as checking whether visitors are taking their temperature and disinfecting themselves in a specified area. That is, the information processing device 1 (for example, the second estimation unit 24, etc.) can set a "warning" flag when the above-mentioned application examples 1 to 4, etc., are met.
[0079] [Information processing method] Next, an information processing method according to an embodiment will be described. FIG. 14 is a flowchart illustrating an information processing method according to an embodiment.
[0080] In step ST101, the imaging unit 11 captures an image of a subject and generates image information.
[0081] In step ST102, the recognition unit 22 recognizes the skeleton of the subject based on the image information generated in step ST101. In this case, the recognition unit 22 may recognize whether the subjects are the same subject based on the skeleton of the subject.
[0082] In step ST103, the first estimation unit 23 estimates the contact position between the subject and the ground based on the skeleton of the subject recognized in step ST102. In this case, the first estimation unit 23 may estimate the contact position between the subject and the ground based on the distance between at least two characteristic points of the skeleton of the subject. The first estimation unit 23 may estimate, as the ground contact position of the subject, the position where a virtual line connecting at least two characteristic points of the skeleton of the subject's upper body touches the ground. As an example, the first estimation unit 23 may estimate the ground contact position of the subject using a characteristic point of the neck and a characteristic point of the waist as the at least two characteristic points. In this case, when the state of the subject estimated by the second estimation unit 24 in step ST104 described later is estimated to be at least one of a case where the subject is standing, a case where the subject is sitting, and a case where the subject is crouching, the first estimation unit 23 may estimate the ground contact position of the subject using the characteristic point of the neck and the characteristic point of the waist. Furthermore, the first estimation unit 23 may estimate, as the ground contact position of the subject, the position where an imaginary line set downward from the waist, which is a characteristic part of the subject's skeleton, touches the ground. In this case, when the state of the subject estimated by the second estimation unit 24 in step ST104 described later is estimated to be at least one of a case where the subject is kneeling, a case where the subject is crouching, and a case where the subject is lying down, the first estimation unit 23 may estimate the ground contact position of the subject by using the characteristic point of the waist.
[0083] In step ST104, the second estimation unit 24 estimates the state of the subject based on the skeletal shape of the subject recognized in step ST102. That is, the second estimation unit 24 may estimate events such as the posture, movement, hand movement, and facial movement of the subject as the state of the subject. In this case, the second estimation unit 24 may set a "warning" flag, for example, when a predetermined event continues for a predetermined time or longer.
[0084] In step ST105, the third estimation unit 25 recognizes the face of the subject based on the image information generated in step ST101, and estimates the gender and age of the subject.
[0085] In step ST106, the identification unit 27 identifies the position of the subject based on the ground contact position of the subject estimated in step ST103. In this case, the identification unit 27 may identify the position of the subject based on the ground contact position corresponding to the imaging unit 11, based on the image position information acquired by the acquisition unit 26. In this case, the identification unit 27 may track the positions of subjects recognized as the same subject by the recognition unit 22.
[0086] In step ST107, the output control unit 28 may control the output unit to output, for example, at least one of the position of the subject identified in step ST106 and the state of the subject estimated in step ST104. The output unit may be, for example, the communication unit 12, the storage unit 13, the display unit 14, etc.
[0087] Each unit of the information processing device 1 described above may be realized as a function of a computer's arithmetic processing device or the like. That is, the reception unit 21, recognition unit 22, first estimation unit 23, second estimation unit 24, third estimation unit 25, acquisition unit 26 (first to ninth acquisition units 261 to 269), identification unit 27, and output control unit 28 (control unit 20) of the information processing device 1 may be realized as a reception function, a recognition function, a first estimation function, a second estimation function, a third estimation function, an acquisition function (first to ninth acquisition functions), an identification function, and an output control function (control function), respectively, by a computer's arithmetic processing device or the like. The information processing program can cause a computer to realize each of the above-described functions. The information processing program may be recorded on a computer-readable non-transitory recording medium, such as a memory, a solid-state drive, a hard disk drive, or an optical disk. Furthermore, as described above, each unit of the information processing device 1 may be realized by an arithmetic processing device of a computer or the like. The arithmetic processing device or the like is configured by, for example, an integrated circuit or the like. Therefore, each unit of the information processing device 1 may be realized as a circuit constituting the arithmetic processing device or the like. That is, the reception unit 21, the recognition unit 22, the first estimation unit 23, the second estimation unit 24, the third estimation unit 25, the acquisition unit 26 (first to ninth acquisition units 261 to 269), the identification unit 27, and the output control unit 28 (control unit 20) of the information processing device 1 may be realized as a reception circuit, a recognition circuit, a first estimation circuit, a second estimation circuit, a third estimation circuit, an acquisition circuit (first to ninth acquisition circuits), an identification circuit, and an output control circuit (control circuit) constituting the arithmetic processing device of a computer or the like. The imaging unit 11, communication unit 12, storage unit 13, and display unit 14 (output unit) of the information processing device 1 may be realized as, for example, an imaging function including the functions of an arithmetic processing device, as well as a communication function, a storage function, and a display function (output function). The imaging unit 11, communication unit 12, storage unit 13, and display unit 14 (output unit) of the information processing device 1 may be realized as, for example, an imaging circuit, a communication circuit, a storage circuit, and a display circuit (output circuit) by being configured using integrated circuits, etc. The imaging unit 11, communication unit 12, storage unit 13, and display unit 14 (output unit) of the information processing device 1 may be realized as, for example, an imaging device, a communication device, a storage device, and a display device (output device) by being configured using multiple devices.
[0088] The information processing device 1 can be configured by combining one or any combination of the above-described multiple units. In this disclosure, the term "information" is used, but the term "information" can be replaced with "data" and the term "data" can be replaced with "information."
[0089] [Aspects and Effects of the Present Embodiment] Next, one aspect of this embodiment and the effects of each aspect will be described. Note that this embodiment is not limited to the aspects described below, and may be realized by appropriately combining the above-mentioned parts. Furthermore, the effects described below are examples, and the effects of each aspect are not limited to those described below.
[0090] (Aspect 1) An information processing device of one embodiment includes a recognition unit that recognizes the skeleton of a subject based on image information in which the subject is recorded, a first estimation unit that estimates the contact position of the subject with the ground based on the skeleton of the subject recognized by the recognition unit, and an identification unit that identifies the position of the subject based on the contact position of the subject estimated by the first estimation unit. This allows the information processing device to identify the position of the subject relative to the imaging unit at the ground contact position of the subject. Furthermore, if a preset area (predetermined area) exists in the actual area recorded in the image, the information processing device can identify whether the subject is within the predetermined area.
[0091] (Aspect 2) An information processing device according to one aspect may include an imaging unit that captures an image of a subject and generates image information, and the identifying unit may identify the position of the subject based on a ground position corresponding to the imaging unit. This allows the information processing device to identify the position of the subject relative to the imaging unit at the ground contact position of the subject.
[0092] (Aspect 3) The information processing device of one aspect may include a second estimation unit that estimates the state of the subject based on the skeletal shape of the subject recognized by the recognition unit. This allows the information processing device to estimate the posture and behavior of the subject, and to estimate the ground position of the subject according to the posture, etc. Furthermore, once the information processing device recognizes the skeleton of the subject, it can estimate the posture and behavior of the subject by tracking characteristic points of the subject's skeleton, such as the neck and nose.
[0093] (Aspect 4) An information processing device according to one embodiment may include an output control unit that controls the output of at least one of the position of the subject identified by the identification unit and the output of a warning according to the state of the subject estimated by the second estimation unit. This allows the information processing device to allow the user of the information processing device to know the position and state of the person, and to make the user aware of a warning according to the state of the person.
[0094] (Aspect 5) In the information processing device of one aspect, the first estimation unit may estimate a contact position between the subject and the ground based on a distance between at least two characteristic points of the skeleton of the subject. This allows the information processing device to estimate the ground contact position of the subject and the ground based on the subject's skeleton even when a part of the subject (for example, the subject's feet) is hidden by an object or the like. The information processing device can identify the position of the subject relative to the imaging unit at the estimated ground contact position of the subject. As a specific example, the information processing device can estimate the ground contact position (position of the subject's feet) based on the positions of the subject's skeletal feature points, such as the waist and neck.
[0095] (Aspect 6) In the information processing device of one aspect, the first estimation unit may estimate, as the ground contact position of the subject, a position where an imaginary line connecting at least two characteristic points of the skeleton of the upper body of the subject touches the ground. This allows the information processing device to estimate the ground contact position (foot position) of the subject by utilizing the skeletal feature points of the subject's upper body.
[0096] (Aspect 7) In the information processing device of one aspect, the first estimation unit may estimate the ground contact position of the subject by using a center part of both shoulders and a waist part as the at least two characteristic parts. This allows the information processing device to estimate the ground contact position (foot position) of the subject by using skeletal feature points of the subject's upper body (the center of both shoulders (neck area) and the waist area).
[0097] (Aspect 8) In one aspect of the information processing device, when the state of the subject estimated by the second estimation unit is at least one of a case where the subject is standing, a case where the subject is sitting, and a case where the subject is crouching, the first estimation unit may estimate the ground contact position of the subject using the center of both shoulders and the waist area. This allows the information processing device to estimate the ground contact position (foot position) of the subject depending on the state of the subject.
[0098] (Aspect 9) In the information processing device of one aspect, the first estimation unit may estimate, as the ground contact position of the subject, a position where an imaginary line set downward from a waist part as a characteristic part of the subject's skeleton touches the ground. This allows the information processing device to estimate the ground contact position (foot position) of the subject by utilizing the feature points (waist area) of the subject's skeleton.
[0099] (Aspect 10) In one embodiment of the information processing device, the first estimation unit may estimate the ground contact position of the subject using the waist area when the state of the subject estimated by the second estimation unit is at least one of the following: the subject is kneeling, crouching, or fallen. This allows the information processing device to estimate the ground contact position (foot position) of the subject depending on the state of the subject.
[0100] (Aspect 11) In the information processing device of one aspect, the recognition unit may recognize whether the subjects are the same subject based on the bone structure of the subject. This allows the information processing device to identify each of the multiple subjects even when there are multiple subjects.
[0101] (Aspect 12) In the information processing device of one aspect, the identification unit may track the positions of the subjects recognized as the same subject by the recognition unit. This allows the information processing device to allow the user to understand the trajectory of the subject's movement, and also to identify the subject as the same subject even when the subject moves.
[0102] (Aspect 13) The information processing device of one aspect may include a third estimation unit that recognizes the face of a subject based on image information and estimates the gender and age of the subject. This allows the information processing device to, for example, once recognize the face of a subject, estimate the gender and age of the subject by tracking feature points of the subject's skeleton, such as the neck and nose, even if the subject's face becomes invisible.
[0103] (Aspect 14) In one aspect of the information processing method, a computer executes a recognition step of recognizing the skeleton of a subject based on image information in which the subject is recorded, a first estimation step of estimating the contact position of the subject with the ground based on the skeleton of the subject recognized by the recognition step, and an identification step of identifying the position of the subject based on the contact position of the subject estimated by the first estimation step. As a result, the information processing method can achieve the same effects as the information processing device of the above-described aspect.
[0104] (Aspect 15) An information processing program according to one embodiment causes a computer to realize a recognition function that recognizes the skeleton of a subject based on image information in which the subject is recorded, a first estimation function that estimates the contact position of the subject with the ground based on the skeleton of the subject recognized by the recognition function, and an identification function that identifies the position of the subject based on the contact position of the subject estimated by the first estimation function. As a result, the information processing program can achieve the same effects as the information processing device of the above-described aspect. [Explanation of symbols]
[0105] 1. Information processing equipment 11 Imaging unit 12 Communications Department 13 Storage section 14 Display section 20 Control Unit 21 Reception 22 Recognition part 23 1st estimation part 24 Second estimation part 25 Third estimation part 26 Acquisition Department 27 Specific section 28 Output control section
Claims
1. a recognition unit that recognizes the skeleton of a subject based on image information in which the subject is recorded; a first estimation unit that estimates a contact position between the subject and the ground based on the skeleton of the subject recognized by the recognition unit, and estimates, as the contact position of the subject, a position where a virtual line connecting at least two characteristic points of the skeleton of the upper body of the subject contacts the ground; an identification unit that identifies a position of the subject based on the ground position of the subject estimated by the first estimation unit; a second estimation unit that estimates a state of the subject based on the skeletal shape of the subject recognized by the recognition unit; an output control unit that controls the output of at least one of the position of the subject identified by the identification unit and the output of a warning according to the state of the subject estimated by the second estimation unit; An information processing device comprising:
2. an imaging unit that captures an image of the subject and generates image information; The identification unit identifies the position of the subject based on a ground position corresponding to the imaging unit. The information processing device according to claim 1 .
3. The first estimation unit estimates the ground position of the subject by using a center part of both shoulders and a waist part as the at least two characteristic parts. The information processing device according to claim 1 .
4. When the state of the subject estimated by the second estimation unit is at least one of a case where the subject is standing, a case where the subject is sitting, and a case where the subject is crouching, the first estimation unit estimates the ground contact position of the subject using the center of both shoulders and the waist. The information processing device according to claim 3 .
5. A recognition unit that recognizes the skeleton of a subject based on image information in which the subject is recorded; a first estimation unit that estimates a contact position between the subject and the ground based on the skeleton of the subject recognized by the recognition unit, and estimates, as the contact position of the subject, a position where an imaginary line set downward from a waist region as a characteristic part of the skeleton of the subject contacts the ground; an identification unit that identifies a position of the subject based on the ground position of the subject estimated by the first estimation unit; a second estimation unit that estimates a state of the subject based on the skeletal shape of the subject recognized by the recognition unit; an output control unit that controls the output of at least one of the position of the subject identified by the identification unit and the output of a warning according to the state of the subject estimated by the second estimation unit; An information processing device comprising:
6. The first estimation unit estimates the ground contact position of the subject using a waist region when the state of the subject estimated by the second estimation unit is at least one of a case where the subject is kneeling, a case where the subject is crouching, and a case where the subject is lying down. The information processing device according to claim 5 .
7. The recognition unit recognizes whether the subject is the same subject based on the skeleton of the subject. The information processing device according to any one of claims 1 to 6.
8. The identification unit tracks the positions of the subjects recognized as the same subject by the recognition unit. The information processing device according to claim 7 .
9. A third estimation unit is provided to recognize the face of the subject based on the image information and estimate the gender and age of the subject. The information processing device according to any one of claims 1 to 8.
10. The computer a recognition step of recognizing the skeleton of the subject based on image information in which the subject is recorded; a first estimation step of estimating a contact position between the subject and the ground based on the skeleton of the subject recognized in the recognition step, in which a position where a virtual line connecting at least two characteristic points of the skeleton of the upper body of the subject contacts the ground is estimated as the contact position of the subject; an identifying step of identifying a position of the subject based on the ground position of the subject estimated in the first estimating step; a second estimation step of estimating a state of the subject based on the skeletal shape of the subject recognized in the recognition step; an output control step of controlling to output at least one of an output of the position of the subject identified in the identifying step and an output of a warning according to the state of the subject estimated in the second estimating step; An information processing method that performs the above.
11. On the computer, a recognition function for recognizing the skeleton of a subject based on image information in which the subject is recorded; a first estimation function that estimates a contact position between the subject and the ground based on the skeleton of the subject recognized by the recognition function, and estimates a position where a virtual line connecting at least two characteristic points of the skeleton of the upper body of the subject contacts the ground as the contact position of the subject; an identification function that identifies a position of the subject based on a ground position of the subject estimated by the first estimation function; a second estimation function that estimates a state of the subject based on the skeletal shape of the subject recognized by the recognition function; an output control function that controls the output of at least one of the position of the subject identified by the identification function and the output of a warning according to the state of the subject estimated by the second estimation function; An information processing program that makes this possible.
Citation Information
Patent Citations
Electronic camera
JP2007280291A
Device, method, and program for detecting attitude change
JP2010237873A
Device, method and program for monitoring
JP2016181786A
External parameter estimation method, estimation device and estimation program of camera
JP2019102877A
Elevator user support system
JP2020059607A