Human posture determination method and mobile machine using the same
By acquiring color and depth data through a rangefinder camera, detecting key points of the human skeleton, and using predefined feature maps to determine the pose, the problem of low efficiency in human pose detection under occlusion conditions is solved, and accurate pose determination is achieved under occlusion conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UBKANG (QINGDAO) TECH CO LTD
- Filing Date
- 2022-04-06
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies are inefficient in detecting human posture, especially when the human body is obscured by obstacles or clothing, and cannot adapt to different sets of key points.
Images with color and depth data are acquired by a ranging camera to detect estimated skeletal key points of the human body. Predefined feature maps are used to determine the human posture and adapt to different sets of key point detection.
It enables accurate determination of human posture even under occlusion conditions, improving the adaptability and efficiency of detection.
Smart Images

Figure CN116830165B_ABST
Abstract
Description
Technical Field
[0001] This application relates to human posture detection technology, and more particularly to a method for judging human posture and a mobile machine using the method. Background Technology
[0002] Leveraging the rapidly developing technology of artificial intelligence (AI), mobile robots have been applied to various aspects of daily life to provide services such as healthcare, housework, and transportation. Taking healthcare as an example, mobility robots are typically designed as walking aids or wheelchairs to assist with walking and other activities, thereby improving the mobility of people with disabilities.
[0003] In addition to automatically navigating and assisting users in a more automated and convenient way, walking assistance robots inevitably need to detect the user's posture in order to provide appropriate services. Skeleton-based posture determination is a common technique for human posture determination in robots. It detects the human posture based on key points identified on the estimated human skeleton.
[0004] With a sufficient number of identified key points, detection can be performed effectively and accurately; otherwise, when key points cannot be identified, such as when the human body is obscured by obstacles or clothing, the detection efficiency will be affected. This is especially true when a person is sitting behind furniture or lying in bed covered by a blanket, as the furniture or blanket may obscure their body, affecting the detection results. Therefore, a human posture determination method that can adapt to different sets of key points is needed. Summary of the Invention
[0005] This application provides a body posture determination method and a mobile machine using the method to detect the posture of an occluded human body, thereby solving the problems existing in the aforementioned prior art human posture determination technology.
[0006] Embodiments of this application provide a method for determining human posture, including: One or more range-measuring images are acquired using a range-measuring camera, wherein the one or more range-measuring images include color data and depth data; Detect multiple key points of the estimated skeleton of the human body in the color data, and calculate the positions of the detected key points based on the depth data, wherein the estimated skeleton has a set of predefined key points; Based on the key points detected in the set of predefined key points, a feature map is selected from a set of predefined feature maps; Based on the location of the detected key points, two features of the human body corresponding to the selected feature map are obtained; and The posture of the human body is determined based on the two features in the selected feature map.
[0007] Embodiments of this application also provide a mobile machine, including: Rangefinder camera; One or more processors; and One or more memories storing one or more computer programs, which are executed by one or more processors, wherein the one or more computer programs include a plurality of instructions for: The ranging camera acquires one or more ranging images, wherein the one or more ranging images include color data and depth data; Detect multiple key points of the estimated skeleton of the human body in the color data, and calculate the positions of the detected key points based on the depth data, wherein the estimated skeleton has a set of predefined key points; Based on the key points detected in the set of predefined key points, a feature map is selected from a set of predefined feature maps; Based on the location of the detected key points, two features of the human body corresponding to the selected feature map are obtained; and The posture of the human body is determined based on the two features in the selected feature map.
[0008] As can be seen from the embodiments of this application above, the human posture determination method provided by this application uses a predefined feature map describing the human posture to determine the human posture based on the detected human key points. Therefore, it can adapt to different sets of key points and solve the problem that the prior art cannot detect occluded human bodies. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments or the prior art are briefly introduced below. In the following drawings, the same reference numerals denote corresponding parts throughout the figures. It should be understood that the drawings in the following description are merely examples of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 These are schematic diagrams illustrating scenarios where a mobile machine is used to detect human posture in some embodiments of this application.
[0011] Figure 2 Is using Figure 1 A diagram illustrating how a camera on a mobile machine detects human posture.
[0012] Figure 3 This is an explanation Figure 1A schematic block diagram of a mobile machine.
[0013] Figure 4 Is using Figure 1 A schematic diagram illustrating an example of using a mobile machine to detect human posture.
[0014] Figure 5A and 5B This is a flowchart of the attitude determination process based on detected key points.
[0015] Figure 6 yes Figure 5A and 5B A schematic diagram of scenario 1 for attitude determination.
[0016] Figure 7 yes Figure 6 Scene characteristics Figure 1 A schematic diagram.
[0017] Figure 8 yes Figure 5A and 5B A schematic diagram of scenario 2 for attitude determination.
[0018] Figure 9 yes Figure 8 Scene characteristics Figure 2 A schematic diagram.
[0019] Figure 10 Is Figure 6 This is a schematic diagram illustrating an example of obtaining the detected interior angles and body proportions in a scene.
[0020] Figure 11 Is Figure 8 This is an illustration of an example of obtaining the detected body proportions and upper body angles in a given scene.
[0021] Figure 12 yes Figure 5A and 5B A schematic diagram of scenario 3 for attitude determination. Detailed Implementation
[0022] To make the objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0023] It should be understood that, when used in this application and the appended claims, the terms “comprising,” “including,” “having,” and variations thereof indicate the presence of the said features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or sets thereof.
[0024] It should also be understood that the terminology used in the specification of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. As used in the specification and appended claims of this application, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0025] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the relevant listed items and all possible combinations, and includes such combinations.
[0026] The terms "first," "second," and "third" used in this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or the number of technical features referred to. Therefore, a feature defined by "first," "second," and "third" may explicitly or implicitly include at least one of those technical features. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly defined.
[0027] The use of terms such as "one embodiment" or "some embodiments" in the specification of this application means that one or more embodiments of this application may include specific features, structures, or characteristics related to the description of that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different places in the specification do not mean that the described embodiments should be referenced by all other embodiments, but rather that they are "referenced by one or more, but not all other embodiments," unless otherwise specifically emphasized.
[0028] This application relates to human posture detection. As used herein, the term "human" refers to the most numerous and widely distributed primates on Earth. The human body includes the head, neck, torso, arms, hands, legs, and feet. The term "posture" (or "gesture") refers to a person's standing, sitting, and lying postures, etc. The term "judgment" refers to the determination of the state (e.g., posture) of a specific object (e.g., a person) based on calculations based on relevant data (e.g., including images of people). The term "mobile machine" refers to a machine capable of moving within its environment, such as a mobile robot or vehicle. The term "sensor" refers to a device, module, machine, or subsystem, such as an ambient light sensor and an image sensor (e.g., a camera), designed to detect events or changes in its environment and transmit information to other electronic devices (e.g., a processor). Figure 1 These are schematic diagrams illustrating scenarios where a mobile machine 100 is used to detect human posture in some embodiments of this application. For example... Figure 1 As shown, a mobile machine 100, navigating its environment (e.g., a room), detects the posture of a person (i.e., user U). The mobile machine 100 is a mobile robot (e.g., a mobile assistive robot) that includes a camera C and wheels E. The camera C can move in a forward direction D, which is a straight line from the mobile machine 100. f The camera is positioned so that the lens of camera C faces directly forward in the direction D. f Camera C has a camera coordinate system, and the coordinates of the moving machine 100 are consistent with those of camera C. In the camera coordinate system, the x-axis is parallel to the horizon, the y-axis is perpendicular to the horizon, and the z-axis is parallel to the forward direction D. f Consistent. It should be noted that mobile machine 100 is just one example of a mobile machine. Mobile machine 100 may have more, fewer, or different parts than the above or below (e.g., having legs instead of wheels E), or may have different part configurations or arrangements (e.g., placing camera C on top of mobile machine 100). In other embodiments, mobile machine 100 may be another type of mobile machine, such as a vehicle.
[0029] Figure 2 Is using Figure 1A schematic diagram of a mobile machine 100 using camera C to detect human posture. 1. Detecting human posture. The field of view (FOV) V of camera C covers user U, the blanket Q covering user U, and the bench B where user U is sitting. The height of camera C on mobile machine 100 (e.g., 1 meter) can be changed according to actual needs (e.g., the greater the height, the larger the field of view V; the smaller the height, the smaller the field of view V). The pitch angle of camera C relative to the ground F can also be changed according to actual needs (e.g., the larger the pitch angle, the closer the field of view V; the smaller the pitch angle, the farther the field of view V). Based on the height and pitch angle of camera C, the relative position of user U (and / or bench B) near mobile machine 100 can be obtained, and thus the posture of user U can be determined.
[0030] In some embodiments, the mobile machine 100 can be navigated in its environment while avoiding hazardous situations such as collisions and unsafe conditions (e.g., falls, extreme temperatures, radiation, and exposure). In this indoor navigation, the mobile machine 100 is navigated from a starting point (e.g., the initial location of the mobile machine 100) to a destination (e.g., the location of the navigation target specified by user U or the navigation / operating system of the mobile machine 100), and obstacles (e.g., walls, furniture, people, pets, and litter) can be avoided to prevent the aforementioned hazardous situations. In navigation, target finding (e.g., person discovery or user identification) can be considered to help locate user U, and the avoidance of occlusions (e.g., by furniture and fabric) can also be considered to improve the efficiency of human posture judgment. The trajectory (e.g., trajectory T) for the mobile machine 100 to move from the starting point to the destination must be planned so that the mobile machine 100 moves according to the trajectory. The trajectory includes a series of poses (e.g., pose S of trajectory T). n-1 -S n In some embodiments, to enable navigation of the mobile machine 100, it is necessary to construct an environmental map, which may require (e.g., using an IMU 1331) determining the current position of the mobile machine 100 in the environment, and a trajectory can be planned based on the constructed map and the determined current position of the mobile machine 100. Desired attitude S d This is the last of the sequence of poses S in trajectory T (only a portion is shown in the figure), i.e., the endpoint of trajectory T. Trajectory T is planned based on, for example, the shortest path to user U in a constructed map. Additionally, collision avoidance for obstacles in the constructed map (e.g., walls and furniture) or obstacles detected in real-time (e.g., humans and pets) can be considered during planning to ensure accurate and safe navigation of the mobile machine 100. It is important to note that the start and destination refer only to the position of the mobile machine 100, not the actual beginning and end of trajectory T (the actual beginning and end of trajectory T should each be a pose). Furthermore, Figure 1The trajectory T shown is only a part of the planned trajectory T.
[0031] In some embodiments, navigation of the mobile machine 100 can be initiated by the mobile machine 100 itself (e.g., a control interface on the mobile machine 100) or by a navigation request provided by a remote control, smartphone, tablet, laptop, desktop computer, or other electronic device. The mobile machine 100 and the control device can communicate via a network, which may include, for example, the Internet, an intranet, an extranet, a local area network (LAN), a wide area network (WAN), a wired network, a wireless network (e.g., Wi-Fi, Bluetooth, and mobile networks), or any suitable network, or any combination of two or more such networks.
[0032] Figure 3 This is an explanation Figure 1A schematic block diagram of a mobile machine 100. The mobile machine 100 may include a processing unit 110, a storage unit 120, and a control unit 130 that communicate via one or more communication buses or signal lines L. It should be noted that the mobile machine 100 is only one example of a mobile machine. The mobile machine 100 may have more or fewer components (e.g., units, sub-units, and modules) than shown above or below, may combine two or more components, or may have different component configurations or arrangements. The processing unit 110 executes various (sets of) instructions stored in the storage unit 120. These instructions may be in the form of software programs to perform various functions of the mobile machine 100 and process related data. The processing unit 110 may include one or more processors (e.g., a central processing unit). The storage unit 120 may include one or more memories (e.g., high-speed random access memory (RAM) and non-transitory memory), one or more memory controllers, and one or more non-transitory computer-readable storage media (e.g., solid-state drives (SSDs) or hard disks). Control unit 130 may include various controllers (e.g., camera controller, display controller, and physical button controller) and peripheral interfaces for coupling input / output peripherals of mobile machine 100 to processing unit 110 and storage unit 120, such as external ports (e.g., USB), wireless communication circuitry (e.g., RF communication circuitry), audio circuitry (e.g., speaker circuitry), and sensors (e.g., inertial measurement unit (IMU)). In some embodiments, storage unit 120 may include navigation module 121 for implementing navigation functions (e.g., map building and trajectory planning) related to navigation (and trajectory planning) of mobile machine 100, which may be stored in one or more memories (and one or more non-transitory computer-readable storage media).
[0033] The navigation module 121 in the storage unit 120 of the mobile device 100 can be a software module (of the operating system of the mobile device 100), which has instructions I for implementing navigation of the mobile device 100. n (For example, instructions for actuating the wheels E of the mobile machine 100 to move the mobile machine 100), a map builder 1211, and a trajectory planner 1212. The map builder 1211 may have instructions I for building a map for the mobile machine 100. b The software module. The trajectory planner 1212 may be a software module with instructions for planning trajectories for the mobile machine 100. p The software module. The trajectory planner 1212 may include a global trajectory planner for planning a global trajectory (e.g., trajectory T) for the mobile machine 100, and a local trajectory planner for planning a local trajectory (e.g., trajectory T) for the mobile machine 100. Figure 1The global trajectory planner is a local trajectory planner (a portion of the trajectory T in the map). The global trajectory planner can be, for example, a trajectory planner based on Dijkstra's algorithm, which plans the global trajectory based on a map constructed by the map builder 1211 using methods such as simultaneous localization and mapping (SLAM). The local trajectory planner can be a trajectory planner based on the TEB (timed elastic band) algorithm, which plans the local trajectory based on the global trajectory and other data collected by the mobile machine 100. For example, images can be acquired by the camera C of the mobile machine 100, and the acquired images can be analyzed to identify obstacles. The identified obstacles can then be used as a reference to plan the local trajectory, and the mobile machine 100 can be moved according to the planned local trajectory to avoid obstacles.
[0034] Map builder 1211 and trajectory planner 1212 can be respectively connected to instructions I for implementing navigation of the mobile machine 100. n Separate submodules or other submodules of navigation module 121, or instruction I n As part of the process unit 110, the trajectory planner 1212 may also have data (e.g., input / output data and temporary data) related to the trajectory planning of the mobile machine 100, which may be stored in one or more memories and accessed by the processing unit 110. In some embodiments, each trajectory planner 1212 may be a module in the storage unit 120 that is separate from the navigation module 121.
[0035] In some embodiments, instruction I n This may include instructions for implementing collision avoidance (e.g., obstacle detection and trajectory replanning) of the mobile machine 100. Furthermore, the global trajectory planner can replan the global trajectory (i.e., plan a new global trajectory) in response to situations where, for example, the original global trajectory is blocked (e.g., by an unexpected obstacle) or insufficient to avoid collisions (e.g., unable to avoid detected obstacles at the time of adoption). In other embodiments, the navigation module 121 may be a navigation unit that communicates with the processing unit 110, storage unit 120, and control unit 130 via one or more communication buses or signal lines L, and may also include one or more memories (e.g., high-speed random access memory (RAM) and non-temporary memory) for storing instructions I. n A map builder 1211 and a trajectory planner 1212; and one or more processors (e.g., microprocessors (MPUs) and microcontrollers (MCUs)) for executing stored instructions I n I b and I p To enable navigation of the mobile machine 100.
[0036] The mobile machine 100 may further include a communication subunit 131 and an actuation subunit 132. The communication subunit 131 and the actuation subunit 132 communicate with the control unit 130 via one or more communication buses or signal lines. These one or more communication buses or signal lines may be the same as or at least partially different from the aforementioned one or more communication buses or signal lines L. The communication subunit 131 is coupled to a communication interface of the mobile machine 100, such as a network interface 1311 for the mobile machine 100 to communicate with a control device via a network, an I / O interface 1312 (e.g., a physical button), etc. The actuation subunit 132 is coupled to a component / device for realizing the movement of the mobile machine 100, such as a motor 1321 driving the wheels E and / or joints of the mobile machine 100. The communication subunit 131 may include a controller for the aforementioned communication interface of the mobile machine 100, and the actuation subunit 132 may include a controller for the aforementioned component / device for realizing the movement of the mobile machine 100. In other embodiments, the communication subunit 131 and / or the actuation subunit 132 may simply be abstract components used to represent the logical relationships between components of the mobile machine 100.
[0037] The mobile machine 100 may also include a sensor subunit 133. The sensor subunit 133 may include a set of sensors and associated controllers, such as a camera C and an IMU 1331 (or an accelerometer and gyroscope), for detecting its environment to enable navigation. The camera C may be a ranging camera that generates ranging images. The ranging images include color data representing the colors of pixels in the image and depth data representing the distance to objects in the scene within the image. In some embodiments, the camera C is an RGB-D camera that generates RGB-D image pairs. Each image pair includes an RGB image and a depth image, wherein the RGB image includes pixels represented as red, green, and blue, respectively, while the depth image includes pixels, each pixel having a value representing the distance to a scene object (e.g., a person or furniture). The sensor subunit 133 communicates with the control unit 130 via one or more communication buses or signal lines, which may be the same as or at least partially different from the one or more communication buses or signal lines L described above. In other embodiments, when the navigation module 121 is the navigation unit described above, the sensor subunit 133 can communicate with the navigation unit through one or more communication buses or signal lines, which may be the same as or at least partially different from the one or more communication buses or signal lines L described above. Furthermore, the sensor subunit 133 may simply be an abstract component used to represent the logical relationships between the components of the mobile machine 100.
[0038] In some embodiments, the map builder 1211, trajectory planner 1212, sensor subunit 133, and motor 1321 (as well as the wheels E and / or joints of the mobile machine 100 connected to the motor 1321) together constitute a (navigation) system to realize map building, (global and local) trajectory planning, and motor drive to achieve navigation of the mobile machine 100. Furthermore, Figure 3 The various components shown can be implemented in hardware, software, or a combination of hardware and software. Two or more of the processing unit 110, storage unit 120, control unit 130, navigation module 121, and other units / subunits / modules can be implemented on a single chip or circuit. In other embodiments, at least some of them can be implemented on separate chips or circuits.
[0039] Figure 4 Is using Figure 1 A schematic block diagram illustrating an example of a mobile machine 100 detecting human posture. In some embodiments, instructions (sets) I corresponding to a human posture determination method are used, for example. n The navigation module 121 is stored in the storage unit 120, and the stored instructions I are executed by the processing unit 110. n The human posture determination method is implemented in mobile device 100, and then mobile device 100 can use camera C to detect and determine the posture of user U. The human posture determination method can be executed in response to a request from, for example, mobile device 100 itself or a control device (navigation / operating system) to detect the posture of user U, and can then be re-executed, for example, at predetermined time intervals (e.g., 1 second), to detect changes in the posture of user U. According to this human posture determination method, processing unit 110 can acquire RGB-D image pairs G(…) through camera C. Figure 4 (frame 410). Camera C captures image pairs G, where each image pair G includes an RGB image G. r and depth image G d RGB image G r Includes color information used to represent the colors that make up the image, depth image G d This includes depth information representing the distance to scene objects (e.g., user U or bench B) in the image. In some embodiments, multiple image pairs G may be acquired so that one image pair G (e.g., satisfying a certain quality requirement) can be selected for use.
[0040] Processing unit 110 can further detect RGB image G r Estimated skeleton N of user U in (see Figure 6 Key point P () Figure 4 (Frame 420). Can recognize RGB image G. rThe keypoints P in the image are used to obtain the 2D (two-dimensional) position of keypoint P on the estimated skeleton N of user U. The estimated skeleton N is a pseudo-human skeleton used to determine human posture (e.g., standing, sitting, and lying down), which has a set of predefined keypoints P. Each predefined keypoint represents a joint (e.g., knee) or an organ (e.g., ear). In some embodiments, the predefined keypoints P may include 18 keypoints, namely two eye keypoints, one nose keypoint, two ear keypoints, one neck keypoint, and two shoulder keypoints P. s (See Figure 6 Two key elbow points, two key hand points, and two key hip points (P) h (See Figure 6 ), two key points of the knee P k (See Figure 6 ), and two foot keypoints. These two eye keypoints, the nose keypoint, and the two ear keypoints are also called head keypoints P. d (Not shown in the figure). The processing unit 110 can further process the depth image G. d Calculate the location of the detected key point P ( Figure 4 (Boundary 430). The depth image G can be... d The depth information corresponding to key point P is combined with the obtained 2D position of key point P to obtain the 3D (three-dimensional) position of key point P.
[0041] Processing unit 110 can further select a feature map M from a set of predefined feature maps M based on the detected key point P in the predefined key point P. Figure 4 (Block 440). A set of predefined feature maps M can be pre-stored in storage unit 120. Each feature map M (e.g., Figure 7 Features Figure 1 ) is a custom graph that includes features corresponding to distinguishing poses (e.g. Figure 7 The values of the interior angles and body proportions in the feature map (M) are used to classify the detected keypoints P, thereby achieving pose differentiation. Clear pose segmentation in the feature map M can be achieved by using the maximum margin between each pose region in the feature map M and the target region. Figure 5A and 5BThis is a flowchart of the pose determination process based on detected keypoint P. In the pose determination process, the pose of user U is determined based on the detected keypoint P. Besides classifying the detected keypoint P using feature map M (the process of classifying poses using feature map M on keypoint P in individual image pairs G is called a "classifier") (steps 441-446, 452-453, and 462-463) when all or a specific portion of predefined keypoints are detected, this pose determination process also distinguishes poses when not all or a specific portion of predefined keypoints are detected, but the head keypoint P has been detected. d In this case, the detected head keypoint P is also used. d To distinguish postures (steps 470-490). This posture judgment process can be combined with the human posture judgment method described above.
[0042] In some embodiments, in order to select feature map M ( Figure 4 (in box 440), in step 441, it is determined whether the detected keypoint P includes at least one shoulder keypoint P from the predefined keypoints P. s At least one hip key point P h and at least one knee key point P k If the detected keypoint P includes at least one shoulder keypoint P s At least one hip key point P h and at least one knee key point P k If the condition is met, then step 442 is executed; otherwise, the method terminates. In step 442, features are selected from a set of predefined feature maps M. Figure 2 . Figure 9 yes Figure 8 Scene characteristics Figure 2 A schematic diagram. Features Figure 2 This involves two features: body proportion and upper body angle, which include threshold curves β and γ. Threshold curve β is a multinormal curve used to distinguish between standing and sitting postures, while threshold curve γ is a multinormal curve used to distinguish between lying, standing, and sitting postures. Since interior angles are no longer effective in this case, and lying and standing / sitting postures overlap in body proportion values, upper body angle is introduced in feature map 2. It is important to note that, to improve the efficiency of posture judgment, the features... Figure 2 The feature values displayed are also normalized values with mean and scale. After step 442, steps 451 and 461 are executed sequentially to determine the user U's pose. In step 443, it is determined whether the determined pose is a lying position. If not, step 444 is executed; otherwise, the method ends.
[0043] In step 444, it is determined whether the detected key point P includes all predefined key points P. Figure 6 yes Figure 5A and 5B A schematic diagram of scenario 1 for pose determination. In scenario 1, all predefined keypoints P are detected, and Figure 5A and 5B Steps 445, 452, and 462 will be executed to determine the posture of user U. For example, due to Figure 6 In a scenario where user U stands upright in front of camera C on mobile machine 100, with no obstacles between them, all 18 predefined keypoints P mentioned above can be detected. It's important to note that the 5 head keypoints P... d (That is, the two eye keypoints, the nose keypoint, and the two ear keypoints) are not shown in the image. If the detected keypoint P includes all predefined keypoints P (e.g., the detected keypoint P includes all 18 predefined keypoints mentioned above), then step 445 is executed; otherwise, step 446 is executed. In step 445, features are selected from the set of predefined feature maps M. Figure 1 . Figure 7 yes Figure 6 Scene characteristics Figure 1 A schematic diagram. Features Figure 1 For the two features of interior angle and body proportion, a threshold curve α is included. The threshold curve α is a multinormal curve used to distinguish between standing and sitting postures. Because lying down is much more likely to be occluded than standing or sitting postures, only standing and sitting postures are considered in Scenario 1 where all predefined keypoints P are detected. It is important to note that, to improve the efficiency of pose determination, the features... Figure 1 The feature values displayed are normalized values with mean and scale. In step 446, it is determined whether the detected keypoint P includes at least one shoulder keypoint P from the predefined keypoints P. s At least one hip key point P h and at least one knee key point P k .
[0044] Figure 8 yes Figure 5A and 5B The diagram illustrates scenario 2 for pose determination. In scenario 2, one or two shoulder keypoints P are detected. s One or two hip key points P h One or two key points P of the knee k Instead of all predefined keypoints P, and Figure 5A and 5B Steps 453 and 463 will be executed to determine the pose of user U. For example, due to Figure 8In the image, user U is sitting in a chair H in front of camera C of mobile machine 100, and a table T partially obscures part of user U's body. Only two of the 18 predefined keypoints P mentioned above are shoulder keypoints P. s Two key hip points P h and two knee key points P k (with head key point P) d Keypoints at the shoulder, hip, or knee can be detected. In one embodiment, if only one keypoint at the shoulder, hip, or knee is detected, the locations of undetected keypoints can be calculated based on the locations of the detected keypoints, for example, using the locations of the detected keypoints (and other detected keypoints). It should be noted that there are 5 head keypoints P. d (i.e., the two eye keypoints, the nose keypoint, and the two ear keypoints) are not shown in the diagram. If the detected keypoint P includes at least one shoulder keypoint P... s At least one hip key point P h and at least one knee key point P k If the condition is met, proceed to step 453; otherwise, proceed to step 470.
[0045] Processing unit 110 can further obtain two human features (i.e., feature 1 and feature 2) corresponding to the selected feature map M based on the position of the detected key point P. Figure 4 (box 450). In some embodiments, in order to obtain the feature corresponding to Figure 2 Features, in Figure 5A and 5B In step 451, the detected upper body angle of user U is obtained based on the position of the detected key point P (i.e., feature 2); in order to obtain the corresponding feature Figure 1 The two features are used to obtain the detected interior angle A of user U's body in step 452 based on the position of the detected key point P. i (i.e., feature 1, see) Figure 10 ) and the detected body proportion R (i.e., feature 2, see Figure 10 In order to obtain the corresponding features Figure 2 In step 453, based on the location of the detected key point P, the detected body proportion R of the detected user U is obtained (i.e., feature 1).
[0046] Figure 10 Is Figure 6 This is an illustration of an example of obtaining the detected interior angle and the detected body proportion in a scenario. In some embodiments, in order to obtain the detected interior angle A... i Based on the detected body proportion R (i.e., feature 1) and the detected body proportion R (i.e., feature 2), the processing unit 110 can acquire two hip key points P from the detected key points P.h The middle position p between h (See Figure 6 () Figure 10 (Box 4521). Processing unit 110 can further be based on the intermediate position p. h And two shoulder keypoints P among the detected keypoints s To obtain the upper body plane L position u (See Figure 6 () Figure 10 (Box 4522). Processing unit 110 can further be based on the intermediate position p. h The lower body plane L1 is obtained by using the positions of two knee keypoints among the detected keypoints. Figure 10 (Frame 4523). Processing unit 110 can further obtain the angle between the upper body plane and the lower body plane as the interior angle A. i ( Figure 10 (frame 4524). In one embodiment, the upper body plane L u The normal vector is up The plane of the lower body, L1, is the plane with normal vector L1. low plane, normal vector up With normal vector low The angle between them, taken as interior angle A i The processing unit 110 can further obtain the lower body height. h low (See Figure 6 ) and upper body height h up (See Figure 6 The ratio between ) to serve as the body proportion R ( Figure 10 (frame 4525). In one embodiment, the upper body height h up = p s .y- p h .y and lower body height h low = p h .y- p k .y, where p s There are two key points on the shoulders, P. s The middle position, p s .y is p s y-coordinate, ph .y is p h y-coordinate, p k The two key points of the knees, P k The middle position, p k .y is p k The y-coordinate.
[0047] Figure 11 Is Figure 8 In the scene (i.e., scene 2), obtain the detected body proportion R and the detected upper body angle A. u A schematic diagram of an example. In some embodiments, in order to obtain the detected body proportion R (i.e., feature 1) and the detected upper body angle A... u (i.e., feature 2), the processing unit 110 can acquire two hip key points P from the detected key points P. h The middle position between the positions p h ( Figure 11 (Block 4531). Processing unit 110 can further obtain lower body height. h low with upper body height h up The ratio between them, used as the body proportion R ( Figure 11 (box 4532), where h up = p s .y- p h .y and h low = p h .y - p k .y, p s There are two key points on the shoulders, P. s The middle position, p s .y is p s y-coordinate, p h .y is p h y-coordinate, p k The two key points of the knees, P k The middle position between the positions, p k .y is p kThe y-coordinate. Processing unit 110 can further obtain the intermediate position. p h and the middle position p s Vectors between h s The angle between the upper body and the unit vector in the y-direction is taken as the upper body angle A. u ( Figure 11 (4533)
[0048] Processing unit 110 can further determine the pose of user U based on two features (i.e., feature 1 and feature 2) in the selected feature map M. Figure 4 (box 460). In some embodiments, in order to determine the feature Figure 2 The upper body angle is used to determine the user U's posture (only reclining posture is distinguished here). In step 461, the user U's posture A is determined based on the detected upper body angle. u (i.e., feature 2). In order to base on feature Figure 1 The two features (i.e., the detected interior angle A) i The pose of user U is determined based on the detected body proportion R, and in step 462, the pose is determined based on the detected interior angle A. i The user U's pose is determined by the location of the intersection point between feature 1 (i.e., feature 1) and the detected body proportion R (i.e., feature 2). For example, for Figure 6 The standing user U, due to the detected interior angle A i The location of the intersection of (e.g., 0.613) and the detected body proportion R (e.g., 0.84) (both normalized values) is in Figure 7 Features Figure 1 At a point above the threshold curve α, user U will be judged as standing. To determine the posture based on features... Figure 2 The two features (i.e., the detected body proportion R and the detected upper body angle A) u Determine the pose of user U in step 463 based on features. Figure 1 The user U's pose is determined by the location of the intersection of the detected body proportions R (i.e., feature 1). For example, for... Figure 8 The seated user U, due to the detected body proportion R (e.g., -0.14) and the detected upper body angle A u The intersection of (e.g., 0.44) (all standardized values) is located to the left of the threshold curve β and below the threshold curve γ in feature map 2. User U will be judged as sitting.
[0049] exist Figure 5A and 5B During the pose determination process, in step 470, it is determined whether the detected keypoint P includes multiple head keypoints P in the predefined keypoint P.d . Figure 12 yes Figure 5A and 5B The diagram illustrates scenario 3 for pose determination. In scenario 3, more than half of user U's body is obscured or outside the field of view V of camera C, but (partial or complete) head keypoint P is detected. d , Figure 5A and 5B Steps 480 and 490 will be executed to determine the pose of user U. For example, since user U is lying on bench B in front of camera C of mobile machine 100, and part of the user's body is obscured by a blanket Q, only the head keypoint P among the 18 predefined keypoints P mentioned above is considered. d (For example, two or more of the two eye keypoints, the nose keypoint, and the two ear keypoints) (and two shoulder keypoints P) s Two elbow keypoints and two hand keypoints can be detected. It is important to note that five head keypoints (P...) are also detected. d (That is, the two eye keypoints, the nose keypoint, and the two ear keypoints) are not shown in the image. If the detected keypoint P includes multiple head keypoints P... d If the condition is met, proceed to step 480; otherwise, the method will terminate. In step 480, based on the head keypoint P... d Calculate head height H head Head height H head This is the height of the head (including the eyes) relative to the ground. Since the relative height and angle of camera C relative to the ground are known, the head height H can be calculated based on the head's position in the camera coordinate system. head For example, when user U is sitting in chair C, the head height H head It is the height of the head relative to the floor F (see Figure 8 When user U stands on floor F, head height H head It is also the height of the head relative to the floor F (see Figure 6). In step 490, by adjusting the head height H head A first head height threshold (e.g., 120 cm) and a second head height threshold (e.g., 50 cm) are used to distinguish between standing and sitting postures, and to distinguish between sitting and lying postures. At head height H... head If the head height is greater than the first head height threshold, the user is determined to be in a standing posture (U); if the head height is greater than the first head height threshold (H), the user is determined to be in a standing posture (U). head If the head height falls between the first and second head height thresholds, the user U is determined to be in a seated posture. If the head height H is between the first and second head height thresholds, the user U is determined to be in a seated posture. head If the height is less than the second head height threshold, the user is determined to be in a lying position (U). For example, for Figure 12 The user U, who is lying down, has a head height of H. headIf the height of the first head is less than the second head height threshold (e.g., 30 cm), then user U will be judged to be in a lying position. In some embodiments, the first and second head height thresholds can be changed according to, for example, the height of user U. In other embodiments, all steps of the three scenarios described above (i.e., steps 445, 452, and 462 for scenario 1, steps 444, 453, and 463 for scenario 2, and steps 480 and 490 for scenario 3) can be performed without judgment steps (i.e., steps 441, 446, and 470). Then, the judgment results of the three scenarios corresponding to each frame (i.e., image pair G) can be combined with the maximum vote to determine the posture of user U.
[0050] In other embodiments, in human pose determination methods (incorporating pose determination processes), since humans do not frequently change their pose within a millisecond timescale, a time window (e.g., 1 second) can be added to filter out invalid results, achieving more accurate and robust pose determination. For example, for multiple adjacent frames within a time window, the determination results of the three scenarios with the most votes corresponding to the three scenarios mentioned above are combined to make a human pose decision. Then, the pose of user U is determined when all decisions within the time window are the same. It should be noted that the size of the time window (representing how many adjacent frames will be accumulated) can be defined according to actual needs. For example, the size of the time window can be dynamically adjusted based on the closeness of the current feature value to the corresponding threshold curve; the closer it is to the threshold curve, the smaller its size, and vice versa.
[0051] This human pose determination method can adapt to different sets of detected keypoints because it uses a predefined feature map describing the human pose based on the detected keypoints. By using unique predefined features and feature maps, it can accurately determine human pose. This human pose determination method can be implemented in real time, requires minimal computational resources, and is cost-effective because it only requires one RGB-D camera instead of multiple sensors for detection. When the mobile machine implementing this human pose determination method is a walking robot, it can determine the person's pose and choose an appropriate way to interact with them. For example, when a person is determined to be an elderly person lying down, the mobile robot can ask them to sit down first before providing further assistance.
[0052] Those skilled in the art will understand that all or part of the methods in the above embodiments can be implemented by one or more computer programs instructing the relevant hardware. Furthermore, one or more programs can be stored in a non-transitory computer-readable storage medium. When one or more programs are executed, all or part of the corresponding methods in the above embodiments are performed. Any reference to storage, memory, database, or other media can include non-transitory and / or transient memory. Non-transitory memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, solid-state drives (SSDs), etc. Volatile memory can include random access memory (RAM), external cache memory, etc.
[0053] Processing unit 110 (and the processor described above) may include a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gates, transistor logic devices, and discrete hardware components. The general-purpose processor may be a microprocessor or any conventional processor. Storage unit 120 (and the memory described above) may include internal storage units such as hard disk drives and internal memory. Storage unit 120 may also include external storage devices such as plug-in hard disk drives, smart media cards (SMCs), secure digital cards (SD cards), and flash memory cards.
[0054] The exemplary units / modules and methods / steps described in the embodiments can be implemented by software, hardware, or a combination of software and hardware. Whether these functions are implemented by software or hardware depends on the specific application and design constraints of the technical solution. The above-mentioned human lying posture detection method and mobile machine 100 can be implemented in other ways. For example, the division of units / modules is only a logical functional division. In actual implementation, other division methods can be used, that is, multiple units / modules can be combined or integrated into another system, or certain features can be ignored or not executed. In addition, the above-mentioned mutual coupling / connection can be direct coupling / connection or communication connection, or indirect coupling / connection or communication connection through some interfaces / devices, or it can be electrical, mechanical or other forms.
[0055] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the technical solutions of the present invention. Although the present invention has been described in detail with reference to the above embodiments, the technical solutions in the above embodiments can still be modified, or some technical features can be equivalently replaced, so that these modifications or replacements do not cause the substance of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the scope of protection of the present invention.
Claims
1. A method for judging human posture, characterized in that, include: One or more range-measuring images are acquired using a range-measuring camera, wherein the one or more range-measuring images include color data and depth data; Detect multiple key points of the estimated skeleton of the human body in the color data, and calculate the positions of the detected key points based on the depth data, wherein the estimated skeleton has a set of predefined key points; Based on the key points detected in the set of predefined key points, feature maps are selected from a set of predefined feature maps; Based on the location of the detected key points, two features of the human body corresponding to the selected feature map are obtained; as well as The posture of the human body is determined based on the two features in the selected feature map; The step of selecting feature maps from a set of predefined feature maps based on the detected key points in the set of predefined key points includes: When the detected key points include at least one shoulder key point, at least one hip key point and at least one knee key point among the predefined key points, a second feature map with the two features of body proportion and upper body angle is selected from the set of predefined feature maps. The second feature map includes a first threshold curve for distinguishing between standing and sitting postures and a second threshold curve for distinguishing between lying, standing and sitting postures. When the detected posture is not a lying posture and the detected key points include all the predefined key points, a first feature map with the two features of interior angle and body proportion is selected from the set of predefined feature maps. The first feature map includes a threshold curve for distinguishing between standing and sitting postures.
2. The method as described in claim 1, characterized in that, When the selected feature map is the first feature map, the step of obtaining two features of the human body corresponding to the selected feature map based on the position of the detected key points includes: In response to the detected keypoints including all the predefined keypoints, the detected interior angles and body proportions of the human body are obtained based on the positions of the detected keypoints; and Determining the human body's posture based on the two features in the selected feature map includes: In response to the detected key points, which include all the predefined key points, the pose of the human body is determined based on the position of the intersection of the detected interior angle and the detected body proportion in the first feature map.
3. The method as described in claim 2, characterized in that, Based on the location of these key points, the inner angle and body proportions are obtained, including: Obtain the midpoint of this location between two hip keypoints among the detected keypoints. p h ; Based on this intermediate position p h The upper body plane is obtained by taking the positions of two shoulder key points among the detected key points; Based on this intermediate position p h The lower body plane is obtained by taking the positions of two knee key points among the detected key points; Obtain the angle between the upper body plane and the lower body plane and use it as the interior angle; and Get lower body height h low and upper body height h up The ratio between them is used as the body proportion, where h up = p s .y - p h .y and h low = p h .y - p k .y, p s It is the midpoint between these two acromion points. p s .y is p s y-coordinate, p h .y is p h y-coordinate, p k It is the midpoint between these two key knee points. p k .y is p k The y-coordinate; and Determining the human body's posture based on the location of the intersection of the detected interior angle and the detected body proportion in the first feature map includes: The posture of the human body is determined based on the position of the intersection of the interior angle and the body proportion in the first feature map relative to the threshold curve used to distinguish between the standing posture and the sitting posture.
4. The method as described in claim 1, characterized in that, When the selected feature map is the second feature map, the step of obtaining two features of the human body corresponding to the selected feature map based on the position of the detected key points includes: In response to the detected key points, including at least one shoulder key point, at least one hip key point, and at least one knee key point among these predefined key points, the detected body proportions and detected upper body angles are obtained based on the positions of the detected key points; and Determining the human body's posture based on the two features in the selected feature map includes: In response to the detected key points, including at least one shoulder key point, at least one hip key point, and at least one knee key point among the predefined key points, the posture of the human body is determined based on the position of the intersection of the detected body proportion and the detected upper body angle in the second feature map.
5. The method as described in claim 4, characterized in that, Based on the detected locations of these key points, the body proportions and upper body angle are obtained, including: Obtain the midpoint of this location between two hip keypoints among the detected keypoints. p h ; Get lower body height h low and upper body height h up The ratio between them is used as the body proportion, where h up = p s .y- p h .y and h low = p h .y- p k .y, p s It is the midpoint between the two shoulder key points. p s .y is p s y-coordinate, p h .y is p h y-coordinate, p k It is the midpoint between the two knee joints, and p k .y is p k The y-coordinate; and Get the intermediate position p h and the middle position p s Vectors between hs unit vector in the y-direction The angle between them, and as the angle of the upper body; and Determining the posture of the human body based on the position of the intersection of the body proportions and the upper body angle in the second feature map includes: The posture of the human body is determined based on the position of the intersection of the body proportion and the upper body angle relative to the position of the first threshold curve used to distinguish the standing posture and the sitting posture and the second threshold curve used to distinguish the lying posture, the standing posture and the sitting posture in the second feature map.
6. The method as described in claim 5, characterized in that, The method further includes: In response to the detected keypoint including a shoulder keypoint, the position of the other of the two shoulder keypoints is obtained based on the position of the detected shoulder keypoint; In response to the detected keypoints including a hip keypoint, the location of the other of the two hip keypoints is obtained based on the location of the detected hip keypoint; and In response to the detected keypoint including a knee keypoint, the position of the other of the two knee keypoints is obtained based on the position of the detected knee keypoint.
7. The method as described in claim 1, characterized in that, The method further includes: In response to the detected keypoints, which include multiple head keypoints from among these predefined keypoints, the head height H is calculated based on these head keypoints. head ;as well as By adjusting the head height H head The posture of the human body is determined by comparing it with a first head height threshold used to distinguish between standing and sitting postures and a second head height threshold used to distinguish between sitting and lying postures.
8. A mobile machine, comprising: Rangefinder camera; One or more processors; as well as One or more memories storing one or more computer programs, which are executed by one or more processors, wherein the one or more computer programs include a plurality of instructions for: The ranging camera acquires one or more ranging images, wherein the one or more ranging images include color data and depth data; Detect multiple key points of the estimated skeleton of the human body in the color data, and calculate the positions of the detected key points based on the depth data, wherein the estimated skeleton has a set of predefined key points; Based on the key points detected in the set of predefined key points, feature maps are selected from a set of predefined feature maps; Based on the location of the detected key points, two features of the human body corresponding to the selected feature map are obtained; as well as The posture of the human body is determined based on the two features in the selected feature map; The step of selecting feature maps from a set of predefined feature maps based on the detected key points in the set of predefined key points includes: When the detected key points include at least one shoulder key point, at least one hip key point and at least one knee key point among the predefined key points, a second feature map with the two features of body proportion and upper body angle is selected from the set of predefined feature maps. The second feature map includes a first threshold curve for distinguishing between standing and sitting postures and a second threshold curve for distinguishing between lying, standing and sitting postures. When the detected posture is not a lying posture and the detected key points include all the predefined key points, a first feature map with the two features of interior angle and body proportion is selected from the set of predefined feature maps. The first feature map includes a threshold curve for distinguishing between standing and sitting postures.