Body part correction method and device based on physical calibration and key point recognition under monocular vision

By combining a monocular camera and a key point detection model with depth estimation, a physical height mapping relationship between body parts is established, which solves the problem of inaccurate positioning in existing technologies, achieves contactless, fast, and accurate positioning of body parts, and reduces equipment costs.

CN120604974APending Publication Date: 2025-09-09TAOYING (SHIJIAZHUANG) MEDICAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510556834.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies for body part positioning based on monocular vision cannot accurately guide X-ray equipment to automatically move to the target position, and methods that rely on camera posture and skeletal key points increase costs.

Method used

By capturing images with a monocular camera, a mapping relationship between pixel positions and actual physical positions is established. By combining key point detection and depth estimation models, the physical height of body parts is calculated and corrected, reducing dependence on depth cameras.

Benefits of technology

It realizes contactless and rapid measurement, improves the accuracy and stability of body part positioning, adapts to different body postures, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120604974A_ABST
    Figure CN120604974A_ABST
Patent Text Reader

Abstract

The invention discloses a body part correction method and device based on physical calibration and key point recognition under monocular vision, and relates to the field of medical instruments, and the method comprises the steps: building a mapping relation between a human body image pixel position and an actual physical position based on physical calibration; inputting the image into a key point detection model to detect the pixel position of the target part; calculating the physical height of the target part; inputting the image into a depth estimation model to estimate the depth information of the pixel position of the target part, and calculating the offset generated by the target part relative to the reference part; the physical heights of the human body under different depths are corrected according to the offset, the height of the corrected target part in the actual physical space is calculated, non-contact rapid measurement is achieved, body height positioning can be rapidly completed based on monocular image data without body contact, universality is high, and the method can adapt to height positioning under different postures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical imaging technology, and in particular to a body part correction method and device based on physical calibration and key point recognition under monocular vision. Background Art

[0002] X-ray scanning is an important auxiliary method for medical examinations. Traditional imaging equipment requires manual movement of the device to the target body part. By accurately positioning the patient's body part, the device can be automatically guided to the target body part, helping to improve examination efficiency.

[0003] The camera-based automatic positioning solution has the advantages of simple configuration, contactless, and reusable. Combined with physical calibration, it can achieve accurate positioning of body parts and further correct the positioning results at different distances through depth estimation.

[0004] At present, in order to realize the positioning of body parts, the following solutions are usually adopted: 1) Patent application number CN111611928A, titled "A Method for Measuring Height and Body Dimensions Based on Monocular Vision and Keypoint Recognition," uses a monocular image to detect key points using a human keypoint model and extracts body contours using a semantic segmentation model, thereby locating body parts. Furthermore, by matching the user with a template, the user's height calculation model parameters are adaptively updated. The updated height calculation model calculates the user's true height, and the obtained user key points and edge contours are combined to calculate the dimensions of each body part. However, this method suffers from the following drawbacks: Calculating body dimensions based on a trained height estimation model cannot directly determine the physical location of each body part, and cannot accurately guide the X-ray device to automatically move to the target location.

[0005] 2) Patent application number CN114022532A, entitled "Height Measurement Method, Height Measurement Device, and Terminal," includes an image of a target object and the camera's position when capturing the image; obtaining pixel coordinates of at least two skeletal key points of the target object in the image; obtaining the three-dimensional coordinates of at least two skeletal key points using a collision detection algorithm based on the camera's position and the pixel coordinates of the skeletal key points, combined with three-dimensional point cloud information acquired by a distance sensor; and determining the target object's height data based on the three-dimensional coordinates of the at least two skeletal key points. Its drawback is that the three-dimensional coordinates of the skeletal key points are obtained based on the camera's position and the pixel coordinates of the skeletal key points, combined with three-dimensional point cloud information acquired by a distance sensor, to determine the target object's height data. This solution relies on skeletal images and cannot locate body parts based on natural human images. It relies on camera position information and a distance sensor, increasing costs. Summary of the Invention

[0006] The purpose of the present invention is to provide a body part correction method and system based on physical calibration and key point recognition under monocular vision, which can realize contactless and rapid measurement. Based on monocular vision image data, body height positioning can be quickly completed without body contact. It has strong versatility and can adapt to height positioning in different body postures, reduces the number of depth cameras, and greatly reduces costs.

[0007] The present invention provides a body part correction method based on physical calibration and key point recognition under monocular vision, comprising the following steps: A monocular camera is used to capture a monocular image at the same position, and a mapping relationship between the pixel position of the monocular image and the actual physical position is established based on physical calibration; Inputting the captured monocular vision image into a key point detection model to detect the pixel position of the target part in the human body image; Mapping the detected pixel height to the physical height on the calibration plane and calculating the physical height of the target part based on the result of the physical calibration; Inputting the captured monocular visual image into a depth estimation model to estimate the depth information of the pixel position of the target part in the monocular visual image, and comparing the depth estimation value of the target part with the depth estimation value of the reference part to obtain the offset of the target part relative to the reference part; The physical height of the human body at different depths is corrected according to the offset, and the height of the corrected target part in the actual physical space is calculated.

[0008] Preferably, the method of shooting a monocular visual image at the same position by a monocular camera and establishing a mapping relationship between the pixel position of the monocular visual image and the actual physical position based on physical calibration includes: Setting the distance between the monocular camera and the target object to form a calibration plane and capturing a monocular vision image for physical calibration; Data points are selected from the calibration image according to equal pixel distances, and the vertical pixel positions of the data points in the calibration image are recorded to form a pixel position sequence. ; Read the actual physical height corresponding to the data point to form a physical position sequence ; Create an Array by polynomial fitting pixel to Array physical The mapping relationship is obtained to obtain the polynomial fitting coefficients (c0, c1, ... c 15 ); Among them, c0, c1, ... c 15 Represents the polynomial fit coefficients for physical calibration.

[0009] Preferably, the step of inputting the captured monocular visual image into a key point detection model to detect the pixel position of the target part in the human body image includes: Collect monocular vision images of a target object standing naturally facing the direction of the monocular camera; Inputting the monocular visual image into a key point detection model; The pixel locations of target parts including the nose, lower center of the neck, shoulders, elbows, wrists, hips, knees and ankles are output by the keypoint detection model.

[0010] Preferably, calculating the physical height of the target part includes: The physical height L corresponding to the target part is calculated using the polynomial fitting coefficient of the physical calibration physical : ; Among them, L pixel is the pixel height of the target part, , ,..., Represents the polynomial fit coefficients for physical calibration.

[0011] Preferably, the monocular visual image obtained by capture is input into a depth estimation model to estimate the depth information of the pixel position of the target part in the monocular visual image, and the depth estimation value of the target part is compared with the depth estimation value of the reference part to obtain the offset of the target part relative to the reference part, which includes: Collect the monocular vision image InputImg of the target object standing naturally facing the direction of the monocular camera; Inputting the monocular vision image InputImg into the depth estimation model; Output the depth estimation value OutputImg of the target part at the pixel position through the depth estimation model; Get the depth estimation value OutputImg[X target ,Y target ] and the depth estimation value OutputImg[X ankle ,Y ankle ]; The difference between the depth estimates of the target part and the reference part is calculated as the offset: .

[0012] Preferably, the correcting the physical height of the human body at different depths according to the offset and calculating the height of the corrected target part in the actual physical space includes: ; Among them, L physical_correct is the height of the corrected target part in the actual physical space, L p is the distance between the calibration plane and the camera lens during calibration, L c The target object's standing plane is the distance between the standing position and the camera lens, and Offset is the difference between the depth estimation values ​​of the target part and the reference part, that is, the offset.

[0013] The present invention also provides a body part correction device based on physical calibration and key point recognition under monocular vision, which implements the body part correction method based on physical calibration and key point recognition under monocular vision as described in an embodiment of the present invention. The device includes: A physical calibration module is used to capture a monocular visual image at the same position using a monocular camera and establish a mapping relationship between the pixel position of the monocular visual image and the actual physical position based on physical calibration; A key point detection module is used to input the monocular visual image obtained by shooting into a key point detection model to detect the pixel position of the target part in the human body image; A physical height calculation module is used to map the detected pixel height to the physical height on the calibration plane and calculate the physical height of the target part based on the result of the physical calibration; A depth estimation module is used to input the captured monocular visual image into a depth estimation model to estimate the depth information of the pixel position of the target part in the monocular visual image, and compare the depth estimation value of the target part with the depth estimation value of the reference part to obtain the offset of the target part relative to the reference part; The physical height correction module is used to correct the physical height of the human body at different depths according to the offset, and calculate the height of the corrected target part in the actual physical space.

[0014] Preferably, the physical calibration module includes: A parameter setting unit, used to set the distance between the monocular camera and the target object to form a calibration plane and capture a monocular vision image for physical calibration; A data point selection unit is used to select data points from the calibration image according to equal pixel distances, record the vertical pixel positions of the data points in the calibration image, and form a pixel position sequence. ; The physical height reading unit is used to read the actual physical height corresponding to the data point to form a physical position sequence ; Fitting unit, used to create an array through polynomial fitting pixel to Array physicalThe mapping relationship is obtained to obtain the polynomial fitting coefficients (c0, c1, ... c 15 ); Among them, c0, c1, ... c 15 Represents the polynomial fit coefficients for physical calibration.

[0015] The present invention also provides an electronic device, comprising: a memory for storing a processing program; A processor, which implements the body part correction method based on physical calibration and key point recognition under monocular vision as described in an embodiment of the present invention when executing the processing program.

[0016] The present invention also provides a readable storage medium having a processing program stored thereon. When the processing program is executed by a processor, the body part correction method based on physical calibration and key point recognition under monocular vision as described in an embodiment of the present invention is implemented.

[0017] With respect to the prior art, the present invention has the following beneficial effects: The body part correction method based on physical calibration and key point recognition under monocular vision provided by the present invention has contactless rapid measurement and can quickly complete body height positioning based on image data without physical contact; by establishing a mapping relationship between the camera image and the actual physical position through physical calibration, two-dimensional image information can be more accurately converted into three-dimensional spatial information, thereby improving the accuracy of posture estimation; using a key point detection model to locate body parts can better identify the target parts and enhance the robustness of the system; by considering individual differences such as height and body shape during physical calibration, posture estimation is more personalized, suitable for different groups of people, with strong versatility, and can adapt to height positioning under different body postures; introducing a depth estimation model to obtain depth information of the target part further enriches the data source that can be used for posture estimation, which helps to improve the estimation result; the method can work under the conditions of a monocular camera, has simple equipment and low cost, is easy to apply and promote in practice, that is, reduces the depth camera and greatly reduces the cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Schematic diagram of the steps of the body part correction method based on physical calibration and key point recognition under monocular vision according to the first embodiment of the present invention. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0020] Example 1 like Figure 1 As shown, the present invention provides a body part correction method based on physical calibration and key point recognition under monocular vision, comprising the following steps: S1: A monocular vision image is captured at the same position using a monocular camera, and a mapping relationship between the pixel position of the monocular vision image and the actual physical position is established based on physical calibration. The monocular vision image refers to an image captured from a single perspective, which only contains two-dimensional information and loses depth information.

[0021] S2: The monocular visual image obtained by shooting is input into the key point detection model to detect the pixel position of the target part in the human body image; based on physical calibration, the pixel position of the body part obtained by key point detection can be mapped to the physical height at the same shooting position.

[0022] S3: Mapping the detected pixel height to the physical height on the calibration plane and calculating the physical height of the target part according to the result of the physical calibration; S4: inputting the captured monocular visual image into a depth estimation model to estimate the depth information of the pixel position of the target part in the monocular visual image, and comparing the depth estimation value of the target part with the depth estimation value of the reference part to obtain the offset of the target part relative to the reference part; S5: The physical height of the human body at different depths is corrected based on the offset, and the corrected height of the target part in actual physical space is calculated. Utilizing imaging principles and depth estimation, the physical height of the mapped body part can be corrected at any shooting position, without any restrictions on the shooting position. A series of pixel positions and corresponding physical heights are collected, and a polynomial fitting is performed to obtain a mapping relationship between pixel positions and physical heights. Although the image captured by the monocular camera loses depth information, the keypoint detection model can be used to obtain the pixel positions of the target human part in the image. During physical calibration, a mapping relationship between the pixel positions of the monocular image and the actual physical positions is established. The pixel heights are then mapped to physical heights and corrected using the depth information and offsets obtained from the depth estimation model. This multi-information fusion approach comprehensively considers the two-dimensional pixel positions, actual physical height, and depth information of the image. Compared to using any single information for pose estimation, it can more comprehensively and accurately recover the human body's pose in actual three-dimensional space. For example, in complex scenes, relying solely on pixel positions may lead to pose estimation errors due to factors such as viewing angle. Incorporating depth information can effectively correct such errors, thereby improving the accuracy of pose estimation.

[0023] The above scheme establishes a mapping relationship between pixel positions and actual physical positions through physical calibration, ensuring the accurate correspondence between two-dimensional image information and three-dimensional physical space, and uses a depth estimation model to obtain the depth information of the target part and correct the physical height at different depths, effectively improving the accuracy of height measurement. The depth estimation model supplements the lost depth information in the monocular vision image, so that the two-dimensional image can reflect the position relationship in the three-dimensional space. The key point detection model is used to accurately identify the target part in the human body image, improving the stability and reliability of the measurement. By comparing the depth estimation value of the target part with the reference part and calculating the offset, the adaptability of the measurement result to depth changes is further enhanced. Non-contact fast measurement, based on monocular vision image data, can quickly complete body height positioning without body contact, has strong versatility, can adapt to height positioning under different body postures, reduces the number of depth cameras, and greatly reduces costs.

[0024] In one embodiment, in step S1, capturing a monocular visual image at the same position by a monocular camera, and establishing a mapping relationship between the pixel position of the monocular visual image and the actual physical position based on physical calibration includes: Parameter Setting: Set the distance between the monocular camera and the target object to form a calibration plane and capture a monocular vision image for physical calibration. For example, set the monocular camera's installation height to H, the calibration plane's distance from the camera lens to Lp, and the target object's standing distance from the camera lens to Lc. The calibration plane is a vertical plane with a known physical distance Lp, and an image of a vertical scale is captured. The purpose of forming a calibration plane by setting the distance and height between the monocular camera and the target object is to create a reference plane with a known geometric relationship. The actual physical height of the data points on this plane can be accurately obtained through measurement and other methods, providing basic data for establishing a mapping relationship between image pixel positions and actual physical positions.

[0025] Data point selection: select data points from the calibration image at equal pixel distances, and record the vertical pixel positions of the data points in the calibration image. For example, 15 data points are evenly selected from the bottom to the top of the image, and their vertical pixel positions are recorded sequentially to form a pixel position sequence. ; Selecting data points at equal pixel distances ensures the uniform distribution of data points on the image. By recording the vertical pixel positions of these data points to form a pixel position sequence, and reading the corresponding actual physical heights to form a physical position sequence, the dataset established in this way can more comprehensively and accurately reflect the correspondence between the pixel positions of the monocular vision image and the actual physical positions. This is also to ensure that the data points on the image can evenly cover the entire field of view, and more comprehensively reflect the correspondence between the pixel positions and physical positions of different areas of the image. For example, in a monocular vision image for human posture estimation, selecting data points at equal pixel distances from the head to the feet can cover the pixel information of various parts of the human body, so as to better establish the overall mapping relationship.

[0026] Physical height reading: read the actual physical height corresponding to the data point to form a physical position sequence ; For example, read the physical height of 15 selected data points from the ruler.

[0027] Polynomial fitting: Create an Array by polynomial fitting pixel to Array physical The mapping relationship is obtained to obtain the polynomial fitting coefficients (c0, c1, ... c 15 ), c0, c1, … c 15The polynomial fitting coefficients represent the physical calibration. Polynomial fitting can effectively adapt to these complex nonlinear relationships, resulting in a more accurate mapping. Compared to simple linear calibration methods, polynomial fitting can more accurately describe the geometric deformation characteristics of camera imaging, thereby improving the calibration accuracy from monocular image pixel positions to actual physical locations. For example, if the camera lens has a certain degree of distortion, polynomial fitting can better account for this distortion, making the calibration results closer to reality.

[0028] In one embodiment, the step S2 of inputting the captured monocular vision image into a key point detection model to detect the pixel position of the target part in the human body image and the step S3 of calculating the physical height of the target part include: Collect the monocular vision image InputImg of the target object standing naturally facing the direction of the monocular camera; Inputting the monocular visual image InputImg into the key point detection model; The pixel positions of the target parts are output through the key point detection model. The target parts include the nose, the center of the lower neck, shoulders, elbows, wrists, hips, knees and ankles. It can be understood that the purpose of collecting the monocular visual image InputImg of the target object standing naturally facing the direction of the monocular camera is to obtain a reference image containing complete posture information of the human body from the front. In this image, the relative position and distribution of various parts of the human body on the image plane can reflect its posture characteristics. The image InputImg is input into the key point detection model. The model recognizes the characteristic patterns of key parts of the human body in the image, such as the nose, the center of the lower neck, shoulders, elbows, wrists, hips, knees and ankles, through learning and analyzing the image, and outputs the pixel positions of these key parts.

[0029] The physical height L corresponding to the target part is calculated using the polynomial fitting coefficient of the physical calibration physical : ; Among them, L pixel is the pixel height of the target part, , ,..., The polynomial fitting coefficients represent the physical calibration. By properly applying the polynomial fitting coefficients, the solution accurately calculates the physical height of the target object, regardless of whether the image is captured from close or long distances, or when the camera angle deviates. This allows for effective pose estimation in complex real-world environments, such as crowded scenes and various indoor and outdoor settings, enhancing the solution's versatility and practicality.

[0030] The pixel positions of target parts output by the keypoint detection model are combined with the physically calibrated polynomial fitting coefficients to calculate the physical height of the target parts. This approach leverages the accurate mapping between image features and actual physical dimensions. This combination provides a more refined representation of the true size and proportional relationships of the human body in different postures. For example, as the human body changes posture, the height changes of different parts can be accurately reflected through the accurate pixel-to-physical height conversion, providing more precise scale information for pose estimation. Compared to pose estimation based solely on pixel information or simple proportional relationships, this approach can produce a more accurate pose description. For example, it provides more accurate data support for determining the amplitude of body movements and body tilt angles, making the pose estimation results closer to the real situation. By directly using the pixel position information output by the keypoint detection model and the known polynomial fitting coefficients for physical height calculation, the entire process is relatively simple and efficient, without the need for complex iterative algorithms or significant additional computing resources to determine the physical dimensions of various human parts.

[0031] In one embodiment, in step S4, the monocular visual image obtained by capture is input into a depth estimation model to estimate the depth information of the pixel position of the target part in the monocular visual image, and the depth estimation value of the target part is compared with the depth estimation value of the reference part to obtain the offset of the target part relative to the reference part, which includes: Collect the monocular vision image InputImg of the target object standing naturally facing the direction of the monocular camera; Inputting the monocular vision image InputImg into the depth estimation model; The depth estimation model outputs the depth estimation value OutputImg at the pixel position of the target part; the learning ability of the model is used to analyze the position information of the object corresponding to each pixel in the image in three-dimensional space. The depth estimation model is usually built based on deep learning algorithms (such as convolutional neural networks). By training a large amount of image data with depth annotations, the mapping relationship from image features to depth information is learned. Get the depth estimation value OutputImg[X target ,Y target ] and the depth estimation value OutputImg [X ankle ,Y ankle ], these depth estimates represent the distance information of the object surface where each pixel in the image is located; The difference between the depth estimates of the target part and the reference part is calculated as the offset: ; This offset reflects the relative depth relationship of the target part relative to the reference part in three-dimensional space. For example, if the depth estimate of the knee is less than the depth estimate of the ankle, then the offset is negative, indicating that the knee is closer to the camera than the ankle in the depth direction, which may mean that the leg is in a bent state. The depth difference between different parts can reflect the local posture changes of the body. For example, the body's forward leaning, backward leaning, twisting, and other movements can be accurately described by the continuous changes in the depth of each part of the body. By comprehensively considering the depth offsets of multiple parts, a richer and more detailed human posture model can be constructed.

[0032] The depth estimation model obtains depth estimates for the pixel positions of the target and reference parts, and calculates their difference as an offset. This depth information provides the position of the human body in three-dimensional space, more comprehensively reflecting changes in human posture than relying solely on two-dimensional image information. The depth differences between different body parts in three-dimensional space can help more accurately determine the relative position and posture of different body parts. For example, when determining the degree of leg bending, the depth difference between the knee and ankle can provide important clues as to whether the leg is bent or straight. Pose estimation based on depth offsets reduces information loss caused by two-dimensional image projection, thereby improving the accuracy of human pose estimation. This is especially true for complex poses and movements, effectively distinguishing subtle differences between similar poses. This solution maintains good pose estimation performance in various lighting environments (such as strong light, low light, and backlight) and when the camera view angle varies to a certain extent, reducing errors caused by environmental factors and improving system stability and reliability.

[0033] In one embodiment, in step S5, the physical height of the human body at different depths is corrected according to the offset, and the height of the target part in the actual physical space after correction is calculated includes: ; Among them, L physical_correct is the height of the corrected target part in the actual physical space, L p is the distance between the calibration plane and the camera lens during calibration, L c The target object's standing plane is the distance from the camera lens to the standing position. Offset is the difference between the depth estimates of the target and reference parts, or the offset. Compared to pose estimation methods without depth correction, this solution can obtain a more realistic description of human pose, reduce errors caused by depth changes, and significantly improve the accuracy of pose estimation, especially when processing images captured at long or close distances.

[0034] The basic principle of the correction method is to use the principle of similar triangles. In the camera imaging model, the size of the object in the image has a certain proportional relationship with the actual size. This proportional relationship is related to the distance between the object and the camera. By using the known calibration plane distance L p Distance L from standing position c , a conversion relationship from the target part in the image to the actual physical height can be established. When the depth offset Offset between the target part and the reference part is taken into account, this conversion relationship is actually being fine-tuned. For example, if the target part is closer to the camera than the reference part, such as Offset is negative, then when calculating the actual physical height, it is necessary to adjust it according to the degree of proximity so that the calculated height is more consistent with the actual position of the object in three-dimensional space. Specifically, the difference Offset between the depth estimation values ​​of the target part and the reference part is first obtained through image analysis, and then combined with the calibration distance L p and actual standing distance L c , calculate the corrected physical height L according to the above formula physical_correct This can combine the two-dimensional information in the image with the actual three-dimensional spatial information, and more accurately reflect the position and posture of the target part in the actual physical space.

[0035] By comparing the depth estimate with the reference depth estimate, an offset is calculated to correct the physical height of the human body at different depths. Because pixels at different depths are affected by factors such as camera characteristics and shooting angle, their actual physical heights may differ from those simply calculated based on the calibration plane. By calculating the offset and applying the correction, the true height at different depths can be reflected. This correction method effectively reduces the height estimation error caused by depth variations, making the corrected height of the target part in real physical space closer to the true value, thereby improving the accuracy of human pose estimation. This is particularly true when the human body is at different depths, such as at different distances, allowing for more precise positioning of various body parts in three-dimensional space. In real-world scenarios, the distance between the human body and the camera may vary, and individuals may have different heights and standing positions. By using the offset and related distance parameters for height correction, this scheme can adapt to different scene conditions. Whether in densely populated scenes with complex occupant positions, or at different shooting angles and distances, as long as accurate offset and distance information are available, effective height correction can be performed, ensuring the stability and accuracy of pose estimation. This enables the human body posture estimation method to operate reliably in various complex practical application scenarios, expands the scope of application of the method, and improves its value in practical applications.

[0036] Example 2 The present invention provides a body part correction device based on physical calibration and key point recognition under monocular vision, comprising: A physical calibration module is used to capture a monocular visual image at the same position using a monocular camera and establish a mapping relationship between the pixel position of the monocular visual image and the actual physical position based on physical calibration; A key point detection module is used to input the monocular visual image obtained by shooting into a key point detection model to detect the pixel position of the target part in the human body image; A physical height calculation module is used to map the detected pixel height to the physical height on the calibration plane and calculate the physical height of the target part based on the result of the physical calibration; The depth estimation module is used to input the monocular visual image obtained by capture into the depth estimation model to estimate the depth information of the pixel position of the target part in the monocular visual image, and compare the depth estimation value of the target part with the depth estimation value of the reference part to obtain the offset of the target part relative to the reference part; the physical height correction module is used to correct the physical height of the human body at different depths according to the offset, and calculate the height of the corrected target part in the actual physical space.

[0037] In one embodiment, the physical calibration module includes: A parameter setting unit, used to set the distance between the monocular camera and the target object to form a calibration plane and capture a monocular vision image for physical calibration; A data point selection unit is used to select data points from the calibration image according to equal pixel distances, record the vertical pixel positions of the data points in the calibration image, and form a pixel position sequence. ; The physical height reading unit is used to read the actual physical height corresponding to the data point to form a physical position sequence ; Fitting unit, used to create an array through polynomial fitting pixel to Array physical The mapping relationship is obtained to obtain the polynomial fitting coefficients (c0, c1, ... c 15 ), c0, c1, … c 15 Represents the polynomial fit coefficients for physical calibration.

[0038] The specific contents and implementation methods of the above modules and units are as described in Example 1 and will not be repeated here.

[0039] The present invention also provides an electronic device including a memory and a processor, wherein the memory stores computer-readable instructions. When the processor executes the computer-readable instructions, it implements the body part correction method based on physical calibration and key point recognition under monocular vision as described in the first embodiment of the present invention.

[0040] Likewise, the present application also protects a computer-readable storage medium loaded with a computer program for implementing the body part correction method based on physical calibration and key point recognition under monocular vision.

[0041] In each of the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program. The computer program includes one or more computer programs. When the computer program is loaded and executed on a computer, the process or function described in the embodiment of the present disclosure is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (such as coaxial cable, optical fiber, DDL (Digital Subscriber Line, Digital Subscriber Line)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a high-density DVD (Digital Video DiDD, digital video disc)), or a semiconductor medium (eg, a solid-state drive (SSD)).

[0042] It should be noted that, in the above-mentioned embodiment, the terms used herein are only for the purpose of describing specific exemplary embodiments, and are not intended to be restrictive. As used herein, the singular forms "one", "an" and "the or described" may be intended to also include plural forms, unless the context clearly indicates otherwise. The terms "comprise", "include" and "have" are inclusive and therefore specify the presence of stated features, wholes, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, parts and / or their groups. The method steps, processes and operations described herein should not be interpreted as necessarily requiring the method steps, processes and operations to be performed in the specific order discussed or shown, unless otherwise specified, they must be performed in the set step order. It should also be understood that additional or alternative steps may be adopted.

[0043] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A body part correction method based on physical calibration and key point recognition under monocular vision, characterized in that: The steps include: A monocular camera is used to capture a monocular image at the same position, and a mapping relationship between the pixel position of the monocular image and the actual physical position is established based on physical calibration; Inputting the captured monocular vision image into a key point detection model to detect the pixel position of the target part in the human body image; Mapping the detected pixel height to the physical height on the calibration plane and calculating the physical height of the target part based on the result of the physical calibration; Inputting the captured monocular visual image into a depth estimation model to estimate the depth information of the pixel position of the target part in the monocular visual image, and comparing the depth estimation value of the target part with the depth estimation value of the reference part to obtain the offset of the target part relative to the reference part; The physical height of the human body at different depths is corrected according to the offset, and the height of the corrected target part in the actual physical space is calculated.

2. The body part correction method based on physical calibration and key point recognition under monocular vision according to claim 1, characterized in that: The method of capturing a monocular visual image at the same position by a monocular camera and establishing a mapping relationship between the pixel position of the monocular visual image and the actual physical position based on physical calibration includes: Setting the distance between the monocular camera and the target object to form a calibration plane and capturing a monocular vision image for calibration; Data points are selected from the calibration image according to equal pixel distances, and the vertical pixel positions of the data points in the calibration image are recorded to form a pixel position sequence. ; Read the actual physical height corresponding to the data point to form a physical position sequence ; Create an Array by polynomial fitting pixel to Array physical The mapping relationship is obtained to obtain the polynomial fitting coefficients (c0, c1, ... c 15 ); Among them, c0, c1, ... c 15 Represents the polynomial fit coefficients for physical calibration.

3. The method for body part correction based on physical calibration and key point recognition under monocular vision according to claim 1, characterized in that: The step of inputting the captured monocular vision image into a key point detection model to detect the pixel position of the target part in the human body image includes: Collect monocular vision images of a target object standing naturally facing the direction of the monocular camera; Inputting the monocular visual image into a key point detection model; The pixel locations of target parts including the nose, lower center of the neck, shoulders, elbows, wrists, hips, knees and ankles are output by the keypoint detection model.

4. The method for body part correction based on physical calibration and key point recognition under monocular vision according to claim 1, characterized in that: Calculating the physical height of the target part includes: The physical height L corresponding to the target part is calculated using the polynomial fitting coefficient of the physical calibration physical : ; Among them, L pixel is the pixel height of the target part, , ,..., Represents the polynomial fit coefficients for physical calibration.

5. The method for body part correction based on physical calibration and key point recognition under monocular vision according to claim 1, characterized in that: The monocular visual image obtained by shooting is input into the depth estimation model to estimate the depth information of the pixel position of the target part in the monocular visual image, and the depth estimation value of the target part is compared with the depth estimation value of the reference part to obtain the offset of the target part relative to the reference part. Collect the monocular vision image InputImg of the target object standing naturally facing the direction of the monocular camera; Inputting the monocular vision image InputImg into the depth estimation model; Output the depth estimation value OutputImg of the target part at the pixel position through the depth estimation model; Get the depth estimation value OutputImg[X target ,Y target ] and the depth estimation value OutputImg [X ankle ,Y ankle ]; The difference between the depth estimates of the target part and the reference part is calculated as the offset: 。 6. The method for body part correction based on physical calibration and key point recognition under monocular vision according to claim 1, characterized in that: Correcting the physical height of the human body at different depths according to the offset and calculating the height of the corrected target part in the actual physical space includes: ; Among them, L physical_correct is the height of the corrected target part in the actual physical space, L p is the distance between the calibration plane and the camera lens during calibration, L c The target object's standing plane is the distance between the standing position and the camera lens, and Offset is the difference between the depth estimation values ​​of the target part and the reference part, that is, the offset.

7. A body part correction device based on physical calibration and key point recognition under monocular vision, characterized in that: A method for correcting body parts based on physical calibration and key point recognition under monocular vision as described in any one of claims 1 to 6 is implemented, wherein the device comprises: A physical calibration module is used to capture a monocular visual image at the same position using a monocular camera and establish a mapping relationship between the pixel position of the monocular visual image and the actual physical position based on physical calibration; A key point detection module is used to input the monocular visual image obtained by shooting into a key point detection model to detect the pixel position of the target part in the human body image; A physical height calculation module is used to map the detected pixel height to the physical height on the calibration plane and calculate the physical height of the target part based on the result of the physical calibration; A depth estimation module is used to input the captured monocular visual image into a depth estimation model to estimate the depth information of the pixel position of the target part in the monocular visual image, and compare the depth estimation value of the target part with the depth estimation value of the reference part to obtain the offset of the target part relative to the reference part; The physical height correction module is used to correct the physical height of the human body at different depths according to the offset, and calculate the height of the corrected target part in the actual physical space.

8. The body part correction device based on physical calibration and key point recognition under monocular vision according to claim 7, characterized in that: The physical calibration module includes: A parameter setting unit, used to set the distance between the monocular camera and the target object to form a calibration plane and capture a monocular vision image for physical calibration; A data point selection unit is used to select data points from the calibration image according to equal pixel distances, record the vertical pixel positions of the data points in the calibration image, and form a pixel position sequence. ; The physical height reading unit is used to read the actual physical height corresponding to the data point to form a physical position sequence ; Fitting unit, used to create an array through polynomial fitting pixel to Array physical The mapping relationship is obtained to obtain the polynomial fitting coefficients (c0, c1, ... c 15 ); Among them, c0, c1, ... c 15 Represents the polynomial fit coefficients for physical calibration.

9. An electronic device, characterized in that: include: a memory for storing a processing program; A processor, wherein when executing the processing program, the processor implements the body part correction method based on physical calibration and key point recognition under monocular vision as described in any one of claims 1 to 6.

10. A readable storage medium, characterized in that: The readable storage medium stores a processing program, which, when executed by the processor, implements the body part correction method based on physical calibration and key point recognition under monocular vision as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Height and body size measuring method based on monocular vision and key point recognition

    CN111611928A

  • Height measuring method, height measuring device and terminal

    CN114022532A