Personnel positioning method, device and equipment based on detection frame correction
By correcting the detection frame based on the detection model and key points of the human body, the problems of detection frame error and offset in traditional methods are solved, centimeter-level personnel positioning accuracy is achieved in complex industrial scenarios, and hardware costs are reduced.
Patent Information
- Application Number
- CN202510823332.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-10
AI Technical Summary
Traditional personnel positioning methods have detection frame errors and offset problems in complex scenarios, resulting in insufficient positioning accuracy, especially in industrial scenarios such as substations. It is difficult to achieve accurate positioning.
The detection model is used to detect image data, obtain the initial detection frame and human key points, and use the human key points to determine whether the detection frame meets the positioning reference constraints. If there is an offset, the detection frame is corrected and the target homography matrix is used for coordinate transformation to achieve centimeter-level positioning accuracy.
Without the need for additional hardware, the accuracy of personnel positioning is significantly improved, making it particularly suitable for complex industrial scenarios and reducing deployment costs and complexity.
Smart Images

Figure CN120765723A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a method, apparatus, and device for positioning a person based on detection frame correction. Background Art
[0002] With the rapid development of computer vision technology, visual personnel positioning has become one of the core technologies in fields such as intelligent security, smart cities, and industrial monitoring. Especially in scenarios such as substations, accurate positioning of personnel is extremely important for substation operation and maintenance.
[0003] Traditional personnel positioning technologies rely primarily on multi-sensor fusion, such as combining depth cameras and binocular cameras to enhance three-dimensional spatial positioning capabilities, or implementing personnel positioning through 3D modeling. These methods are expensive to deploy, making them difficult to promote in actual industrial scenarios. Personnel positioning solutions based on monocular cameras often rely on depth information, but depth estimation in depth detection models often has certain errors, and model deployment requires extensive data training.
[0004] Traditional personnel positioning solutions based on detection frames have poor accuracy in complex scenarios, such as substations with numerous devices and situations where two people are working, where people may be partially obscured. The detection frame may also be unilaterally offset from the main body of the human body due to the extension of the person's arms during equipment operation or maintenance. Summary of the Invention
[0005] The present application provides a personnel positioning method, apparatus and device based on detection frame correction, which is used to solve the problem of personnel positioning deviation caused by errors in the detection frame.
[0006] A first aspect of the present application provides a person positioning method based on detection frame correction, comprising: detecting a person to be positioned in image data using a preset detection model to obtain an initial detection frame and a plurality of key points of the human body;
[0007] Determining, based on the key points of the human body located within the initial detection frame, whether the initial detection frame satisfies a preset positioning reference constraint condition, wherein the positioning reference constraint condition at least includes that both the left frame line and the right frame line of the initial detection frame do not have unilateral offset;
[0008] If the left or right frame line has a unilateral offset, the shoulder joint point on the same side as the offset side is selected as the target key point;
[0009] Determining a target offset value based on a distance between the target key point and a corresponding frame line in the initial detection frame, and moving the corresponding frame line and correcting the detection frame according to the target offset value so that the corrected target detection frame satisfies the positioning reference constraint condition;
[0010] The pixel coordinates of the midpoint of the bottom frame line of the target detection frame are determined as the pixel coordinates to be located, and the pixel coordinates to be located are converted to the world coordinate system based on the preset target homography matrix to obtain the target positioning information corresponding to the person to be located.
[0011] A second aspect of the present application provides a person positioning device based on detection frame correction, comprising: a detection module, configured to detect a person to be positioned in image data using a preset detection model to obtain an initial detection frame and a plurality of key points of the human body;
[0012] a determination module, configured to determine, based on key points of the human body located within the initial detection frame, whether the initial detection frame satisfies a preset positioning reference constraint condition, wherein the positioning reference constraint condition at least includes that neither the left frame line nor the right frame line of the initial detection frame exhibits unilateral offset;
[0013] a correction module configured to select, if the left or right frame line has a unilateral offset, a shoulder joint point on the same side as the offset side as a target key point; determine a target offset value based on a distance between the target key point and a corresponding frame line in the initial detection frame, and move the corresponding frame line according to the target offset value and correct the detection frame so that the corrected target detection frame satisfies the positioning reference constraint condition;
[0014] The positioning module is used to determine the pixel coordinates of the midpoint of the bottom frame line of the target detection frame as the pixel coordinates to be positioned, and convert the pixel coordinates to be positioned into the world coordinate system based on a preset target homography matrix to obtain the target positioning information corresponding to the person to be positioned.
[0015] In the technical scheme provided in the present application, the initial detection frame and multiple human body key points are obtained by detecting the personnel to be positioned by the detection model, and whether the initial detection frame meets the preset positioning reference constraint condition is determined by the human body key points in the frame, such as whether the left frame line and the right frame line of the initial detection frame are not offset on one side, so as to avoid that when the personnel in the transformer substation and the like area overhauls, the arms are stretched to cause the side frame line to be offset on one side along with the stretching of the arms, thereby affecting the positioning accuracy, and an automatic correction mechanism of the detection frame is constructed, which can more accurately identify the abnormal state of the detection frame, and when the offset on one side is detected, the system automatically triggers the correction mechanism: the shoulder joint on the same side as the offset side is selected as the reference key point, the offset amount is determined by calculating the spatial distance between the key point and the corresponding frame line, and the frame line position is dynamically adjusted accordingly. The design is particularly suitable for asymmetric posture scenes such as sideways, the adjustment range is defined by the relative position relationship between the key point on the offset side and the frame line, the corrected detection frame meets the positioning reference requirement, and finally the midpoint of the bottom edge of the corrected detection frame is taken as the positioning reference point, the coordinate system mapping is completed in combination with the scene homography matrix, and the optimized spatial positioning information is output. The whole process does not need to rely on positioning labels or additional hardware devices, and the centimeter-level positioning accuracy is realized by a pure visual scheme, and the personnel positioning error in a complex industrial scene is significantly reduced. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 The first embodiment of the personnel positioning method based on detection frame correction in the present application is shown in the figure.
[0017] Figure 2 The second embodiment of the personnel positioning method based on detection frame correction in the present application is shown in the figure.
[0018] Figure 3 The third embodiment of the personnel positioning method based on detection frame correction in the present application is shown in the figure.
[0019] Figure 4 The first embodiment of the personnel positioning device based on detection frame correction in the present application is shown in the figure.
[0020] Figure 5 The second embodiment of the personnel positioning device based on detection frame correction in the present application is shown in the figure.
[0021] Figure 6 The first embodiment of the personnel positioning device based on detection frame correction in the present application is shown in the figure. DETAILED DESCRIPTION
[0022] The present application provides a personnel positioning method, device and equipment based on detection frame correction, which is used to solve the technical problem of low personnel positioning accuracy due to detection frame error.
[0023] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0024] In this application, image data is used to indicate an image in a target area monitored by an image acquisition device or an image obtained by extracting key frames from a surveillance video, wherein the image acquisition device can be various surveillance cameras in the target area, such as a monocular camera, etc., without specific limitation.
[0025] In practical applications, the position of the person to be located can be framed by a detection frame, and the position of the person from the camera can be obtained through similar triangles. This places high demands on the accuracy of the detection frame. Since the person stretches his arms during equipment operation or maintenance, the detection frame moves with the arm, that is, the left and / or right frame lines are unilaterally offset, causing the detection frame to deviate from the main body of the human body. At this time, the detection frame cannot provide accurate horizontal coordinate positioning information.
[0026] To solve the above problems, this embodiment constructs an automatic correction mechanism for the detection frame based on the key points of the human body of the person to be located to ensure that the target detection frame meets the positioning reference constraint conditions to improve the positioning accuracy. It can be applied to application scenarios such as substations that have high requirements for accurate positioning of personnel.
[0027] See also Figure 1 , providing a first embodiment of the present application's method for positioning a person based on detection frame correction, including:
[0028] 101. Detect the person to be located in the image data using a preset detection model to obtain an initial detection frame and multiple human key points.
[0029] It can be understood that the executor of the present application can be a personnel positioning device based on detection frame correction, or it can be a terminal, monitoring system or server; this embodiment can be applicable to single-target positioning scenarios (there is only one person to be located) or multi-target positioning scenarios (there are two or more people to be located), and the specific details are not limited here.
[0030] In this embodiment, the initial detection frame is used to represent the location area of the person to be identified in the image data; the human body key points are used to indicate the key skeleton points or joint points based on the human body. In this embodiment, both the initial detection frame and the human body key points can be automatically detected using a preset detection model.
[0031] It should be understood that target detection and human key point detection are usually two tasks, generally requiring different models to complete. In this scenario, some human key points may be outside the detection box. See the third embodiment below. However, some detection models can complete both tasks simultaneously. For example, they can first identify human key points and then output a detection box based on the human key points; or first identify the detection box and then mark the human key points within the detection box. In this scenario, all human key points are located inside the detection box. See the second embodiment below. You can choose according to the actual situation and there is no specific limitation.
[0032] The above-mentioned human body key points may include ankle points, knee joints, hip joints, shoulder joints, and may further include human skeletal frameworks such as the head, hands, and elbows. This embodiment does not limit the number of detected human body key points and their corresponding part types.
[0033] 102. Determine whether the initial detection frame satisfies a preset positioning reference constraint condition based on the key points of the human body located in the initial detection frame. The positioning reference constraint condition at least includes that both the left frame line and the right frame line of the initial detection frame do not have a unilateral offset.
[0034] In this embodiment, the positioning reference constraint is used to determine whether the detection frame can provide an accurate positioning reference. Specifically, the position and length of the positioning reference frame accurately represent the position of the person to be located. It should be understood that the positioning reference constraint can be set based on actual conditions. In this embodiment, the positioning reference constraint at least requires that both the left and right frame lines of the initial detection frame do not exhibit unilateral offset, to address errors caused by arm extension in maintenance scenarios.
[0035] The positioning reference constraint conditions of this embodiment can also be set as the positioning reference condition that the bottom frame line of the detection frame is in contact with the ground where the person is located; the positioning reference condition can be set as the detection frame completely frames the person to be located; the positioning reference condition can be set as the detection frame width covers the widest part of the human body, etc., without specific restrictions.
[0036] It should be noted that there are many ways to determine whether the initial detection frame meets the positioning reference constraint condition according to the human body key points in the frame, which can be set according to the set positioning reference condition, and can include but is not limited to: judging according to whether the frame includes human body key points of the target part type, judging according to the part type and number of human body detection points in the frame, and judging based on the distance difference between the human body key points in the frame and the detection frame according to the human body proportion relationship, without limitation.
[0037] It should be understood that if the initial detection frame meets the preset positioning reference constraint condition, the positioning based on the midpoint of the bottom frame line of the detection frame is more accurate at this time, and it is determined that the initial detection frame that meets the preset positioning reference constraint condition is determined as the target detection frame without detection frame correction. If the initial detection frame does not meet the preset positioning reference constraint condition, the detection frame is first corrected, that is, step 103 is executed, and the corrected detection frame is positioned as the target detection frame.
[0038] In this embodiment, the judgment reference of the automatic correction mechanism of the detection frame is the human body key points in the initial detection frame, that is, the human body key points in the frame, and the human body key points identified in step 101 can be located entirely in the initial detection frame or partially in the initial detection frame.
[0039] 103, if the left frame line or the right frame line appears unilateral deviation, the shoulder joint point on the same side of the deviation side is selected as the target key point.
[0040] In this embodiment, the target key point is used as the reference reference for detection frame correction, and the initial detection frame is corrected according to the relative position between the human body key points and the positioning reference frame line. The human body key points can be located in the initial detection frame, or can be located outside the initial detection frame, and the human body key points inside and outside the frame can also be used to correct the detection frame.
[0041] It should be understood that the positioning reference frame line is used to indicate the frame line that provides the positioning reference in the detection frame, which mainly indicates the left frame line and the right frame line of the detection frame, and in appropriate cases, it can be the bottom frame line of the detection frame. The left frame line and the right frame line usually define the width of the detection frame, which is closely related to the horizontal coordinate accuracy of the positioning information. If the left frame line and / or the right frame line deviates from the body trunk, it will cause deviation in horizontal positioning.
[0042] It should be noted that if two shoulder joints or two hip joints or two ankle points can be detected, the symmetric key points (such as two shoulder joints) can be directly used to correct unilateral deviation of the detection frame. Since the symmetric key points of the human body can better represent the backbone, positioning based on the coordinate mean of the symmetric key points can obtain a more accurate horizontal positioning reference.
[0043] However, during equipment operation or maintenance, personnel are usually in a sideways position due to the equipment placement and camera position, and only one side of the key point is detected (such as only one shoulder joint point is detected). In order to improve the robustness to the scene, this embodiment uses the shoulder joint point on the same side of the offset side as the target key point as the reference point for correction. For example, if the left frame line is offset, the left shoulder joint point is selected as the target key point; for example, if the right frame line is offset, the right shoulder joint point is selected as the target key point.
[0044] 104. Determine a target offset value based on the distance between the target key point and the corresponding frame line in the initial detection frame, move the corresponding frame line according to the target offset value, and correct the detection frame so that the corrected target detection frame meets the positioning reference constraint condition.
[0045] In this embodiment, the target offset value is used to indicate the degree of offset of the positioning reference frame line to be corrected, so as to adjust the initial detection frame to meet the positioning reference constraint conditions. The target offset value can be set according to the set positioning reference constraint conditions. The target offset value may include a lateral offset value, wherein the lateral offset value is used for the lateral translation adjustment of the side frame line of the detection frame, that is, the horizontal coordinate of the side frame point corresponding to the detection frame is adjusted.
[0046] It should be understood that the target offset value can also include a longitudinal offset value, wherein the longitudinal offset value is used to perform longitudinal translation adjustment on the bottom frame line of the initial detection frame along the longitudinal axis direction, such as the vertical upward or downward movement of the bottom frame line of the detection frame, that is, to adjust the vertical coordinates of all bottom frame line points, which will not be described in detail here.
[0047] The following target offset value uses the horizontal offset value as an example to provide a feasible automatic correction mechanism for the detection frame. If a unilateral offset is detected on the left or right frame line, that is, the person's arm is extended:
[0048] 1) If two shoulder joints are detected, the average of their horizontal coordinates is taken as the horizontal coordinate of the ground contact point; if only one shoulder joint or hip joint is detected (for example, when a person is leaning sideways), and one ankle is detected, the horizontal coordinate of the ankle is used as the horizontal coordinate of the ground contact point; if two ankles are detected, the average of the horizontal coordinates of the ankles is used as the horizontal coordinate of the ground contact point.
[0049] 2) If only one key point is detected (e.g., only one shoulder joint or one hip joint is identified), the stretching direction is determined by the relative position of the wrist, elbow, and shoulder. The position of the detection frame border in the same stretching direction is moved to the shoulder joint point, and a preset number of pixels are extended in that direction. The average of the horizontal coordinate of the pixel point and the horizontal coordinate of the border on the other side of the detection frame is taken as the horizontal coordinate of the ground contact point.
[0050] 105. Determine the pixel coordinates of the midpoint of the bottom frame line of the target detection frame as the pixel coordinates to be located, and convert the pixel coordinates to be located into the world coordinate system based on a preset target homography matrix to obtain the target positioning information corresponding to the person to be located.
[0051] It can be understood that the pixel coordinate system (u, v) has a unit scale of one pixel, is a discrete image coordinate or pixel coordinate, and has its origin at the upper left corner of the image.
[0052] In this embodiment, the target homography matrix is used to indicate the coordinate mapping relationship between the pixel coordinate system of the image acquisition device and the world coordinate system. The target homography matrix is used to implement the coordinate mapping between the two coordinate systems. The target homography matrix can be used to determine the coordinate information of any pixel point in the world coordinate system to achieve personnel positioning. This embodiment uses the target homography matrix to achieve coordinate mapping. After obtaining the mapping relationship, the positioning of the personnel in the camera can be achieved without deploying additional hardware or 3D models. The present invention is faster and less costly in engineering deployment. Compared with traditional monocular vision positioning solutions, it does not rely on depth information and is faster in personnel positioning.
[0053] In this embodiment, a detection model is used to detect the person to be located, generating an initial detection frame and multiple human key points. The key points within the frame are then used to determine whether the initial detection frame meets the preset positioning reference constraints. For example, the left and right frame lines of the initial detection frame exhibit no unilateral offset. This prevents unilateral offset of the side frame lines caused by extended arms during maintenance work in areas such as substations, which could affect positioning accuracy. An automatic correction mechanism for the detection frame is established, enabling more accurate identification of abnormal detection frame conditions. When a unilateral offset is detected, the system automatically triggers the correction mechanism: the shoulder joint point on the same side as the offset is selected as a reference key point. The offset is determined by calculating the spatial distance between this key point and the corresponding frame line, and the frame line position is dynamically adjusted accordingly. This design is particularly suitable for asymmetric posture scenarios, such as sideways posture. The relative positional relationship between the offset joint point and the frame line defines the adjustment range, ensuring that the corrected detection frame meets the positioning reference requirements. Finally, the midpoint of the bottom edge of the corrected detection frame is used as the positioning reference point. The coordinate system mapping is completed in combination with the scene homography matrix, and the optimized spatial positioning information is output. The entire process does not require reliance on positioning tags or additional hardware equipment. It achieves centimeter-level positioning accuracy through a pure visual solution, significantly reducing personnel positioning errors in complex industrial scenarios.
[0054] In practical applications, traditional computer vision-based technologies often rely on depth information to achieve the positioning of the person to be located in the world coordinate system. However, depth estimation often has errors, and the estimated depth information usually relies on hardware with high deployment costs such as depth cameras and binocular cameras, which makes it difficult to promote on a large scale. The depth estimation solution based on monocular cameras requires the training of a dedicated depth detection model, and its implementation is also relatively complex. Considering the convenience and cost of hardware deployment in actual industrial operation scenarios, refer to Figure 2 The second embodiment of the personnel positioning method based on detection frame correction of the present application is provided, which proposes a personnel positioning solution that does not rely on depth information. It only needs to obtain the pixel position of the person in the camera image, and the world coordinate information of the person can be directly obtained through the constructed target homography matrix mapping relationship.
[0055] 201. Construct a target homography matrix using multiple ground contact points in the image data and the actual coordinates of each ground contact point in the world coordinate system.
[0056] Specifically, multiple ground contact points are selected from the image data to obtain the pixel coordinates corresponding to each ground contact point; the actual coordinates of each ground contact point in the world coordinate system are determined based on a preset plan view of the target area; and a target homography matrix is constructed based on the pixel coordinates corresponding to each ground contact point and its actual coordinates. This embodiment achieves positioning by constructing a target homography matrix, which can reduce the hardware requirements for locating personnel in the target area.
[0057] Taking a substation as an example, the station typically includes equipment. The points where cabinets and walls contact the ground at the personnel's location can be selected as calibration points, resulting in multiple ground contact points. The substation floor plan typically includes the corresponding actual distances, such as the specific locations of each substation cabinet, which can be used to further determine the world coordinate position of each selected ground contact point. The target homography matrix is constructed using the pixel coordinates and actual coordinates of multiple ground contact points to complete the initial parameter configuration for personnel positioning. Subsequently, the actual coordinates of the contact points of the personnel to be located (such as the pixel coordinates of the feet or the pixel coordinates of the midpoint between the feet) can be inferred using the target homography matrix to achieve personnel positioning.
[0058] 202. Identify the person to be located by detecting the image data using a preset target model, and obtain an initial detection frame and key points of the human body.
[0059] In this embodiment, target model detection can simultaneously complete target detection and human key point detection. The target model can extract common features through a shared backbone network, and then locate the human body in the target detection stage (such as YOLO or Faster R-CNN), and then perform pose estimation (such as OpenPose or DensePose) on the detected human body area for key point detection.
[0060] In this embodiment, all detected human key points are located within the initial detection frame. Therefore, the pixel coordinates of the human key points are generally within the range constrained by the pixel coordinates of the four corner points of the initial detection frame. Due to various factors such as occlusion, diverse and dynamic human postures, the human key points output by the target model may not cover all human parts, and the initial detection frame may only frame a portion of the detected human parts, resulting in the initial detection frame being unable to provide an accurate positioning reference. To address this issue, this embodiment uses the human key point information within the frame to determine whether the initial detection frame meets the positioning reference constraints.
[0061] 203. Determine whether the initial detection frame satisfies a preset positioning reference constraint condition based on the type of the human body key point located in the initial detection frame.
[0062] In actual applications, in addition to the unilateral offset of the left and right frame lines mentioned above, which makes it impossible to provide accurate horizontal positioning and precise position, the initial detection frame output due to occlusion, algorithm errors and other reasons may only frame part of the human body, such as only framing the head or body of the person to be identified, causing the detection frame to "float", that is, the bottom frame line is not actually in contact with the ground where the person is located, and the bottom frame line is usually closely related to the accuracy of positioning information. In particular, in the positioning scheme where the midpoint of the bottom frame line is used as the position of the human body, if the bottom frame line is not in contact with the actual ground (such as the bottom frame line is biased up or down), it will cause deviations in the vertical positioning.
[0063] It should be understood that there are many ways to determine whether the initial detection frame meets the positioning reference constraint conditions based on the type of the body key points in the frame. This can be set according to the set positioning reference conditions, or further determined based on the type and number of the body key points in the frame. Some examples are provided below:
[0064] Exemplarily, whether the initial detection frame meets the positioning reference condition is determined based on whether the initial detection frame includes all types of human key points. In this embodiment, the initial detection frame including all types of human key points can provide accurate positioning results.
[0065] Exemplarily, whether the initial detection frame meets the positioning reference condition is determined based on whether the initial detection frame includes two ankle points.
[0066] It should be noted that if both ankle points fall within the initial detection frame, the midpoint of the bottom frame line of the initial detection frame can more accurately reflect the position of the person, and direct positioning can be achieved by identifying that the number of ankle points in the frame meets the requirements. At this time, even if other key points of the human body are not in the initial detection frame, since the key points of other parts of the human body (such as the head, etc.) have little impact on the positioning benchmark, the initial detection frame can achieve positioning even if it does not cover the entire human body. This embodiment can quickly determine whether the initial detection frame meets the positioning benchmark conditions by determining whether the initial detection frame includes two ankle points, thereby improving the efficiency of positioning processing.
[0067] Exemplarily, based on the human body proportion relationship, according to the part type of the human body key points located in the initial detection frame, and the pixel coordinate difference between the human body key points and the initial detection frame, it is determined whether the initial detection frame meets the preset positioning reference constraint conditions.
[0068] In this embodiment, the horizontal positioning reference and the vertical positioning reference can be determined based on the same key points of the human body, or can be determined based on different key points of the human body. The following describes the process of determining the horizontal positioning reference:
[0069] In some examples, a first vertical coordinate difference between the wrist point and the shoulder joint point on the same side, and a second vertical coordinate difference between the elbow point and the shoulder joint point on the same side in the initial detection frame are determined; when both the first vertical coordinate difference and the second vertical coordinate difference are less than a preset first difference threshold, it is determined that the corresponding frame line has a unilateral offset; when the first vertical coordinate difference or the second vertical coordinate difference is greater than or equal to the first difference threshold, it is determined that the corresponding frame line has no unilateral offset.
[0070] It should be understood that when a person's arm is extended, the relative values of the vertical coordinates of the wrist, elbow and shoulder will be relatively small. If only the vertical coordinate between two points is considered, such as only the vertical coordinate between the wrist and shoulder joint, human postures such as arm bending will be mistakenly judged as arm extension. This embodiment avoids false detection of unilateral offset of the frame line through comprehensive judgment of the wrist point and shoulder joint point on the same side, and the elbow point and shoulder joint point on the same side.
[0071] The following describes the process of determining the longitudinal positioning benchmark:
[0072] In some examples, when the ankle point is included in the initial detection frame and the difference in the third vertical coordinate between the ankle point and the bottom frame line is less than a preset second difference threshold, it is determined that the bottom frame line is in contact with the ground at the person's location; when the ankle point is not included in the initial detection frame, or the difference in the third vertical coordinate is greater than or equal to the second difference threshold, it is determined that the bottom frame line is not in contact with the ground at the person's location.
[0073] In this embodiment, the first difference threshold and the second difference threshold can be set according to actual conditions. For example, in a 2k resolution image, the first difference threshold can be set to 30 pixels. If it is determined that the first vertical coordinate difference and the second vertical coordinate difference are both less than 30 pixels, it is determined that the arm is stretched, that is, there is unilateral frame line deviation.
[0074] 204. If not, a target key point is selected from the plurality of human key points, and a target offset value is determined based on a distance between the target key point and a corresponding frame line in the initial detection frame.
[0075] In this embodiment, the target offset value is used to adjust the initial detection frame to meet the positioning reference constraint condition. When the horizontal positioning reference is not met, the target offset value is a horizontal offset value, which is used to indicate the distance of moving the corresponding frame line along the horizontal axis direction; when the vertical positioning reference is not met, the target offset value is a vertical offset value, which is used to indicate the distance of moving the corresponding frame line along the vertical axis direction; of course, the initial detection frame may not be able to provide accurate positioning reference in both horizontal and vertical directions, in which case the target offset value includes the horizontal offset value and the vertical offset value, which are not limited in this embodiment.
[0076] In some examples, when the initial detection frame does not meet the positioning reference constraint condition, a plurality of target points are selected based on the human proportion relationship, and the target offset value is determined according to the part type of each target point, the pixel coordinate difference between any two target points, and the distance between each target point and the positioning reference frame line.
[0077] It can be understood that the height and body shape of the person to be positioned can be determined based on the pixel coordinate difference between any two target points, and the distance between each target point and the positioning reference frame line is combined to correct the bottom frame line of the detection frame to touch the ground and the left and right frame lines not to deviate from the main body of the human body.
[0078] In some examples, taking the calculation of the horizontal offset value as an example: the pixel horizontal coordinate difference between the offset side frame line and the same side shoulder joint point is determined as a first offset value, and the difference between the first offset value and a preset horizontal offset compensation value is determined as the horizontal offset value; the offset side frame line is moved along the horizontal axis direction by the horizontal offset value, and the length of the bottom frame line in the initial detection frame is adjusted to be connected to the moved offset side frame line to obtain a target detection frame.
[0079] For example, if only one side key point is detected, the stretching direction is determined by the relative position relationship of the wrist, elbow and shoulder; the position of the detection frame frame with the same stretching direction is moved to the shoulder joint point, and is extended by 20 pixel points in this direction, and then the horizontal coordinate of the pixel point and the horizontal coordinate of the other side frame of the detection frame are averaged as the horizontal coordinate of the ground contact point.
[0080] In some examples, the calculation of the vertical offset value is used as an example: if the target key point is located inside the initial detection frame, the pixel vertical coordinate difference between the target key point and the reference key point is determined as the second offset value, and the second offset value is used as the vertical offset value; if the target key point is located outside the initial detection frame, the pixel vertical coordinate difference between the bottom frame line of the initial frame line and the target key point is determined as the third offset value, and the sum of the second offset value and the third offset value is determined as the vertical offset value. The bottom frame line of the initial detection frame is moved along the vertical axis by the vertical offset value, and the lengths of the left and right frame lines in the initial detection frame are adjusted to connect with the moved bottom frame line to obtain the target detection frame.
[0081] For example, 1) when an ankle point is detected: if the pixel vertical coordinate of the bottom frame line of the detection frame is greater than the vertical coordinate of any ankle point, no correction is required; otherwise, the pixel vertical coordinate of the bottom frame line of the detection frame is extended to the vertical coordinate position of the ankle point;
[0082] 2) If the ankle point is not detected but the part above it is detected: Get the distance from the knee joint keypoint to the hip joint. If the bottom line of the detection frame is lower than the knee joint keypoint, add the distance from the knee joint to the vertical coordinate of the bottom line of the detection frame. If the bottom line of the detection frame is higher than the knee joint, first extend the bottom line of the detection frame to the knee joint and add the distance from the knee joint to the hip joint.
[0083] 3) If the knee joint is not detected but the body parts above it are detected: Get the distance from the shoulder joint to the hip joint. If the bottom line of the detection frame is lower than the hip joint, add the distance from the shoulder joint to the hip joint. Otherwise, extend the bottom line of the detection frame to the hip joint and add the distance from the shoulder joint to the hip joint.
[0084] 4) If the hip joint is not detected but the above parts are detected: Get the distance between the top of the detection frame and the shoulder joint. If the bottom frame line of the detection frame is lower than the shoulder joint, add n times the distance between the top of the detection frame and the shoulder joint to the vertical coordinate of the bottom frame line of the detection frame. Otherwise, extend the bottom frame line of the detection frame to the shoulder joint position and then add n times the distance between the top of the detection frame and the shoulder joint, where n is usually 4 or 5.
[0085] 205. Adjust the initial detection frame according to the target offset value to obtain the target detection frame.
[0086] Specifically, the positioning reference frame line of the initial detection frame is moved axially according to the target offset value, and the target detection frame is corrected to satisfy the positioning reference constraint condition.
[0087] In this embodiment, the axial direction refers to the horizontal axis direction and the vertical axis direction corresponding to the pixel coordinate system. If the target offset value is the vertical offset value, the moving direction is the vertical axis direction. If the target offset value is the horizontal offset value, the moving direction is the horizontal axis direction.
[0088] The positioning reference frame line to be adjusted includes at least one of the left frame line, the right frame line and the bottom frame line. The corresponding frame line to be adjusted is determined according to the corresponding offset situation so that the target detection frame meets the positioning reference constraint condition.
[0089] It can be understood that the bottom frame line can be moved vertically by adding or subtracting the vertical offset value from the pixel vertical coordinates of each bottom frame line point located in the initial detection frame to adjust the position of the bottom frame line of the initial detection frame up and down, so that the adjusted bottom frame line is exactly in contact with the ground where the person is located; similarly, the target side frame line is moved horizontally by the pixel horizontal coordinates of each target side frame line point to adjust the target side frame line left and right, so that the human body trunk is exactly centered in the adjusted detection frame.
[0090] 206. Determine the pixel coordinates of the midpoint of the bottom frame line of the target detection frame as the pixel coordinates to be located, and convert the pixel coordinates to be located into the world coordinate system based on a preset target homography matrix to obtain target positioning information corresponding to the person to be located.
[0091] Specifically, the pixel coordinates of the two endpoints of the bottom frame line of the target detection frame are obtained, and the average value is taken to obtain the coordinates of the pixel to be located; the pixel coordinates to be located are homogenized and mapped to the world coordinate system through the preset target homography matrix to obtain the candidate positioning coordinates; the candidate positioning coordinates are normalized to obtain the target positioning information corresponding to the person to be located.
[0092] In this embodiment, the target homography matrix is constructed by mapping the pixel coordinates of multiple ground contact points to their actual coordinates.
[0093] Homogenization expands the dimension of the coordinate representation by adding a dimension, which in this embodiment refers to expanding from two-dimensional pixel coordinates to three-dimensional ones.
[0094] For example, the pixel coordinates (u, v) are expanded to (u, v, 1) through homogenization, and the coordinate system is transformed through the target homography matrix to obtain the candidate positioning coordinates (X, Y, Z). After normalization, the target positioning information is obtained as (X', Y'), where X'=X / Z and Y'=Y / Z.
[0095] In this embodiment, a target homography matrix is constructed by constructing multiple ground contact points in the image and their world coordinates to ultimately provide a more accurate spatial reference for positioning. The person to be located is initially identified using a preset target model initial detection frame, and an automatic correction mechanism for the detection frame is constructed based on the human body key points within the initial detection frame. Whether the detection frame can provide an accurate positioning reference is determined by the type of the human body key points within the frame. When it is determined that the detection frame needs to be corrected, the target offset value is dynamically calculated based on the human body proportion relationship and the human body key points so that the target detection frame meets the preset positioning reference constraints, ensuring that accurate positioning can be achieved based on the target detection frame. This solves the problem of the detection frame being suspended, unilaterally offset, or not meeting the positioning reference constraints. The detection frame is accurately corrected through the target key points and the corresponding frame lines, thereby improving the positioning accuracy of the detection frame. The pixel coordinates of the midpoint of the bottom frame line of the target detection frame are used as the position of the person to be located, and the target positioning information is determined by performing coordinate system transformation in combination with the target homography matrix, which can improve the accuracy of person positioning.
[0096] In practical applications, in the solution of achieving positioning by the midpoint of the bottom frame line of the detection frame, whether the detection frame can provide an accurate positioning reference is closely related to the accuracy of the vertical coordinate of the bottom frame line of the detection frame and the horizontal coordinates of the left and right frame lines. The following describes the third embodiment of the personnel positioning method based on detection frame correction of this application:
[0097] 301. Identify the person to be located in the image data through a preset target detection network to obtain an initial detection frame, and identify the person to be located in the image data through a preset human key point detection network to obtain human key points.
[0098] In this embodiment, the target detection network can be various network models suitable for object area recognition, including but not limited to Faster-RCNN, YOLO and other models that can detect the location area of a person; and the human key point detection network can be various posture estimation models, including but not limited to OpenPose, AlphaPose and other models, without specific limitation.
[0099] 302. Based on the key points of the human body located in the initial detection frame, determine whether the left frame line and the right frame line of the initial detection frame have no unilateral offset, and obtain a horizontal reference judgment result.
[0100] In this embodiment, it is determined whether the detection frame has no unilateral offset by the key points of the human body in the frame, so as to ensure that the detection frame can provide an accurate horizontal coordinate positioning reference.
[0101] Optionally, when the first vertical coordinate difference between the wrist point and the shoulder joint point on the same side in the initial detection frame is greater than a preset first difference threshold, or the first vertical coordinate difference between the elbow joint point and the shoulder joint point on the same side is greater than the first difference threshold, the lateral reference judgment result is determined to be that the initial detection frame has a unilateral offset; when the first vertical coordinate difference and the first vertical coordinate difference are both less than or equal to the first difference threshold, the lateral reference judgment result is determined to be that the initial detection frame has no unilateral offset.
[0102] 303. Based on the key points of the human body in the initial detection frame, determine whether the bottom frame line of the initial detection frame is in contact with the ground where the person is located, and obtain a longitudinal reference judgment result.
[0103] Optionally, when the ankle point is included in the initial detection frame, and the difference in the third vertical coordinate between the ankle point and the bottom frame line of the initial detection frame is within a preset second difference threshold range, the longitudinal reference judgment result is determined to be that the bottom frame line is in contact with the ground at the person's location; when the ankle point is not included in the initial detection frame, or the difference in the third vertical coordinate exceeds the second difference threshold range, the longitudinal reference judgment result is determined to be that the bottom frame line is not in contact with the ground at the person's location.
[0104] 304. If both the longitudinal reference judgment result and the lateral reference judgment result are yes, then it is determined that the initial detection frame meets the preset positioning reference constraint condition; otherwise, it is determined that the initial detection frame does not meet the positioning reference constraint condition.
[0105] Specifically, if both the longitudinal benchmark judgment result and the lateral benchmark judgment result are yes, it is determined that the initial detection frame meets the preset positioning benchmark constraint conditions, and the initial detection frame is directly determined as the target detection frame; if the longitudinal benchmark judgment result and / or the lateral benchmark judgment result is no, it is determined that the initial detection frame does not meet the positioning benchmark constraint conditions.
[0106] 305. If the result of the lateral reference judgment is no, the shoulder joint point on the same side as the offset side is selected as the target key point, and the lateral offset value is determined according to the positioning reference frame line of the target key point and the initial frame line.
[0107] Specifically, the ipsilateral shoulder joint point is selected as the target key point, and the pixel horizontal coordinate difference between the offset side frame line and the ipsilateral shoulder joint point is determined as the first offset value; the difference between the first offset value and the preset lateral offset compensation value is determined as the lateral offset value.
[0108] Among them, the lateral offset compensation value is a preset empirical value, which is used to adjust the distance between the shoulder joint and the frame line on the same side. Its specific value can be set according to actual conditions. For example, for a 2K resolution image, it can be set to 20 pixels.
[0109] 306. If the result of the longitudinal reference judgment is no, the human body key point with the largest pixel longitudinal coordinate is selected as the target key point, and the reference key point is determined according to the part type of the target key point.
[0110] Exemplarily, the human body key point with the largest pixel vertical coordinate is selected as the target key point; the reference key point is determined according to the part type of the target key point, and the target offset value is determined based on the pixel vertical coordinate difference between the reference key point and the target key point, wherein the pixel vertical coordinate of the reference key point is smaller than the pixel vertical coordinate of the target key point.
[0111] In this embodiment, the vertical coordinates of the bottom frame line points of the detection frame are adjusted based on the human body proportion relationship and the vertical distance between the key points of the human body, so that the target detection frame can just touch the ground where the person is located, providing an accurate vertical coordinate positioning reference; by dynamically adjusting the longitudinal offset value in the key point combination, the robustness of the algorithm is enhanced, it can adapt to different heights, and reduce dependence on key points of a single category.
[0112] In this embodiment, the target key point is located inside or outside the initial detection frame. The above-mentioned determination of the target offset value based on the pixel vertical coordinate difference between the reference key point and the target key point includes: if the target key point is located inside the initial detection frame, then the second offset value is determined as the target offset value, and the second offset value is the pixel vertical coordinate difference between the reference key point and the target key point; if the target key point is located outside the initial detection frame, then the sum of the second offset value and the third offset value is determined as the target offset value, and the third offset value is the pixel vertical coordinate difference between the bottom frame line point of the initial detection frame and the target key point.
[0113] In this embodiment, the determination of the reference key point is associated with the part type of the target key point. The pixel vertical coordinate of the reference key point is smaller than the pixel vertical coordinate of the target key point, that is, the reference key point is located above the target key point. It may be a human body key point located above the target key point. Under appropriate circumstances, the reference key point may also be the top (top frame line) point of the initial detection frame.
[0114] Optionally, the above-mentioned determination of the reference key point based on the part type of the target key point includes: if the target key point is the ankle point, the bottom frame line point of the initial detection frame is determined as the reference key point; if the target key point is the knee joint point, the hip joint point is determined as the reference key point; if the target key point is the hip joint point, the shoulder joint point is determined as the reference key point; if the target key point is the shoulder joint point, the top frame line point of the initial detection frame is determined as the reference key point.
[0115] 307. Move the corresponding positioning reference frame line according to the target offset value and correct the detection frame so that the corrected target detection frame meets the positioning reference constraint condition.
[0116] Specifically, if the left frame line and / or right frame line of the initial detection frame is offset on one side, the offset side frame line is moved along the horizontal axis by the lateral offset value, and the bottom frame line of the initial frame line is adjusted until the bottom frame line is connected to the offset side frame line after the move to obtain the target detection frame; if the bottom frame line of the initial detection frame does not contact the ground where the person is located, the bottom frame line is moved along the vertical axis by the longitudinal offset value and connected to the left and right frame lines to obtain the target detection frame.
[0117] The above adjustment includes extending or shortening the frame line corresponding to the initial frame line until the corresponding frame line is connected to the moved frame line. It can be understood that the initial detection frame may have only a unilateral offset on the left frame line, or only a unilateral offset on the right frame line, or only the bottom frame line may be suspended. It is also possible that more than two frame lines among the left frame line, bottom frame line, and right frame line cannot provide an accurate positioning reference. Among them, the unilateral offset direction can be the left or right side of the corresponding frame line, and the suspension direction may also be the upper or lower side, without specific restrictions.
[0118] 308. Determine the pixel coordinates of the midpoint of the bottom frame line of the target detection frame as the pixel coordinates to be located, and map the pixel coordinates to be located to the world coordinate system based on a preset target homography matrix to obtain target positioning information corresponding to the person to be located.
[0119] Step 308 can be performed with reference to step 206 and will not be described again here.
[0120] In this embodiment, an initial detection frame of the person to be located is first acquired through a target detection network. A human keypoint detection network is then used to extract the coordinates of keypoints, providing a data foundation for subsequent corrections. Based on the distribution of keypoints within the initial detection frame, an automatic correction logic is constructed. The validity of the detection frame is determined by analyzing the location and type of keypoints. When abnormal conditions such as floating or unilateral offset are detected, a targeted correction process is initiated. During the longitudinal correction phase, the system selects pixel extremes of the vertical coordinate to form target keypoint pairs. By calculating the vertical spacing between these keypoint pairs and the bottom line of the detection frame, as well as the vertical distance between these keypoint pairs, the longitudinal offset is quantified and compensated, effectively addressing vertical coordinate deviations caused by the detection frame not being aligned with the ground. During the transverse correction phase, a reference side frame line is determined based on the offset direction. The horizontal spacing between this side frame line and the ipsilateral shoulder joint is calculated. The transverse coordinate deviation is corrected using the offset compensation, mitigating the problem of trunk area offset caused by limb extension. Finally, the midpoint of the corrected bottom line of the detection frame is used as the positioning reference point. Combined with the target homography matrix, the coordinate system is mapped to the target, and optimized spatial positioning information is output. This solution significantly improves the accuracy of personnel positioning in complex scenarios through a key point-driven detection frame adaptive adjustment mechanism.
[0121] The above describes the personnel positioning method based on the detection frame correction in this application. The following describes the personnel positioning device based on the detection frame correction in this application. Figure 4 In this application, an embodiment of a personnel positioning device based on detection frame correction includes:
[0122] Detection module 401, used to detect the person to be located in the image data using a preset detection model to obtain an initial detection frame and multiple human key points;
[0123] Determining module 402, for determining whether the initial detection frame satisfies a preset positioning reference constraint based on the key points of the human body located within the initial detection frame, the positioning reference constraint at least including that both the left and right frame lines of the initial detection frame do not exhibit unilateral offset;
[0124] Correction module 403 is configured to select the shoulder joint point on the same side as the offset side as the target key point if the left or right frame line is unilaterally offset; determine a target offset value based on the distance between the target key point and the corresponding frame line in the initial detection frame; move the corresponding frame line according to the target offset value and correct the detection frame so that the corrected target detection frame meets the positioning reference constraint condition;
[0125] The positioning module 404 is used to determine the pixel coordinates of the midpoint of the bottom frame line of the target detection frame as the pixel coordinates to be positioned, and convert the pixel coordinates to be positioned to the world coordinate system based on a preset target homography matrix to obtain the target positioning information corresponding to the person to be positioned.
[0126] In this embodiment, a detection model is used to detect the person to be located, generating an initial detection frame and multiple human key points. The key points within the frame are then used to determine whether the initial detection frame meets the preset positioning reference constraints. For example, the left and right frame lines of the initial detection frame exhibit no unilateral offset. This prevents unilateral offset of the side frame lines caused by extended arms during maintenance work in areas such as substations, which could affect positioning accuracy. An automatic correction mechanism for the detection frame is established, enabling more accurate identification of abnormal detection frame conditions. When a unilateral offset is detected, the system automatically triggers the correction mechanism: the shoulder joint point on the same side as the offset is selected as a reference key point. The offset is determined by calculating the spatial distance between this key point and the corresponding frame line, and the frame line position is dynamically adjusted accordingly. This design is particularly suitable for asymmetric posture scenarios, such as sideways posture. The relative positional relationship between the offset joint point and the frame line defines the adjustment range, ensuring that the corrected detection frame meets the positioning reference requirements. Finally, the midpoint of the bottom edge of the corrected detection frame is used as the positioning reference point. The coordinate system mapping is completed in combination with the scene homography matrix, and the optimized spatial positioning information is output. The entire process does not require reliance on positioning tags or additional hardware equipment. It achieves centimeter-level positioning accuracy through a pure visual solution, significantly reducing personnel positioning errors in complex industrial scenarios.
[0127] See also Figure 5 Another embodiment of the personnel positioning device based on detection frame correction in the present application includes:
[0128] The detection module 401 is configured to detect a person to be positioned in the image data by using a preset detection model to obtain an initial detection frame and a plurality of human body key points.
[0129] The determination module 402 is configured to determine whether the initial detection frame satisfies a preset positioning reference constraint condition according to the human body key points located in the initial detection frame, the positioning reference constraint condition at least including that neither a left frame line nor a right frame line of the initial detection frame appears unilateral deviation.
[0130] The correction module 403 is configured to, if the unilateral deviation appears in the left frame line or the right frame line, select a shoulder joint point on the same side as the deviated side as a target key point, determine a target offset value based on a distance between the target key point and a corresponding frame line in the initial detection frame, and move the corresponding frame line and correct the detection frame according to the target offset value, so that the target detection frame satisfies the positioning reference constraint condition.
[0131] The positioning module 404 is configured to determine a bottom frame line midpoint pixel coordinate of the target detection frame as a pixel coordinate to be positioned, and convert the pixel coordinate to be positioned to a world coordinate system based on a preset target homography matrix to obtain target positioning information corresponding to the person to be positioned.
[0132] Optionally, the determination module 402 includes:
[0133] The transverse reference determination unit 4021 is configured to determine a first vertical coordinate difference value between a same-side wrist point and a shoulder joint point in the initial detection frame, and a second vertical coordinate difference value between a same-side elbow point and the shoulder joint point, and determine that the corresponding frame line appears unilateral deviation when the first vertical coordinate difference value and the second vertical coordinate difference value are both less than a preset first difference threshold value, and determine that the corresponding frame line does not appear unilateral deviation when the first vertical coordinate difference value or the second vertical coordinate difference value is greater than or equal to the first difference threshold value.
[0134] Optionally, the vertical reference determination unit 4022 is configured to determine that the bottom frame line is in contact with the ground at the position of the person when the initial detection frame includes an ankle point and a third vertical coordinate difference value between the ankle point and the bottom frame line is less than a preset second difference threshold value.
[0135] When the initial detection frame does not include the ankle point or the third vertical coordinate difference value is greater than or equal to the second difference threshold value, it is determined that the bottom frame line is not in contact with the ground at the position of the person.
[0136] Optionally, the correction module 403 includes:
[0137] The transverse correction unit 4031 is configured to determine a pixel transverse coordinate difference value between the deviated side frame line and the same-side shoulder joint point as a first offset value, and determine a difference value between the first offset value and a preset transverse offset compensation value as a transverse offset value.
[0138] The offset side frame line is moved along the horizontal axis by the horizontal offset value, and the length of the bottom frame line in the initial detection frame is adjusted to connect it with the offset side frame line after the movement to obtain the target detection frame.
[0139] Optionally, the longitudinal correction unit 4032 is used to select the human body key point with the largest pixel vertical coordinate as the target key point if the bottom frame line does not contact the ground where the person is located, and determine the reference key point according to the part type of the target key point, wherein the reference key point is a human body key point with a pixel vertical coordinate smaller than the target key point.
[0140] Optionally, the longitudinal correction unit 4032 is specifically configured to determine the bottom frame line point of the initial detection frame as a reference key point if the target key point is an ankle point;
[0141] If the target key point is the knee joint, the hip joint is determined as the reference key point;
[0142] If the target key point is the hip joint, the shoulder joint is determined as the reference key point;
[0143] If the target key point is the shoulder joint point, the top frame line point of the initial detection frame is determined as the reference key point.
[0144] Optionally, the longitudinal correction unit 4032 is further configured to: if the target key point is located inside the initial detection frame, determine the pixel vertical coordinate difference between the target key point and the reference key point as a second offset value, and use the second offset value as the longitudinal offset value;
[0145] If the target key point is outside the initial detection frame, the pixel vertical coordinate difference between the bottom frame line of the initial frame line and the target key point is determined as the third offset value, and the sum of the second offset value and the third offset value is determined as the vertical offset value.
[0146] The bottom frame line of the initial detection frame is moved along the vertical axis by the longitudinal offset value, and the lengths of the left and right frame lines in the initial detection frame are adjusted to connect them with the moved bottom frame line to obtain the target detection frame.
[0147] Optionally, the positioning module 404 is specifically used to homogenize the pixel coordinates to be positioned, and map them to the world coordinate system through a preset target homography matrix to obtain candidate positioning coordinates; normalize the candidate positioning coordinates to obtain target positioning information corresponding to the person to be positioned.
[0148] Optionally, the personnel positioning device based on detection frame correction further includes a construction module 405 for selecting a plurality of ground contact points in the image data to obtain pixel coordinates corresponding to each ground contact point;
[0149] Determine the actual coordinates of each ground contact point in the world coordinate system based on a preset target area plan;
[0150] The target homography matrix is constructed based on the pixel coordinates corresponding to each ground contact point and its actual coordinates.
[0151] In this embodiment, an initial detection frame of the person to be located is first acquired through a target detection network. A human keypoint detection network is then used to extract the coordinates of keypoints, providing a data foundation for subsequent corrections. Based on the distribution of keypoints within the initial detection frame, an automatic correction logic is constructed. The validity of the detection frame is determined by analyzing the location and type of keypoints. When abnormal conditions such as floating or unilateral offset are detected, a targeted correction process is initiated. During the longitudinal correction phase, the system selects pixel extremes of the vertical coordinate to form target keypoint pairs. By calculating the vertical spacing between these keypoint pairs and the bottom line of the detection frame, as well as the vertical distance between these keypoint pairs, the longitudinal offset is quantified and compensated, effectively addressing vertical coordinate deviations caused by the detection frame not being aligned with the ground. During the transverse correction phase, a reference side frame line is determined based on the offset direction. The horizontal spacing between this side frame line and the ipsilateral shoulder joint is calculated. The transverse coordinate deviation is corrected using the offset compensation, mitigating the problem of trunk area offset caused by limb extension. Finally, the midpoint of the corrected bottom line of the detection frame is used as the positioning reference point. Combined with the target homography matrix, the coordinate system is mapped to the target, and optimized spatial positioning information is output. This solution significantly improves the accuracy of personnel positioning in complex scenarios through a key point-driven detection frame adaptive adjustment mechanism.
[0152] above Figure 4 and Figure 5 The personnel positioning device based on detection frame correction in the present application is described in detail from the perspective of modular functional entities. The personnel positioning device based on detection frame correction in the present application is described in detail from the perspective of hardware processing.
[0153] See also Figure 6 As shown, the personnel positioning device based on detection frame correction includes a processor 600 and a memory 601. The memory 601 stores machine executable instructions that can be executed by the processor 600. The processor 600 executes the machine executable instructions to implement the above-mentioned personnel positioning method based on detection frame correction.
[0154] Furthermore, Figure 6 The personnel positioning device based on detection frame correction shown further includes a bus 602 and a communication interface 603 , and the processor 600 , the communication interface 603 and the memory 601 are connected via the bus 602 .
[0155] Among them, the memory 601 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), for example, at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 603 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 602 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0156] The processor 600 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 600. The above processor 600 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 601 , and the processor 600 reads the information in the memory 601 and completes the method steps of the aforementioned embodiment in combination with its hardware.
[0157] The present application also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of a personnel positioning method based on detection frame correction.
[0158] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0159] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0160] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A person positioning method based on detection frame correction, characterized in that: The personnel positioning method based on detection frame correction includes: The person to be located in the image data is detected using a preset detection model to obtain an initial detection frame and multiple key points of the human body; Determining, based on the key points of the human body located within the initial detection frame, whether the initial detection frame satisfies a preset positioning reference constraint condition, wherein the positioning reference constraint condition at least includes that both the left frame line and the right frame line of the initial detection frame do not have unilateral offset; If the left or right frame line has a unilateral offset, the shoulder joint point on the same side as the offset side is selected as the target key point; Determining a target offset value based on a distance between the target key point and a corresponding frame line in the initial detection frame, and moving the corresponding frame line and correcting the detection frame according to the target offset value so that the corrected target detection frame satisfies the positioning reference constraint condition; The pixel coordinates of the midpoint of the bottom frame line of the target detection frame are determined as the pixel coordinates to be located, and the pixel coordinates to be located are converted to the world coordinate system based on the preset target homography matrix to obtain the target positioning information corresponding to the person to be located.
2. The method for positioning people based on detection frame correction according to claim 1, characterized in that: The determining, based on the key points of the human body located within the initial detection frame, whether the initial detection frame satisfies a preset positioning reference constraint condition includes: Determining a first ordinate difference between the wrist point and the shoulder joint point on the same side in the initial detection frame, and a second ordinate difference between the elbow point and the shoulder joint point on the same side; When the first vertical coordinate difference and the second vertical coordinate difference are both smaller than a preset first difference threshold, it is determined that the corresponding frame line has a unilateral offset; When the first vertical coordinate difference or the second vertical coordinate difference is greater than or equal to the first difference threshold, it is determined that the corresponding frame line does not have a unilateral offset.
3. The method for positioning people based on detection frame correction according to claim 1, characterized in that: When the left frame line and / or the right frame line has a unilateral offset, determining a target offset value based on the distance between the target key point and the corresponding frame line in the initial detection frame, moving the corresponding frame line according to the target offset value, and correcting the detection frame so that the corrected target detection frame satisfies the positioning reference constraint condition, including: Determine the pixel horizontal coordinate difference between the offset side frame line and the ipsilateral shoulder joint point as a first offset value, and determine the difference between the first offset value and a preset lateral offset compensation value as a lateral offset value; The offset side frame line is moved along the horizontal axis by the horizontal offset value, and the length of the bottom frame line in the initial detection frame is adjusted to connect the shifted offset side frame line to obtain the target detection frame.
4. The method for positioning people based on detection frame correction according to claim 2, characterized in that: The positioning reference constraint condition also includes that the bottom frame line of the initial detection frame is in contact with the ground where the person is located; The determining, based on the key points of the human body located in the initial detection frame, whether the initial detection frame satisfies a preset positioning reference constraint condition further includes: When the initial detection frame includes an ankle point and a third vertical coordinate difference between the ankle point and the bottom frame line is less than a preset second difference threshold, determining that the bottom frame line is in contact with the ground at the position of the person; When the initial detection frame does not include the ankle point, or the third vertical coordinate difference is greater than or equal to the second difference threshold, it is determined that the bottom frame line is not in contact with the ground at the position of the person.
5. The method for positioning people based on detection frame correction according to claim 1, characterized in that: The positioning reference constraint condition also includes that the bottom frame line of the initial detection frame is in contact with the ground where the person is located; After determining whether the initial detection frame satisfies a preset positioning reference constraint condition based on the human body key points located in the initial detection frame, and before determining a target offset value based on the distance between the target key points and the corresponding frame line in the initial detection frame, and moving the corresponding frame line according to the target offset value and correcting the detection frame so that the corrected target detection frame satisfies the positioning reference constraint condition, the method further includes: If the bottom frame line does not contact the ground where the person is located, the human body key point with the largest pixel vertical coordinate is selected as the target key point, and the reference key point is determined according to the part type of the target key point, wherein the reference key point is a human body key point with a pixel vertical coordinate smaller than the target key point.
6. The method for positioning people based on detection frame correction according to claim 5, characterized in that: Determining the reference key point according to the location type of the target key point includes: If the target key point is the ankle point, the bottom frame line point of the initial detection frame is determined as the reference key point; If the target key point is the knee joint, the hip joint is determined as the reference key point; If the target key point is the hip joint point, the shoulder joint point is determined as the reference key point; If the target key point is a shoulder joint point, the top frame line point of the initial detection frame is determined as the reference key point.
7. The method for positioning people based on detection frame correction according to claim 5, characterized in that: When the bottom frame line is not in contact with the ground at the location of the person, determining a target offset value based on the distance between the target key point and the corresponding frame line in the initial detection frame, moving the corresponding frame line according to the target offset value, and correcting the detection frame so that the corrected target detection frame satisfies the positioning reference constraint condition, including: If the target key point is located inside the initial detection frame, the pixel vertical coordinate difference between the target key point and the reference key point is determined as a second offset value, and the second offset value is used as the vertical offset value; If the target key point is located outside the initial detection frame, the pixel vertical coordinate difference between the bottom frame line of the initial frame line and the target key point is determined as the third offset value, and the sum of the second offset value and the third offset value is determined as the vertical offset value. The bottom frame line of the initial detection frame is moved along the longitudinal axis by the longitudinal offset value, and the lengths of the left and right frame lines in the initial detection frame are adjusted to connect with the moved bottom frame line to obtain the target detection frame.
8. The method for positioning people based on detection frame correction according to claim 1, characterized in that: The step of converting the pixel coordinates to be located into a world coordinate system based on a preset target homography matrix to obtain target location information corresponding to the person to be located includes: Homogenize the pixel coordinates to be located and map them to the world coordinate system using a preset target homography matrix to obtain candidate positioning coordinates, wherein the target homography matrix is a mapping relationship constructed by the pixel coordinates of multiple ground contact points and the actual coordinates; The candidate positioning coordinates are normalized to obtain target positioning information corresponding to the person to be located.
9. A personnel positioning device based on detection frame correction, characterized in that: The personnel positioning device based on detection frame correction includes: The detection module is used to detect the person to be located in the image data using a preset detection model to obtain an initial detection frame and multiple key points of the human body; a determination module, configured to determine, based on key points of the human body located within the initial detection frame, whether the initial detection frame satisfies a preset positioning reference constraint condition, wherein the positioning reference constraint condition at least includes that neither the left frame line nor the right frame line of the initial detection frame exhibits unilateral offset; a correction module configured to select, if the left or right frame line has a unilateral offset, a shoulder joint point on the same side as the offset side as a target key point; determine a target offset value based on a distance between the target key point and a corresponding frame line in the initial detection frame, and move the corresponding frame line according to the target offset value and correct the detection frame so that the corrected target detection frame satisfies the positioning reference constraint condition; The positioning module is used to determine the pixel coordinates of the midpoint of the bottom frame line of the target detection frame as the pixel coordinates to be positioned, and convert the pixel coordinates to be positioned into the world coordinate system based on a preset target homography matrix to obtain the target positioning information corresponding to the person to be positioned.
10. A personnel positioning device based on detection frame correction, characterized in that: The personnel positioning device based on detection frame correction includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instructions in the memory to enable the personnel positioning device based on detection frame correction to execute the personnel positioning method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Attention point mechanism-based locating loss calculation method and system for target detection system
CN108205687A
Safe wearing detection method based on human body key points
CN112560741A
Multi-pedestrian tracking method based on joint global and local features
CN114120188A
Indoor personnel positioning method and device based on image recognition and medium
CN115797445A
Cited By
Data processing method and system applied to multi-target detection
CN121437864A
Safety early warning method, safety early warning device and safety early warning system
CN121583046A