Personnel positioning method, device and equipment based on monocular vision and storage medium

By collecting images on a monocular camera and using polynomial regression and the principle of similar triangles to correct depth estimation errors, combined with the coordinate transformation matrix, low-cost, efficient and high-precision personnel positioning is achieved, solving the problems of high cost and low precision of traditional monocular vision positioning.

CN120765724APending Publication Date: 2025-10-10ZHUHAI UNITECH POWER TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510823333.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Traditional monocular vision-based personnel positioning technology is costly, has low accuracy, and is complex to deploy. The depth detection model has poor generalization, making it difficult to promote in actual industrial scenarios.

Method used

By collecting the correction position and working position images of the monocular camera, the depth information of the pixel points is determined using the depth detection model, and the depth estimation error is corrected in combination with the polynomial regression equation. The mapping relationship of the pixel coordinates in the camera reference coordinate system is established, the focal length projection distance is calculated using the principle of similar triangles, and multi-angle positioning is achieved through the coordinate system transformation matrix.

Benefits of technology

It reduces hardware deployment costs, improves positioning accuracy, reduces the demand for model training samples, avoids overfitting problems, and achieves low-cost, efficient, and high-precision personnel positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765724A_ABST
    Figure CN120765724A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a personnel positioning method and device based on monocular vision, equipment and a storage medium, which are used for reducing the hardware deployment cost of personnel positioning and improving the personnel positioning accuracy. The monocular vision-based personnel positioning method comprises the following steps of: estimating depth information by utilizing a depth detection model, and determining an actual vertical distance from a ground contact point to a monocular camera and an actual horizontal distance of a horizontal point pair; fitting the depth information and the actual vertical distance through a polynomial regression equation so as to determine a vertical coordinate mapping relation of the pixel coordinates in a camera reference coordinate system; determining a horizontal coordinate mapping relation according to the actual horizontal distance of the horizontal point pair and the pixel pitch of the horizontal point pair; and determining target positioning information through a coordinate system conversion matrix between the working position and the correction position in combination with the vertical coordinate mapping relation, the horizontal coordinate mapping relation and the relative position information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a personnel positioning method, apparatus, device and storage medium based on monocular vision. Background Art

[0002] Computer vision-based monocular vision-based personnel positioning is one of the core technologies in the field of computer vision and intelligent perception. It aims to detect, identify and determine the location information of people in the scene in real time through image or video data. Especially in scenarios such as substations, accurate positioning of personnel is extremely important for substation operation and maintenance.

[0003] Among related technologies, traditional vision-based personnel positioning based on monocular vision mainly relies on multi-sensor fusion technology, such as combining depth cameras and binocular cameras to enhance three-dimensional spatial positioning capabilities. However, the cost of deploying such cameras is high, making it difficult to promote in actual industrial scene operations; the monocular camera solution that relies on depth estimation, due to the generalization problem of the depth detection model, needs to collect a large amount of data for each scene and conduct model training to improve its detection accuracy, which increases the workload of the debugger. When the number of training samples is small, the depth detection model is prone to overfitting, and large deviations are likely to occur during target inference, resulting in low accuracy of personnel positioning based on monocular vision. Summary of the Invention

[0004] The present application provides a personnel positioning method, apparatus, device and storage medium based on monocular vision, which are used to reduce the hardware deployment cost of personnel positioning and improve the accuracy of personnel positioning based on monocular vision.

[0005] A first aspect of the present application provides a personnel positioning method based on monocular vision, comprising: acquiring a calibration position image when a monocular camera is in a calibration position and a working position image when the monocular camera is in a working position, and determining relative position information of the monocular camera in a target area;

[0006] Determining the depth information corresponding to each pixel in the corrected image based on a preset depth detection model, and determining the actual vertical distances from multiple ground contact points in the corrected image to the monocular camera, as well as the actual horizontal distances between pairs of horizontal points;

[0007] Fitting the correspondence between the depth information of the multiple ground contact points and the actual vertical distances through a polynomial regression equation to correct the depth estimation error of the depth detection model, and determining the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system based on the depth detection model and the fitted polynomial regression equation;

[0008] The horizontal point pair and the monocular camera are used to construct a space triangle and an imaging triangle, a focal length projection distance is calculated based on a similar triangle proportion relationship and an actual horizontal distance of the horizontal point pair, and a horizontal coordinate mapping relationship of a pixel coordinate in a camera reference coordinate system is determined based on the focal length projection distance, wherein the focal length projection distance is a conversion proportion factor of a pixel distance and an actual distance;

[0009] The initial pixel coordinates of the personnel to be positioned in the working position image are converted to a pixel coordinate system in the correction position based on a coordinate system conversion matrix between the working position and the correction position, to obtain correction position pixel coordinates;

[0010] The target positioning information of the personnel to be positioned in a world coordinate system is determined according to the correction position pixel coordinates, the vertical coordinate mapping relationship, the horizontal coordinate mapping relationship and the relative position information.

[0011] A second aspect of the present application provides a personnel positioning device based on monocular vision, comprising: a collection module configured to collect a correction position image when a monocular camera is in a correction position and a working position image when the monocular camera is in a working position, and determine relative position information of the monocular camera in a target region;

[0012] A determination module is configured to determine depth information corresponding to each pixel point in the correction position image based on a preset depth detection model, and determine actual vertical distances of a plurality of ground contact points to the monocular camera and actual horizontal distances of a horizontal point pair;

[0013] A first construction module is configured to fit a corresponding relationship between depth information and actual vertical distances of the plurality of ground contact points by a polynomial regression equation, to correct a depth estimation error of the depth detection model, and determine a vertical coordinate mapping relationship of a pixel coordinate in a camera reference coordinate system based on the depth detection model and the fitted polynomial regression equation;

[0014] A second construction module is configured to construct a space triangle and an imaging triangle from the horizontal point pair and the monocular camera, calculate a focal length projection distance based on a similar triangle proportion relationship and an actual horizontal distance of the horizontal point pair, and determine a horizontal coordinate mapping relationship of a pixel coordinate in a camera reference coordinate system based on the focal length projection distance, wherein the focal length projection distance is a conversion proportion factor of a pixel distance and an actual distance;

[0015] A conversion module is configured to convert initial pixel coordinates of the personnel to be positioned in the working position image to a pixel coordinate system in the correction position based on a coordinate system conversion matrix between the working position and the correction position, to obtain correction position pixel coordinates;

[0016] A positioning module is used to determine the target positioning information of the person to be positioned in the world coordinate system based on the corrected pixel coordinates, the vertical coordinate mapping relationship, the horizontal coordinate mapping relationship and the relative position information.

[0017] A third aspect of the present application provides a monocular vision-based personnel positioning device, comprising: a memory and at least one processor, wherein the memory stores instructions;

[0018] The at least one processor calls the instructions in the memory to enable the monocular vision-based personnel positioning device to execute the above-mentioned monocular vision-based personnel positioning method.

[0019] A fourth aspect of the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned personnel positioning method based on monocular vision.

[0020] In the technical solution provided by the present application, the depth information in the correction position image of the monocular camera in the correction position is estimated by the depth detection model, thereby reducing the increase in deployment costs caused by the traditional positioning solution relying on binocular cameras or depth cameras to provide accurate depth information. The present technical solution can realize personnel positioning based on monocular vision through a monocular camera, thereby reducing the hardware deployment cost of personnel positioning based on monocular vision; there is a depth error when the depth detection model is used to estimate the depth of the monocular camera image. The present technical solution performs real distance correction on the depth information based on the actual distance parameter of the ground contact point at the correction position, establishes the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, and reduces the accuracy requirements of the depth detection model. The present technical solution does not need to use a large number of training samples for model training, and also avoids the depth detection model from appearing when the number of training samples is small. Overfitting reduces the workload of the debugger. The depth information predicted by the depth detection model is corrected based on the objective function relationship, which can achieve efficient and accurate positioning in the vertical position relationship. Then, based on the principle of similar triangles, the focal length projection distance is calculated using the actual horizontal distance of the horizontal point pair and its pixel spacing to construct the horizontal coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, which can achieve accurate positioning in the horizontal position relationship. Furthermore, the coordinate transformation matrix between the working position and the correction position is combined to realize the conversion between the actual working posture and the correction posture of the monocular camera, which can realize the multi-angle personnel positioning function based on monocular vision of the camera. Finally, the relative position information of the monocular camera in the target area is combined to determine the target positioning information of the person to be positioned in the world coordinate system, which can meet the high-precision positioning requirements of the target area with low cost and high deployment efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a schematic diagram of an embodiment of a method for positioning a person based on monocular vision in this application;

[0022] Figure 2 Schematic diagram of calculating focal length projection distance based on the principle of similar triangles in this application;

[0023] Figure 3 Schematic diagram of the position information of the monocular camera in the target area in this application;

[0024] Figure 4 This is a schematic diagram of another embodiment of the personnel positioning method based on monocular vision in this application;

[0025] Figure 5 This is a schematic diagram of an embodiment of a personnel positioning device based on monocular vision in this application;

[0026] Figure 6 This is a schematic diagram of another embodiment of a personnel positioning device based on monocular vision in this application;

[0027] Figure 7 This is a schematic diagram of an embodiment of a personnel positioning device based on monocular vision in this application. DETAILED DESCRIPTION

[0028] The present application provides a personnel positioning method, apparatus, device and storage medium based on monocular vision, which are used to solve the problems of high cost, low accuracy and complex deployment of traditional positioning technology.

[0029] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] See also Figure 1 , an embodiment of a personnel positioning method based on monocular vision in this application includes:

[0031] 101. Collect a calibration position image of the monocular camera when it is in the calibration position and a working position image when it is in the working position, and determine relative position information of the monocular camera in the target area.

[0032] It is understandable that the execution subject of this application can be a monocular vision-based personnel positioning device, or a terminal, a monitoring system or a server, which is not limited here. This embodiment is described by taking the monitoring system as the execution subject as an example.

[0033] In this embodiment, the correction position is the correction posture of the monocular camera, which can be understood as the monitoring angle in the correction state. It can be represented by the relative position between the camera optical axis and the reference plane. For example, when the monocular camera is in the correction position, the projection of the camera optical axis on the ground is perpendicular to the target wall.

[0034] In this embodiment, the working position refers to the actual working posture of the monocular camera, which can be understood as the actual monitoring angle, or can be determined by PTZ parameters or represented by the deflection angle of the camera's optical axis. Typically, when the target camera is in the working position, the projection of the camera's optical axis on the ground is not perpendicular to the target wall. Furthermore, the working position can be adjusted in practice. Monocular vision-based personnel positioning can be achieved by simply reconstructing the coordinate system transformation matrix between the adjusted working position and the corrected position. This embodiment can meet the requirements of monocular vision-based personnel positioning for multiple camera angles.

[0035] It can be understood that the above-mentioned target wall can be any wall that meets the requirements of the working scene. It can be assumed that the working scene is a rectangular space. Generally speaking, the wall refers to the wall closest to the camera. The target wall will only affect the final mapping of the personnel position from the camera coordinates to the world coordinates. If the camera is installed in the center of the room or other position not against the wall, the correction position can also be determined by other means, such as adjusting the camera optical axis to be parallel to the ground or other reference horizontal plane.

[0036] The above-mentioned camera optical axis is an imaginary straight line that describes the propagation path of the light beam in the optical system. It is usually used as the symmetry axis of the optical element. The optical axis is perpendicular to the camera imaging plane, and the intersection with the image plane is defined as the origin of the image coordinate system. This geometric relationship is the basis of the camera calibration and imaging model. In this embodiment, the camera is first adjusted to the correction position, that is, the projection of the camera optical axis on the ground and the calibration position perpendicular to the wall. For example, a monocular camera is usually installed in a corner against the wall to obtain a larger field of view. The angle between the camera optical axis and the wall on which it is installed is ninety degrees, that is, the correction position.

[0037] The above PTZ (Pan-Tilt-Zoom) is used to indicate the camera posture parameters, which may include the camera's horizontal rotation (Pan) parameter, vertical rotation (Tilt) parameter, and zoom (Zoom) parameter.

[0038] The above-mentioned horizontal rotation parameters are used to indicate the angle range of horizontal rotation of the camera, usually ±170° to 360°; the vertical rotation parameters are used to indicate the angle range of vertical rotation of the camera, usually -90 degrees to +90 degrees; and the zoom parameters can include optical zoom achieved through the optical structure of the lens, and can also include zoom achieved through image processing technology.

[0039] In this embodiment, a monocular camera is used to indicate an image acquisition module that captures two-dimensional images through a single lens. The monocular camera has an image acquisition function, is low in cost, simple in structure, low in power consumption, small in size, and easy to deploy, but has shortcomings in obtaining depth information.

[0040] The above-mentioned correction position image and working position image are used to indicate the images collected when the camera is in different positions. It can be understood that in order to better express the concept, the working position image of this embodiment includes the person to be located. In actual applications, the construction of the coordinate system conversion matrix and the positioning of the person to be located can also be carried out separately, that is, as long as they are in the same working position, the coordinate system conversion matrix can be constructed first, and when the person to be located is identified, the positioning step can be executed.

[0041] In this embodiment, the target area is used to indicate the area monitored by the monocular camera, such as a room in a substation. The relative position information of the target area represents the relative position coordinates of the installation position of the monocular camera in the target area, which can be represented by the relative position of the optical center of the camera in the room. The relative position information can be converted into coordinate information in the world coordinate system.

[0042] 102. Determine the depth information corresponding to each pixel in the corrected image based on a preset depth detection model, and determine the actual vertical distances from multiple ground contact points in the corrected image to the monocular camera, as well as the actual horizontal distances between pairs of horizontal points.

[0043] In this embodiment, the depth detection model can be MiDaS, DORN, or other models suitable for monocular depth estimation, and this embodiment does not impose any specific limitations.

[0044] In this embodiment, the ground contact point indicates the pixel point in the image that is in direct contact with the ground. It can also be called a ground point or ground pixel point. In this embodiment, the number of ground contact points is any number that satisfies the objective function fitting requirements. The actual vertical distance indicates the actual distance between the point and the location of the monocular camera, that is, the length of the projection of the line segment between the point and the optical center of the monocular camera on the ground.

[0045] The above horizontal point pair is used to indicate the horizontal position of two points relative to the camera (perpendicular to the camera optical axis), such as Figure 2The two horizontal points (x1, y1) and (x2, y2) shown have a small difference between y1 and y2, that is, the vertical distance d between the horizontal point pair and the camera is close, that is, the depth information d1 and d2 have a small difference. The horizontal point pair can be two ground contact points or other points, without specific limitation. The actual horizontal distance is used to indicate the actual distance in the horizontal direction between the horizontal point pair, and can be expressed by the absolute value of L = (x1-x2).

[0046] It should be understood that achieving accurate three-dimensional positioning from a two-dimensional image requires constructing a mapping between pixel coordinates and the camera reference coordinate system. That is, it is necessary to determine the vertical coordinate mapping relationship and horizontal coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, that is, the relative position relationship of each pixel point compared to the monocular camera reference system. Among them, the camera reference system refers to the reference system with the optical center of the camera as the origin and the optical axis as the Z axis.

[0047] 103. The correspondence between the depth information of multiple ground contact points and the actual vertical distance is fitted through a polynomial regression equation to correct the depth estimation error of the depth detection model, and the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system is determined based on the depth detection model and the fitted polynomial regression equation.

[0048] In this embodiment, a polynomial regression equation is used to fit a mathematical relationship or mathematical model between depth information and actual vertical distance, which can be determined based on multiple ground contact points.

[0049] It should be understood that in actual applications, although the depth detection model can provide depth information of each pixel to achieve three-dimensional positioning and reduce the demand for camera hardware, there are still errors in the depth detection based on the monocular camera. To ensure its accuracy, it is usually necessary to pre-train for the corresponding camera model and the corresponding working angle to solve the problem of large errors in the depth estimation of the depth monitoring model in different camera images. However, due to the generalization problem of the model, inaccurate detection problems will be encountered. Some models will use self-supervised learning, and some models need to re-measure the scene depth for transfer learning, and collecting data and training models often take more time.

[0050] In order to solve the above problems, this embodiment corrects the depth error based on the actual vertical distance of each ground contact point. It only needs to measure the real data of several ground contact points to directly construct a mapping relationship about the depth, which can improve the accuracy of the vertical coordinate mapping. Even if the depth relationship of the model is not accurate, the corresponding vertical coordinates can be obtained through the measured ground feature point information, thereby improving the accuracy of the vertical coordinates in personnel positioning based on monocular vision.

[0051] Compared with the solution for improving the accuracy of the model architecture itself, this embodiment directly corrects the final mapping relationship, so the requirements for the depth estimation accuracy of the model itself are not high, which will greatly improve the vertical coordinate positioning efficiency of this embodiment. Compared with the optimization solution for the model training process, in the application scenario with a small number of training samples (such as about 20 training points), the training of the depth detection model is very easy to overfit, and large deviations are prone to occur during target reasoning. The construction of the vertical coordinate mapping relationship of this solution is based on direct verification based on the objective function. Due to the simplicity of the model, there will be no overfitting problem. Traditional training solutions usually obtain more training points (such as increasing from 20 to 50) to avoid the impact of overfitting. This embodiment can reduce the workload of debuggers and improve the training efficiency of the model. There is even no need to train the depth detection model separately for each scene. Compared with other positioning solutions, such as those based on ultra-wideband (UWB) technology or 3D modeling technology, this embodiment is based on pure vision and only requires a monocular camera to complete personnel positioning based on monocular vision. It has more advantages than other conventional technical solutions in terms of deployment cost and convenience.

[0052] Specifically, multiple ground contact points are selected in the corrected bitmap image; the objective function relationship is fitted based on the depth information of each ground contact point and its corresponding actual vertical distance, and the regression coefficients of the polynomial regression equation are determined by minimizing the error between the predicted distance and the actual vertical distance.

[0053] 104. Construct a spatial triangle and an imaging triangle using the horizontal point pair and the monocular camera. Calculate the focal length projection distance based on the proportional relationship of similar triangles and the actual horizontal distance of the horizontal point pair. Determine the horizontal coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system based on the focal length projection distance.

[0054] See also Figure 2 The horizontal point pair is described, the actual horizontal distance L of the two horizontal points, the optical center O of the monocular camera, that is, the two horizontal points and the monocular camera form a spatial triangle, wherein the dotted line is the optical axis, and the pixel spacing of the horizontal point pair on the imaging plane of the monocular camera, that is, the two pixel points corresponding to the two horizontal points on the imaging plane and the optical center form an imaging triangle, and d is the mean vertical distance, that is, the average value of the actual vertical distance between the two horizontal points and the monocular camera. In this embodiment, by constructing two similar triangles, the spatial triangle and the imaging triangle, the focal length projection distance f of the camera on the ground can be calculated based on the principle of similar triangles:

[0055]

[0056] After obtaining the projection distance f of the camera focal length on the ground, for a given detection pixel point (u, v), in the camera reference system with a resolution of (H, W), its horizontal coordinate x can be obtained as:

[0057]

[0058] The target vertical distance d can be used to determine the mean vertical distance of the horizontal point pair through the vertical coordinate mapping relationship, so as to estimate the horizontal coordinate based on the corrected vertical coordinate and improve the overall positioning accuracy.

[0059] 105. Based on the coordinate system conversion matrix between the working position and the correction position, the initial pixel coordinates of the person to be located in the working position image are converted to the pixel coordinate system at the correction position to obtain the correction position pixel coordinates.

[0060] Specifically, the initial pixel coordinates of the person to be located in the work position image are determined through a preset target detection network; based on the coordinate system conversion matrix between the work position and the correction position, the initial pixel coordinates are converted to the pixel coordinate system at the correction position to obtain the pixel coordinates of the correction position. This embodiment can obtain the initial pixel coordinates of the person to be located when the monocular camera monitors the person, and convert them based on the coordinate system conversion matrix between the work position and the correction position to determine the pixel coordinates of the correction position. Accurate positioning is performed based on the vertical coordinate mapping relationship and the horizontal coordinate mapping relationship of the correction position, avoiding the need to remodel each work position and realizing the multi-view personnel positioning requirements. The vertical characteristic of the optical axis of the correction position can reduce the perspective distortion of the image and improve the mapping accuracy of the pixel coordinates to the physical coordinates.

[0061] It should be understood that in actual application scenarios, the camera may not be in the corrected position due to installation errors, installation requirements of the actual scene (such as adjusting the camera angle to achieve panoramic monitoring) and rotation requirements when the camera is working. In order to adapt to the multi-perspective personnel positioning requirements, this embodiment uses a coordinate system transformation matrix to achieve coordinate conversion at different preset positions.

[0062] In this embodiment, the coordinate system conversion matrix may be a homography matrix constructed based on feature point matching, or a rotation matrix constructed based on a camera.

[0063] Taking the rotation matrix as an example of a coordinate system conversion matrix, a rotation matrix can be constructed based on the PTZ parameters of the monocular camera at its working position, and the rotation matrix can be determined as a coordinate system conversion matrix. The rotation matrix can convert the coordinate systems under different perspectives into a certain reference coordinate system.

[0064] The three-dimensional rotation matrix calculated by the horizontal angle and vertical angle in the PTZ parameters is used to transform the working position coordinate system at the working position to the correction position coordinate system at the correction position.

[0065] In this embodiment, the initial pixel coordinates of the person to be located can be detected by a target detection network such as Faster-RCNN, YOLO, etc. This embodiment does not limit the specific architecture of the target detection network.

[0066] In this embodiment, the position of the person to be located can be represented by the head, torso or feet of the person to be located and other target positions. For example, the pixel points of the feet can be used as the initial pixel coordinates. For example, when locating the head, the positioning can be further corrected by combining camera distortion and human skeleton information. There is no specific limitation.

[0067] For example, in some scenarios, the person to be located may be partially monitored by a monocular camera due to occlusion or other factors, such as only monitoring the head or only monitoring one side of the foot. In this embodiment, the body skeleton key point detection model (such as OpenPose, MediaPipe) can be combined to identify key points such as the head, shoulders, hips, and knees. Even if the feet are blocked, the initial pixel coordinates of the person to be located can be inferred from the position of the head or hip. If only one side of the foot is detected, the position of the other side can be inferred through the gait symmetry model to improve positioning accuracy through the pixel coordinates of both feet, or positioning can also be achieved by relying on the pixel coordinates of one side of the foot.

[0068] 106. Determine target positioning information of the person to be positioned in the world coordinate system based on the corrected pixel coordinates, the vertical coordinate mapping relationship, the horizontal coordinate mapping relationship, and the relative position information.

[0069] Specifically, the candidate vertical coordinates of the person to be located are obtained according to the mapping relationship between the corrected pixel coordinates and the vertical coordinates, and the candidate horizontal coordinates of the person to be located are determined according to the mapping relationship between the target vertical coordinates and the horizontal coordinates; the target positioning information of the person to be located in the world coordinate system is determined according to the candidate vertical coordinates, the candidate horizontal coordinates and the relative position information of the monocular camera in the target area.

[0070] The candidate vertical coordinates are used to indicate the relative vertical distance between the person to be located and the monocular camera; the candidate horizontal coordinates are used to indicate the relative horizontal distance between the person to be located and the monocular camera, that is, the horizontal offset distance relative to the camera optical axis.

[0071] For example, the target detection network detects the position of a person's feet, takes the center pixel position of both feet (u0, v0) and converts it into the corrected pixel coordinates (u1, v1). First, the depth position d1 is obtained through the depth detection model, and then the y value of the pixel point in the camera reference system is obtained through the fitted objective function relationship (i.e., the vertical coordinate mapping relationship), that is, the candidate vertical coordinate; based on the similar triangle and the focal length projection distance, the x value of the pixel point in the camera reference system is obtained, that is, the candidate horizontal coordinate.

[0072] Reference Figure 3 Schematic diagram of the position information of a monocular camera in a target area. In this embodiment, the position of the camera coordinates in the world coordinate system can be obtained by the relative position information of the camera in the target area (assuming it is a room). The optical center of the camera is at the coordinate (xc, yc) relative to the world coordinate system of the room. Therefore, for the camera coordinate (x0, y0, z), its world coordinate is (x0+xc, y0+yc, z). Therefore, through the above steps, the coordinates (u0, v0) of the person detected in the preset position image can be projected into the world coordinate system through a series of transformations, thereby realizing the position of the person.

[0073] It is understandable that since we only care about the position information of pixels on the ground, we set the third dimension to 0 in the camera reference frame.

[0074] In this embodiment, the depth detection model is used to estimate the depth information in the calibration image of the monocular camera in the calibration position, which reduces the deployment cost caused by the traditional positioning scheme relying on binocular cameras or depth cameras to provide accurate depth information. The technical solution can realize personnel positioning through a monocular camera, reducing the hardware deployment cost of personnel positioning. When the depth detection model is used for monocular camera image depth estimation, there is a depth error. The technical solution corrects the depth information based on the actual distance parameters of the ground contact points in the calibration position, establishes the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, reduces the accuracy requirements of the depth detection model, and avoids overfitting of the depth detection model in the case of a small number of training samples, thereby reducing the workload of the debuggers. Based on the target function relationship, the depth information predicted by the depth detection model is corrected, which can efficiently and accurately realize positioning in the vertical position relationship. Then, based on the principle of similar triangles, the actual horizontal distance of the horizontal point pair and the pixel interval are used to calculate the focal length projection distance, so as to construct the horizontal coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, which can realize accurate positioning in the horizontal position relationship. Further, the conversion between the actual working posture of the monocular camera and the calibration posture is realized by combining the coordinate conversion matrix between the working position and the calibration position, which can realize the personnel positioning function of the camera in multiple angles. Finally, the target positioning information of the personnel to be positioned in the world coordinate system is determined by combining the relative position information of the monocular camera in the target area, which can meet the demand of low-cost, high-deployment-efficiency target area high-precision positioning.

[0075] Please refer to Figure 4 Another embodiment of the personnel positioning method based on monocular vision in the present application includes:

[0076] 401. Collect the calibration image when the monocular camera is in the calibration position and the working image when the monocular camera is in the working position, and determine the relative position information of the monocular camera in the target area.

[0077] 402. Determine the depth information corresponding to each pixel point in the calibration image based on the pre-set depth detection model, and determine the actual vertical distance from the calibration image to the monocular camera, and the actual horizontal distance of the horizontal point pair.

[0078] Steps 401-402 can be performed with reference to steps 101-102, which will not be described here.

[0079] 403. With the depth information as the independent variable and the corresponding actual vertical distance as the dependent variable, fit a plurality of ground contact points by a polynomial regression equation and determine the regression coefficients in the polynomial regression equation to obtain an initial depth correction equation.

[0080] Specifically, the depth information of the multiple ground contact points and the corresponding actual vertical distances are selected as training samples, and each regression coefficient is optimized by a gradient descent algorithm to minimize the mean square error between the vertical distance predicted by the polynomial and the actual vertical distance, to obtain an initial depth correction equation.

[0081] φ(d)=y'=arg min β ||y-β[d 2 ,d,1] T || 2

[0082] wherein φ(d) represents the initial depth correction equation, d represents the depth information; y' represents the predicted vertical distance, y represents the actual vertical distance; β represents a vector of the polynomial regression coefficients; and T represents vector transposition.

[0083] The above ‖·‖ represents an L2 norm, which can be a Euclidean distance, used to calculate the mean square error. For example, assuming β=[β0,β1,β2], then β[d 2 ,d,1] T =β0*d 2 +β1*d+β2.

[0084] It can be understood that the determination of the initial depth correction equation is based on the minimization of the mean square error, to enhance the correction effect of the initial depth correction equation.

[0085] Optionally, each regression coefficient in the polynomial regression equation is determined according to the actual vertical distance of each ground contact point and the corresponding depth information, to obtain the initial depth correction equation, including: taking the depth information of each ground contact point predicted by the depth detection model as the independent variable and the corresponding actual vertical distance as the dependent variable, performing polynomial regression equation fitting based on the least square method, solving each regression coefficient of the polynomial regression equation by matrix operation, and obtaining the initial depth correction equation.

[0086] 404, verifying the initial depth correction equation by a preset number of test points, and determining the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system based on the depth detection model and the successfully verified depth correction equation.

[0087] Specifically, a plurality of ground contact points for verification (i.e., ground points not used to fit the initial depth correction equation) are selected and substituted into the determined initial depth correction equation to determine whether the error is less than a preset threshold, and if so, the verification is successful. When the initial depth correction equation fails to pass the verification, the regression coefficients are adjusted until the adjusted depth correction equation passes the verification; when the verification is successful, the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system is determined based on the depth detection model and the successfully verified depth correction equation.

[0088] 405. Based on the principle of similar triangles, the focal length projection distance is calculated using the actual horizontal distance of the horizontal point pair and its pixel spacing, and the horizontal coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system is determined.

[0089] Optionally, the mean vertical distance of the horizontal point pairs is determined according to the vertical coordinate mapping relationship; based on the principle of similar triangles, the focal length projection distance of the camera focal length on the ground is determined according to the ratio between the mean vertical distance and the actual horizontal distance, and the product of the horizontal pixel spacing of the horizontal point pairs in the corrected bit image; the horizontal coordinate mapping relationship is constructed according to the preset camera resolution, focal length projection distance and actual vertical distance.

[0090] For example, the expression of horizontal coordinate mapping relationship is:

[0091]

[0092] in, represents the horizontal coordinate mapping relationship of the pixel coordinate in the camera reference coordinate system, u represents the horizontal coordinate of the pixel to be predicted, y' represents the vertical distance predicted by the vertical coordinate mapping relationship, x' represents the predicted horizontal distance, W represents the width of the monocular camera resolution, and f represents the focal length projection distance.

[0093] 406. Based on the coordinate system conversion matrix between the working position and the correction position, the initial pixel coordinates of the person to be located in the working position image are converted to the pixel coordinate system at the correction position to obtain the correction position pixel coordinates.

[0094] Specifically, a coordinate system conversion matrix is ​​constructed between the working position and the correction position; the initial pixel coordinates of the person to be located in the working position image are determined through a preset target detection network; and the initial pixel coordinates are converted to the pixel coordinate system at the correction position based on the coordinate system conversion matrix to obtain the pixel coordinates of the correction position.

[0095] In this embodiment, one working position corresponds to one coordinate system transformation matrix. The working position of the monocular camera can be fixed or not in actual applications, that is, the monocular camera has one or more actual working postures. It is only necessary to construct a coordinate transformation matrix between each working position and the correction position to eliminate the perspective distortion and coordinate offset caused by the change of the camera angle, ensure that all positioning calculations are based on a consistent geometric reference, avoid re-modeling for each working position, and reduce the amount of real-time calculations.

[0096] Optionally, constructing a coordinate system conversion matrix between the working position and the correction position includes: constructing a homography matrix based on the correction position image and the working position image through a preset feature point matching algorithm, and determining the homography matrix as the coordinate system conversion matrix. In this embodiment, the homography matrix can realize the conversion between the working position and the correction position. For example, when the camera captures the position of the person and obtains its pixel coordinates (u0, v0), the pixel coordinates of the coordinates in the camera correction position (u1, v1) are obtained through the homography matrix (HM).

[0097] The above-mentioned homography matrix can be constructed by matching key points between the correction position image and the working position image, wherein the matching method can adopt the Scale-Invariant Feature Transform (SIFT) algorithm or the Oriented FAST and Rotated BRIEF (ORB) algorithm. Specifically, the corresponding feature points of the correction position image and the working position image are associated through the feature point matching algorithm; the homography matrix is ​​calculated according to the matching results to realize the coordinate system conversion.

[0098] Optionally, constructing a coordinate system conversion matrix between the working position and the correction position includes: constructing a rotation matrix according to PTZ parameters of the monocular camera at the working position, and determining the rotation matrix as the coordinate system conversion matrix.

[0099] Optionally, the initial pixel coordinates of the person to be located in the work position image are determined using a preset target detection network, including: detecting the foot position of the person to be located in the work position image using the preset target detection network to obtain the pixel coordinates of both feet; and determining the initial pixel coordinates of the person to be located based on the center point of the pixel coordinates of both feet. The vertical coordinate mapping relationship of this embodiment is constructed based on the ground contact point, and combining the center point of the pixel coordinates of both feet as the initial pixel coordinate can further improve positioning accuracy.

[0100] It should be noted that the initial pixel coordinates of this embodiment can also be anchored based on key positions of the skeleton such as the head and torso under appropriate circumstances, and this embodiment does not impose any specific limitation.

[0101] 407. Determine target positioning information of the person to be positioned in the world coordinate system based on the corrected pixel coordinates, the vertical coordinate mapping relationship, the horizontal coordinate mapping relationship, and the relative position information.

[0102] Specifically, based on the target mapping relationship, the corrected pixel coordinates are converted to the camera reference coordinate system to obtain the reference coordinate value; and the target positioning information of the person to be positioned in the world coordinate system is determined according to the reference coordinate value and the relative position information.

[0103] In this embodiment, the depth information estimation in the calibration position image of the monocular camera in the calibration position by the depth detection model reduces the increase in deployment cost caused by the traditional positioning scheme relying on the binocular camera or the depth camera to provide accurate depth information. The technical solution can realize personnel positioning through a monocular camera, thereby reducing the hardware deployment cost of personnel positioning. When the monocular camera image depth estimation is performed on the depth detection model, there is a depth error. The technical solution fits and verifies the depth information through a polynomial regression equation for the actual distance parameters of multiple ground contact points, establishes the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, reduces the accuracy requirement of the depth detection model, and avoids overfitting of the depth detection model in the case of a small number of training samples, thereby reducing the workload of the debuggers. The depth information predicted by the depth detection model is corrected based on the target function relationship, and the positioning in the vertical position relationship can be efficiently and accurately realized. Then, based on the principle of similar triangles, the horizontal distance of the horizontal point pair and the pixel interval are used to calculate the focal length projection distance, so as to construct the horizontal coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, and the accurate positioning in the horizontal position relationship can be realized. Further, the conversion between the actual working posture of the monocular camera and the calibration posture is realized by combining the coordinate conversion matrix between the working position and the calibration position, the multi-angle personnel positioning function of the camera can be realized, and finally the target positioning information of the personnel to be positioned in the world coordinate system is determined by combining the relative position information of the monocular camera in the target area, so that the high-precision positioning demand of the target area with low cost and high deployment efficiency can be met.

[0104] The personnel positioning method based on monocular vision in the present application is described above, and the personnel positioning device based on monocular vision in the present application is described below. Please refer to Figure 5 An embodiment of the personnel positioning device based on monocular vision in the present application includes:

[0105] The acquisition module 501 is configured to acquire the calibration position image when the monocular camera is in the calibration position and the working position image when the monocular camera is in the working position, and determine the relative position information of the monocular camera in the target area;

[0106] The determination module 502 is configured to determine the depth information corresponding to each pixel point in the calibration position image based on the preset depth detection model, and determine the actual vertical distance from the multiple ground contact points in the calibration position image to the monocular camera and the actual horizontal distance of the horizontal point pair;

[0107] A first construction module 503 is configured to fit the correspondence between the depth information of multiple ground contact points and the actual vertical distance using a polynomial regression equation to correct the depth estimation error of the depth detection model, and determine the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system based on the depth detection model and the fitted polynomial regression equation;

[0108] A second construction module 504 is configured to construct a spatial triangle and an imaging triangle using the horizontal point pair and the monocular camera, calculate a focal projection distance based on a proportional relationship between similar triangles and an actual horizontal distance between the horizontal point pair, and determine a horizontal coordinate mapping relationship between pixel coordinates in the camera reference coordinate system based on the focal projection distance, wherein the focal projection distance is a conversion factor between pixel distance and actual distance;

[0109] The conversion module 505 is used to convert the initial pixel coordinates of the person to be located in the working position image into the pixel coordinate system at the correction position based on the coordinate system conversion matrix between the working position and the correction position to obtain the pixel coordinates of the correction position;

[0110] The positioning module 506 is used to determine the target positioning information of the person to be positioned in the world coordinate system according to the corrected pixel coordinates, the vertical coordinate mapping relationship, the horizontal coordinate mapping relationship and the relative position information.

[0111] In this embodiment, the depth information in the calibration image of the monocular camera in the calibration position is estimated by the depth detection model, which reduces the deployment cost caused by the traditional positioning scheme relying on binocular cameras or depth cameras to provide accurate depth information. The technical solution can realize personnel positioning through a monocular camera, reducing the hardware deployment cost of personnel positioning. When the depth detection model is used for monocular camera image depth estimation, there is a depth error. The technical solution corrects the depth information based on the actual distance parameters of the ground contact points in the calibration position, establishes the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, reduces the accuracy requirements of the depth detection model, and avoids overfitting of the depth detection model when the number of training samples is small, reducing the workload of the debuggers. Based on the target function relationship, the depth information predicted by the depth detection model is corrected, which can efficiently and accurately realize positioning in the vertical position relationship. Then, based on the principle of similar triangles, the focal length projection distance is calculated using the actual horizontal distance of the horizontal point pair and the pixel interval, to construct the horizontal coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, which can realize accurate positioning in the horizontal position relationship. Further, the conversion between the actual working posture of the monocular camera and the calibration posture is realized by combining the coordinate conversion matrix between the working position and the calibration position, which can realize the personnel positioning function of the monocular camera in multiple angles. Finally, the target positioning information of the personnel to be positioned in the world coordinate system is determined by combining the relative position information of the monocular camera in the target area, which can meet the demand of low-cost, high-deployment-efficiency target area high-precision positioning.

[0112] Please refer to Figure 6 Another embodiment of the personnel positioning device based on monocular vision in the present application includes:

[0113] The acquisition module 501 is configured to acquire the calibration image when the monocular camera is in the calibration position and the working image when the monocular camera is in the working position, and determine the relative position information of the monocular camera in the target area;

[0114] The determination module 502 is configured to determine the depth information corresponding to each pixel point in the calibration image based on the preset depth detection model, and determine the actual vertical distance from the multiple ground contact points in the calibration image to the monocular camera, and the actual horizontal distance of the horizontal point pair;

[0115] The first construction module 503 is configured to fit the corresponding relationship between the depth information and the actual vertical distance of the multiple ground contact points by a polynomial regression equation, to correct the depth estimation error of the depth detection model, and determine the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system based on the depth detection model and the fitted polynomial regression equation;

[0116] A second construction module 504 is configured to construct a spatial triangle and an imaging triangle using the horizontal point pair and the monocular camera, calculate a focal projection distance based on a proportional relationship between similar triangles and an actual horizontal distance between the horizontal point pair, and determine a horizontal coordinate mapping relationship between pixel coordinates in the camera reference coordinate system based on the focal projection distance, wherein the focal projection distance is a conversion factor between pixel distance and actual distance;

[0117] The conversion module 505 is used to convert the initial pixel coordinates of the person to be located in the working position image into the pixel coordinate system at the correction position based on the coordinate system conversion matrix between the working position and the correction position to obtain the pixel coordinates of the correction position;

[0118] The positioning module 506 is configured to determine the target positioning information of the person to be positioned in the world coordinate system according to the corrected pixel coordinates, the vertical coordinate mapping relationship, the horizontal coordinate mapping relationship, and the relative position information.

[0119] Optionally, the first building module 503 includes:

[0120] A fitting unit 5031 is configured to use the depth information as an independent variable and the corresponding actual vertical distance as a dependent variable, fit multiple ground contact points through a polynomial regression equation, and determine the regression coefficients of each term in the polynomial regression equation to obtain an initial depth correction equation;

[0121] A verification unit 5032, configured to verify the initial depth correction equation using a preset number of test points;

[0122] When the initial depth correction equation fails to be verified, the regression coefficients are adjusted until the adjusted depth correction equation is successfully verified;

[0123] When the verification is successful, the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system is determined based on the depth detection model and the successfully verified depth correction equation.

[0124] Optionally, the fitting unit 5031 is specifically configured to:

[0125] φ(d)=y'=argmin β ||y-β[d 2 ,d,1] T || 2

[0126] Where φ(d) represents the initial depth correction equation, d represents the depth information, y' represents the predicted vertical distance, y represents the actual vertical distance, β represents the vector of polynomial regression coefficients, and T represents the vector transpose.

[0127] Optionally, the second building module 504 is specifically configured to:

[0128]

[0129] in, represents the horizontal coordinate mapping relationship of the pixel coordinate in the camera reference coordinate system, u represents the horizontal coordinate of the pixel to be predicted, y' represents the vertical distance predicted by the vertical coordinate mapping relationship, x' represents the predicted horizontal distance, W represents the width of the monocular camera resolution, and f represents the focal length projection distance.

[0130] Optionally, the conversion module 505 includes:

[0131] The construction unit 5051 is used to: construct a coordinate system conversion matrix between the working position and the correction position;

[0132] The detection unit 5052 is used to determine the initial pixel coordinates of the person to be located in the work position image through a preset target detection network;

[0133] The conversion unit 5053 is used to convert the initial pixel coordinates into the pixel coordinate system at the correction position based on the coordinate system conversion matrix to obtain the corrected pixel coordinates.

[0134] Optionally, the constructing unit 5051 is specifically configured to construct a homography matrix according to the correction position image and the working position image by using a preset feature point matching algorithm, and determine the homography matrix as a coordinate system conversion matrix.

[0135] Optionally, the construction unit 5051 is specifically configured to construct a rotation matrix according to the PTZ parameters of the monocular camera at the working position, and determine the rotation matrix as a coordinate system conversion matrix.

[0136] Optionally, the detection unit 5052 is specifically configured to: detect the foot positions of the person to be located in the work position image through a preset target detection network to obtain pixel coordinates of both feet;

[0137] The initial pixel coordinates of the person to be located are determined based on the center point of the pixel coordinates of both feet.

[0138] In this embodiment, the depth information in the correction position image of the monocular camera in the correction position is estimated by the depth detection model, thereby reducing the increase in deployment costs caused by the traditional positioning solution relying on binocular cameras or depth cameras to provide accurate depth information. The present technical solution can realize personnel positioning based on monocular vision through a monocular camera, thereby reducing the hardware deployment cost of personnel positioning based on monocular vision; there is a depth error when the depth detection model is used to estimate the depth of the monocular camera image. The present technical solution fits and verifies the depth information with the actual distance parameters of multiple ground contact points through a polynomial regression equation, establishes the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, and reduces the accuracy requirements of the depth detection model. The present technical solution does not need to use a large number of training samples for model training, and also avoids the depth detection model when the number of training samples is small. Overfitting occurs, which reduces the workload of the debugger. The depth information predicted by the depth detection model is corrected based on the objective function relationship, which can achieve efficient and accurate positioning in the vertical position relationship; then, based on the principle of similar triangles, the focal length projection distance is calculated using the actual horizontal distance of the horizontal point pair and its pixel spacing to construct the horizontal coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system, which can achieve accurate positioning in the horizontal position relationship; further, the coordinate transformation matrix between the working position and the correction position is combined to realize the conversion between the actual working posture and the correction posture of the monocular camera, which can realize the multi-angle personnel positioning function of the camera, and finally, the relative position information of the monocular camera in the target area is combined to determine the target positioning information of the person to be positioned in the world coordinate system, which can meet the high-precision positioning requirements of the target area with low cost and high deployment efficiency.

[0139] above Figure 5 and Figure 6 The personnel positioning device based on monocular vision in the present application is described in detail from the perspective of modular functional entities. The personnel positioning device based on monocular vision in the present application is described in detail from the perspective of hardware processing.

[0140] See also Figure 7 As shown, the personnel positioning device based on monocular vision includes a processor 700 and a memory 701. The memory 701 stores machine executable instructions that can be executed by the processor 700. The processor 700 executes the machine executable instructions to implement the above-mentioned personnel positioning method based on monocular vision.

[0141] Furthermore, Figure 7 The monocular vision-based personnel positioning device shown further includes a bus 702 and a communication interface 703 , and the processor 700 , the communication interface 703 and the memory 701 are connected via the bus 702 .

[0142] Among them, the memory 701 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), for example, at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 703 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 702 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0143] The processor 700 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits or software instructions in the processor 700. The processor 700 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 701 , and the processor 700 reads the information in the memory 701 and completes the method steps of the aforementioned embodiment in combination with its hardware.

[0144] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, cause the computer to execute the steps of a personnel positioning method based on monocular vision.

[0145] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0146] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0147] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A personnel positioning method based on monocular vision, characterized in that: The personnel positioning method based on monocular vision includes: Collecting a calibration position image of the monocular camera when it is in the calibration position and a working position image when it is in the working position, and determining the relative position information of the monocular camera in the target area; Determining the depth information corresponding to each pixel in the corrected image based on a preset depth detection model, and determining the actual vertical distances from multiple ground contact points in the corrected image to the monocular camera, as well as the actual horizontal distances between pairs of horizontal points; Fitting the correspondence between the depth information of the multiple ground contact points and the actual vertical distances through a polynomial regression equation to correct the depth estimation error of the depth detection model, and determining the vertical coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system based on the depth detection model and the fitted polynomial regression equation; Constructing a spatial triangle and an imaging triangle using the horizontal point pair and the monocular camera, calculating a focal projection distance based on a proportional relationship of similar triangles and an actual horizontal distance of the horizontal point pair, and determining a horizontal coordinate mapping relationship of pixel coordinates in a camera reference coordinate system based on the focal projection distance, wherein the focal projection distance is a conversion scale factor between pixel distance and actual distance; Based on the coordinate system conversion matrix between the working position and the correction position, the initial pixel coordinates of the person to be located in the working position image are converted into the pixel coordinate system at the correction position to obtain the correction position pixel coordinates; The target positioning information of the person to be positioned in the world coordinate system is determined according to the corrected pixel coordinates, the vertical coordinate mapping relationship, the horizontal coordinate mapping relationship and the relative position information.

2. The personnel positioning method based on monocular vision according to claim 1, characterized in that: Fitting the correspondence between the depth information of the plurality of ground contact points and the actual vertical distances by a polynomial regression equation to correct a depth estimation error of the depth detection model, and determining a vertical coordinate mapping relationship between pixel coordinates in a camera reference coordinate system based on the depth detection model and the fitted polynomial regression equation, includes: Using the depth information as an independent variable and the corresponding actual vertical distance as a dependent variable, a polynomial regression equation is used to fit multiple ground contact points and the regression coefficients of each term in the polynomial regression equation are determined to obtain an initial depth correction equation; Verifying the initial depth correction equation through a preset number of test points; When the initial depth correction equation fails to be verified, adjusting each regression coefficient until the adjusted depth correction equation is successfully verified; When the verification is successful, a vertical coordinate mapping relationship between the pixel coordinates in the camera reference coordinate system is determined based on the depth detection model and the successfully verified depth correction equation.

3. The personnel positioning method based on monocular vision according to claim 1 or 2, characterized in that: The depth information is used as the independent variable and the corresponding actual vertical distance is used as the dependent variable. A polynomial regression equation is used to fit multiple ground contact points and the regression coefficients of each item in the polynomial regression equation are determined to obtain the initial depth correction equation: φ(d)=y'=argmin β ||y-β[d 2 ,d,1] T || 2 Where φ(d) represents the initial depth correction equation, d represents the depth information, y' represents the predicted vertical distance, y represents the actual vertical distance, β represents the vector of polynomial regression coefficients, and T represents the vector transpose.

4. The personnel positioning method based on monocular vision according to claim 1, characterized in that: The determining of the horizontal coordinate mapping relationship of the pixel coordinates in the camera reference coordinate system based on the focal length projection distance includes: in, represents the horizontal coordinate mapping relationship of the pixel coordinate in the camera reference coordinate system, u represents the pixel horizontal coordinate of the pixel to be predicted, y' represents the vertical distance predicted by the vertical coordinate mapping relationship, x' represents the predicted horizontal distance, W represents the width of the monocular camera resolution, and f represents the focal length projection distance.

5. The personnel positioning method based on monocular vision according to claim 1, characterized in that: The method of converting the initial pixel coordinates of the person to be positioned in the working position image into the pixel coordinate system at the correction position based on the coordinate system conversion matrix between the working position and the correction position to obtain the pixel coordinates of the correction position includes: Constructing a coordinate system transformation matrix between the working position and the correction position; Determining the initial pixel coordinates of the person to be located in the workstation image through a preset target detection network; The initial pixel coordinates are converted into a pixel coordinate system at the correction position based on the coordinate system conversion matrix to obtain the correction pixel coordinates.

6. The personnel positioning method based on monocular vision according to claim 5, characterized in that: The constructing of a coordinate system conversion matrix between the working position and the correction position includes: Constructing a homography matrix based on the correction position image and the working position image by a preset feature point matching algorithm, and determining the homography matrix as a coordinate system conversion matrix; or, A rotation matrix is ​​constructed according to the PTZ parameters of the monocular camera at the working position, and the rotation matrix is ​​determined as a coordinate system conversion matrix.

7. The personnel positioning method based on monocular vision according to claim 5, characterized in that: The determining of the initial pixel coordinates of the person to be located in the work position image by a preset target detection network includes: The preset target detection network is used to detect the foot position of the person to be located in the work position image and obtain the pixel coordinates of both feet; The center point of the pixel coordinates of both feet is determined as the initial pixel coordinates of the person to be located.

8. A personnel positioning device based on monocular vision, characterized in that: The personnel positioning device based on monocular vision includes: An acquisition module is used to acquire a calibration position image when the monocular camera is in a calibration position and a working position image when the monocular camera is in a working position, and to determine the relative position information of the monocular camera in the target area; a determination module, configured to determine depth information corresponding to each pixel in the calibrated image based on a preset depth detection model, and to determine actual vertical distances from a plurality of ground contact points in the calibrated image to the monocular camera, as well as actual horizontal distances between pairs of horizontal points; a first construction module, configured to fit a correspondence between the depth information of the plurality of ground contact points and the actual vertical distances by using a polynomial regression equation to correct a depth estimation error of the depth detection model, and determine a vertical coordinate mapping relationship between pixel coordinates in a camera reference coordinate system based on the depth detection model and the fitted polynomial regression equation; a second construction module, configured to construct a spatial triangle and an imaging triangle using the horizontal point pair and the monocular camera, calculate a focal projection distance based on a proportional relationship of similar triangles and an actual horizontal distance of the horizontal point pair, and determine a horizontal coordinate mapping relationship of pixel coordinates in a camera reference coordinate system based on the focal projection distance, wherein the focal projection distance is a conversion scale factor between pixel distance and actual distance; a conversion module, configured to convert the initial pixel coordinates of the person to be located in the working position image into the pixel coordinate system at the correction position based on a coordinate system conversion matrix between the working position and the correction position, to obtain pixel coordinates of the correction position; A positioning module is used to determine the target positioning information of the person to be positioned in the world coordinate system based on the corrected pixel coordinates, the vertical coordinate mapping relationship, the horizontal coordinate mapping relationship and the relative position information.

9. A personnel positioning device based on monocular vision, characterized in that: The monocular vision-based personnel positioning device includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instructions in the memory to enable the monocular vision-based personnel positioning device to execute the monocular vision-based personnel positioning method according to any one of claims 1 to 7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instruction is read and executed, the method for positioning personnel based on monocular vision as described in any one of claims 1 to 7 is executed.

Citation Information

Cited By

  • Spatial modeling method driven by multi-modal heterogeneous data

    CN122066729A