Position monitoring method and system based on different-source image registration and binocular positioning
By combining the image information of infrared and visible light cameras, the YOLOv8 algorithm and energy loss function are used for object detection and image registration, high-precision three-dimensional target positioning at the hoisting operation site is achieved, solving the problem of insufficient real-time and accuracy in the existing technology, and improving operation safety and real-time monitoring.
Patent Information
- Application Number
- CN202510282036.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing safety monitoring methods for lifting operations have challenges in real-time, accuracy and intelligence, and it is difficult to achieve efficient and accurate target detection and positioning in complex operating environments.
Using a method based on heterologous image registration and binocular positioning, by combining the image information of infrared cameras and visible light cameras, the target detection and instance segmentation are used to construct a target mask image collection, and image pairing and correction are performed through the energy loss function to achieve high-precision three-dimensional target positioning.
It improves the monitoring accuracy and real-time performance of the lifting operation area, ensures the safety of personnel and equipment on the work site, can work effectively under different lighting conditions, and adapt to complex operating environments.
Smart Images

Figure CN120125650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of position monitoring, and particularly to a position monitoring method and system based on heterologous image registration and binocular positioning. Background Art
[0002] With the increasing complexity and danger of modern hoisting operations, traditional hoisting operation safety monitoring methods are facing challenges in terms of real-time performance, accuracy, and intelligence. Traditional safety monitoring methods rely on manual inspections and single-sensor devices, suffering from problems such as slow response, low accuracy, and difficulty in comprehensively covering the operation area. With the rapid development of computer vision technology and three-dimensional reconstruction technology, intelligent monitoring by combining multiple sensor data, especially monitoring methods based on binocular vision and image registration, has gradually become an effective means for safety monitoring at hoisting operation sites.
[0003] In existing visual monitoring methods, many methods use a single visual image for target detection and positioning, or only utilize one type of information in infrared images or visible light images for target detection. These methods usually can only work effectively under specific lighting or viewing angle conditions and are difficult to adapt to complex operation environments. To address these problems, the use of heterologous image registration technology, that is, combining infrared images and visible light images to form left and right binocular cameras for obtaining spatial information, can effectively improve the detection accuracy and real-time performance of personnel and equipment in the hoisting operation area.
[0004] In addition, existing methods have relatively high requirements for the accuracy of image registration and target positioning, and efficient image registration and target positioning technologies are needed. However, due to the problems of high computational complexity and poor robustness in most traditional binocular disparity calculation or three-dimensional reconstruction technologies, problems such as inaccurate target recognition and imprecise target matching often occur in practical applications. Therefore, how to perform accurate target pairing, precise image registration from multiple perspectives, and how to achieve precise three-dimensional target positioning by combining real-time data have become the current research difficulties.
[0005] Based on heterologous image registration and binocular positioning technologies, the present invention proposes an improved safety monitoring method for the hoisting operation area. By combining the image information of infrared cameras and visible light cameras, using the YOLOv8 algorithm for object detection and instance segmentation, constructing a set of target mask images, and using an energy loss function for image pairing and correction, accurate target paired images are obtained. Further, by extracting image feature points and introducing an angle fusion score, the feature points are screened, and the bit difference degree is used for image registration, finally achieving high-precision three-dimensional target positioning. This method combines binocular disparity calculation and triangulation principle to construct a three-dimensional point cloud model of the hoisting operation site of a crane, real-time monitors the spatial positions of personnel and equipment at the operation site, improves the monitoring accuracy and real-time performance, and effectively ensures the safety of hoisting operations. Summary of the Invention
[0006] Based on the above-mentioned disadvantages of the prior art, the object of the present invention is to provide a position monitoring method and system based on heterologous image registration and binocular positioning to solve the above technical problems.
[0007] To achieve the above object, the present invention provides the following technical solution: A position monitoring method based on heterologous image registration and binocular positioning, including:
[0008] By setting infrared cameras and visible light cameras to form left and right binocular cameras to obtain image data of the dangerous area of the hoisting operation, and using the YOLOv8 algorithm for object detection and instance segmentation of the image data to construct a set of mask images;
[0009] By constructing an energy loss function to pair and process the mask images of each target of the left and right cameras, obtaining target paired and corrected images of the left camera target mask image and the right camera target mask image;
[0010] Extract the image feature points of the target pair and introduce an angle fusion score, screen the extracted feature points to obtain the best feature points, perform image registration on the best feature points according to the bit difference degree, and fuse and display the registered left and right camera images;
[0011] Use the cooperative registration dual energy function to calculate the optimal binocular disparity, calculate the three-dimensional coordinates of the target image in the actual space through the optimal binocular disparity combined with the triangulation principle, and construct a three-dimensional point cloud model of the hoisting operation site of the crane and a point cloud of personnel targets;
[0012] Obtain the spatial position of the crane through image matching, and move the personnel target point cloud image to the background point cloud according to the spatial position coordinates of the crane to form a complete real-time personnel monitoring point cloud image.
[0013] The present invention is further configured to perform object detection and instance segmentation on the acquired left camera infrared image Image_left and right camera visible light image Image_right using the YOLOv8 algorithm to create an object mask contour image, obtain a binary mask image of the object after training and mapping, and construct a mask image set mask_left and mask_right, where mask_left is the left camera mask image set and mask_right is the right camera mask image set.
[0014] The present invention is further configured to calculate the energy loss function logic: Loss match = ∑(mask_left[a] i,j - mask_right[b] i,j ), where Loss match is the energy loss function. By calculating the energy loss for each mask image of the left and right cameras, the right camera target mask image mask_right[b] paired with the left camera target mask image mask_left[a] corresponding to the smallest energy loss value is selected to form an object pair; the paired left and right mask images are corrected using binocular image stereo rectification technology to obtain the rectified left mask image rectified_mask_left and right mask image rectified_mask_right.
[0015] The present invention is further configured to extract feature points from the object pair using the ORB operator to obtain candidate feature points f, and the best feature point screening logic: f best = max(S match (f)), where f best is the best feature point; S match (f) is the angular fusion score of the candidate feature points, and the calculation logic is: where N is the number of feature points; d i is the distance of the i-th matching point; d max is the maximum allowable matching distance; w i is the weight of the i-th matching point; e angle,i is the angular error, and the calculation logic is: where and are the direction vectors of the left and right camera target mask images under the i-th candidate feature point; S localgeo,i is the local geometric consistency score, and the calculation logic is: where K is the number of neighborhood points of the i-th candidate feature point; d ij is the distance between the i-th matching point and the j-th matching point; w j is the weight of the j-th neighborhood point.
[0016] The present invention is further configured to calculate the descriptor of the optimal feature points by using the Brief operator, use a Brute-Force matcher for the obtained feature points and descriptors, set the matching condition to match according to the bit difference degree, and obtain a matching result. Among them, the calculation logic of the bit difference degree is as follows: Where D(A,B) is the bit difference degree, N is the number of bits of the descriptor, A i and B i are the values of the two descriptors at the i-th bit respectively, and ⊕ is the exclusive OR operator.
[0017] The present invention is further configured to cut out the target image from the left camera image according to the result of instance segmentation based on the matching result, and project all target points according to the perspective transformation matrix to obtain the registered left camera target image Image2; the image fusion principle is: Image3(x,y) = αImage2(x,y)+(1-α)Image_right, where Image3 is the fused image, x and y are the horizontal and vertical coordinates of the image respectively, and α is the fusion weight. The calculation logic is as follows: Where G red is the gradient amplitude of the infrared image in the target area, G rgb is the gradient amplitude of the visible light image in the target area, and ∈ is a minimum value to prevent the denominator from being zero.
[0018] The present invention is further configured to calculate the collaborative registration dual energy function logic S(d) = ∑ x,y |rectified_mask_left(x-d,y)-rectified_mask_right(x,y)|, where d is the pixel distance of translation, and S(d) is the collaborative registration dual energy value; the calculation logic of the best binocular disparity is as follows: Where disp is the best binocular disparity.
[0019] The present invention is further configured to complete target positioning according to the principle of triangulation. The target positioning calculation logic is as follows: Among them, Location(X, Y, Z) is the target position, cx and cy are the coordinates of the center point in the pixel coordinate system, B is the actual physical distance between the optical centers of the binocular cameras, and f is the focal length of the camera; the crane is used to carry the camera to collect on-site images to form a binocular camera fusion image Image3, and Colmap is used for multi-view stereo modeling to generate a complete three-dimensional point cloud model of the lifting operation site of the crane; for the part with non-zero pixel values in the rectified_mask_left of the corrected left camera mask image, the best binocular disparity is used to fill it to obtain a disparity image. The binocular stereo vision principle is used to restore the disparity image into a three-dimensional point cloud image in space, and a thickness δ is given according to the empirical value. The point cloud in the generated planar point cloud image is randomly dispersed within the range space of δ before and after the original plane to simulate the thickness of the personnel to obtain the personnel target point cloud.
[0020] The present invention is further configured to use the image registration function of Colmap to match the currently captured image with the background image in the database, set the crane as the target, obtain the real-time spatial position Location(X, Y, Z) of the crane according to the target positioning method, and map the personnel target point cloud image to the background point cloud based on the spatial position of the crane to realize the real-time monitoring and tracking of the personnel.
[0021] The present invention also provides a position monitoring system based on heterologous image registration and binocular positioning, and the system includes:
[0022] Image acquisition and target detection module: The left and right binocular cameras are set up by setting an infrared camera and a visible light camera to obtain image data of the dangerous area of the lifting operation, and the YOLOv8 algorithm is used to perform target detection and instance segmentation on the image data to construct a set of mask images;
[0023] Target pairing and correction module: By constructing an energy loss function, the mask images of each target of the left and right cameras are paired and processed to obtain the target pairing and correction images of the left camera target mask image and the right camera target mask image;
[0024] Image registration and fusion module: Extract the feature points of the target-paired images and introduce an angle fusion score, screen the extracted feature points to obtain the best feature points, perform image registration on the best feature points according to the bit difference degree, and fuse and display the registered left and right camera images;
[0025] Three-dimensional space reconstruction and point cloud generation module: Extract the feature points of the target-paired images and introduce an angle fusion score, screen the extracted feature points to obtain the best feature points, perform image registration on the best feature points according to the bit difference degree, and fuse and display the registered left and right camera images;
[0026] Background integration and real-time monitoring module: Obtain the spatial position of the crane through image matching, and move the point cloud image of the personnel target to the background point cloud according to the spatial position coordinates of the crane to form a complete real-time personnel monitoring point cloud image.
[0027] The present invention provides a position monitoring method and system based on heterologous image registration and binocular positioning. The method obtains image data of the dangerous area of the hoisting operation by setting an infrared camera and a visible light camera to form a left and right binocular camera, and uses the YOLOv8 algorithm to perform object detection and instance segmentation on the image data to construct a mask image set; by constructing an energy loss function, pair and process the mask images of each object of the left and right cameras to obtain the target pairing and correction images of the left camera target mask image and the right camera target mask image; extract the image feature points of the target pairing and introduce an angle fusion score, screen the extracted feature points to obtain the best feature points, perform image registration on the best feature points according to the bit difference degree, and fuse and display the registered left and right camera images; calculate the optimal binocular disparity using the collaborative registration dual energy function, calculate the three-dimensional coordinates of the target image in the actual space through the optimal binocular disparity combined with the triangulation principle, construct a three-dimensional point cloud model of the hoisting operation site of the crane and the point cloud of the personnel target; obtain the spatial position of the crane through image matching, and move the point cloud image of the personnel target to the background point cloud according to the spatial position coordinates of the crane to form a complete real-time personnel monitoring point cloud image. The beneficial effects generated include:
[0028] Improve the safety of hoisting operations: By combining the heterologous images of the infrared camera and the visible light camera, and using binocular vision technology to realize real-time monitoring of the dangerous area of the hoisting operation. This method can still maintain high-efficiency object detection and monitoring capabilities under different lighting conditions, ensuring real-time safety monitoring of personnel and equipment in the operation area. Through precise target pairing and correction, combined with three-dimensional point cloud modeling technology, it can comprehensively monitor the risk factors in the hoisting operation area and timely discover potential risks, effectively improving the operation safety.
[0029] High-precision three-dimensional target positioning: Using the binocular disparity and triangulation principle, calculate the optimal binocular disparity through collaborative registration, and can accurately calculate the three-dimensional coordinates of the target. This method avoids the spatial positioning accuracy problem in traditional monocular vision or single image registration methods, and realizes high-precision three-dimensional positioning of the targets at the hoisting operation site. Through the three-dimensional point cloud model, it can accurately reflect the spatial layout and personnel dynamics of the operation site, further enhancing the safety guarantee of personnel and equipment during the hoisting operation.
[0030] Intelligent and real-time monitoring: Using the YOLOv8 algorithm to efficiently detect and instance segment targets, it can quickly identify and segment targets in complex scenarios, reducing manual intervention and improving the automation level of monitoring. Through technologies such as image registration, target matching, and image fusion, the precise positions of targets in the hoisting operation area can be obtained in real time, ensuring real-time monitoring and tracking of operators and equipment. Especially combined with the real-time acquisition of the crane's spatial position, the point cloud image of the personnel target can be fused with the point cloud of the crane background to form a complete real-time monitoring image, providing data support for subsequent decision-making.
[0031] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the specific embodiments of this application are specifically given below. Brief Description of the Drawings
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0033] Figure 1 It is a flowchart of a position monitoring method based on heterologous image registration and binocular positioning shown in an exemplary embodiment of the present invention;
[0034] Figure 2 It is a schematic structural diagram of a position monitoring system based on heterologous image registration and binocular positioning shown in an exemplary embodiment of the present invention. Detailed Embodiments
[0035] The following will illustrate the embodiments of the present invention with reference to the drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, rather than for limiting the protection scope of the present invention.
[0036] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0037] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0038] Embodiment 1
[0039] A position monitoring method based on heterologous image registration and binocular positioning, as Figure 1 shown, includes:
[0040] Acquire image data of the dangerous area of the hoisting operation by setting an infrared camera and a visible light camera to form a left and right binocular camera, and use the YOLOv8 algorithm to perform object detection and instance segmentation on the image data to construct a set of mask images;
[0041] Pair and process the mask images of each target of the left and right cameras by constructing an energy loss function to obtain the target pairing and correction images of the left camera target mask image and the right camera target mask image;
[0042] Extract the image feature points of the target pairing and introduce an angle fusion score, screen the extracted feature points to obtain the best feature points, perform image registration on the best feature points according to the bit difference degree, and fuse and display the registered left and right camera images;
[0043] Calculate the optimal binocular disparity using the cooperative registration dual energy function, calculate the three-dimensional coordinates of the target image in the actual space through the optimal binocular disparity combined with the principle of triangulation, and construct a three-dimensional point cloud model of the hoisting operation site of the crane and a point cloud of the personnel target;
[0044] Obtain the spatial position of the crane through image matching, and move the personnel target point cloud image to the background point cloud according to the spatial position coordinates of the crane to form a complete real-time personnel monitoring point cloud image.
[0045] The present invention is further configured to perform object detection and instance segmentation on the acquired left camera infrared image Image_left and right camera visible light image Image_right using the YOLOv8 algorithm to create a target mask contour image, obtain a binary mask image of the target after training mapping, and construct a mask image set mask_left and mask_right, where mask_left is the left camera mask image set and mask_right is the right camera mask image set. Specifically, the visible light images and infrared images taken on the crane site are used as the data set, and CVAT is used to annotate the images and output the mask images of the mask data set. The constructed training data set is randomly divided into a training set, a validation set, and a test set. The YOLOv8 network model is built using the PyTorch framework, and the corresponding network training parameters are set. The network training parameters include model settings, training settings, validation settings, and test settings. After the settings are completed, an open-source pre-trained model is loaded, and training starts using the GPU. After training is completed, the YOLOv8 instance segmentation network model with the best performance in the validation set is obtained. The visible light image and the infrared image are respectively input into the trained object detection network, and the information of object recognition is output, including the category of the object, the position information of the detection frame, and the target mask information. The contour point information of the target to be detected is obtained through YOLOv8 instance segmentation and mapped onto a pure black image with the same resolution as the original image prepared in advance. The part between the contour points is filled with a white image in a polygon shape, and finally, a binary mask image of each target of interest in the image is formed and constructed into a mask image set.
[0046] The present invention is further configured to calculate the energy loss function logic: Loss match = ∑(mask_left[a] i,j - mask_right[b] i,j ), where Loss matchis the energy loss function. By calculating the energy loss for each mask image of the left and right cameras, the right camera target mask image mask_right[b] paired with the left camera target mask image mask_left[a] corresponding to the minimum energy loss value is selected to form the target pair; the paired left and right mask images are corrected using binocular image stereo rectification technology to obtain the rectified left mask image rectified_mask_left and the rectified right mask image rectified_mask_right. Specifically, if there are multiple target information in the same image, the mask images of each target of the left and right cameras are paired by constructing an energy loss function. Traverse the set of target mask images of various categories of the left camera. For each target mask image in it, calculate the energy loss function Loss match with each target mask image in the set of target mask images of the same category of the right camera. Select the right camera target mask image corresponding to the minimum energy loss function Loss match value as the target mask image of the left camera. If there is a situation where the same right camera target mask image corresponds to multiple left camera target mask images, compare the energy loss values of all pairing methods containing this right camera target mask image, and select the left camera target mask image with the minimum energy loss Loss match value as its corresponding paired image, and select the right camera target mask image with the second smallest energy loss Loss match in the corresponding group for other left camera target mask images to pair.
[0047] The present invention is further configured to extract feature points from the target pair using the ORB operator to obtain candidate feature points f. The best feature point screening logic: f best = max(S match (f)), where f best is the best feature point; S match (f) is the angular fusion score of the candidate feature point, and the calculation logic is: where N is the number of feature points; d i is the distance of the i-th matching point; d max is the maximum allowable matching distance; w i is the weight of the i-th matching point; e angle,i is the angular error, and the calculation logic is: where, and are the direction vectors of the left and right camera target mask images under the i-th candidate feature point; S localgeo,i is the local geometric consistency score, and the calculation logic is: where K is the number of neighborhood points of the i-th candidate feature point; d ijis the distance between the i-th matching point and the j-th matching point; w j is the weight of the j-th neighborhood point. Specifically, the one with the highest score is selected as the best feature point through the angular fusion scoring of the feature points, e angle,i is the angular error, which measures the direction difference between feature points under two perspectives, is the distance score, which is used to evaluate the distance accuracy between matching points. The smaller the distance, the better the match and the higher the score, is the angular error score, which is used to evaluate the direction consistency between matching points. The smaller the angular error, the stronger the direction consistency and the higher the score, S localgeo,i is the local geometric consistency score, which is used to measure the geometric consistency between feature points and neighborhood points. The higher the local consistency, the higher the score.
[0048] The present invention is further configured to use the Brief operator to calculate the descriptor of the optimal feature point, use the Brute-Force matcher for the obtained feature points and descriptors, and set the matching condition to match according to the bit difference degree to obtain the matching result. Among them, the calculation logic of the bit difference degree: Among them, D(A,B) is the bit difference degree, N is the number of bits of the descriptor, A i and B i are the values of the two descriptors at the i-th bit respectively, and ⊕ is the exclusive OR operator. Specifically, for the left and right target mask images of the same pair, traverse the set of best feature point descriptors detected in the left target mask image, and calculate the one with the smallest distance between this descriptor and each feature point descriptor in the right target mask image as the best matching point to form the matching result.
[0049] The present invention is further configured to cut out the target image from the left camera image according to the result of instance segmentation based on the matching result, and project all target points according to the perspective transformation matrix to obtain the registered left camera target image Image2; the principle of image fusion: Image3(x,y) = αImage2(x,y)+(1-α)Image_right, where Image3 is the fused image, x and y are the horizontal and vertical coordinates of the image respectively, and α is the fusion weight. The calculation logic is: where, G red is the gradient amplitude of the infrared image in the target area, G rgb is the gradient amplitude of the visible light image in the target area, and ∈ is a minimum value to prevent the denominator from being zero. Specifically, calculate the perspective transformation matrix for the registration transformation of the left image to the right image according to the matching result, and the formula is as follows: Among them, x' and y' are the coordinates of the feature points in the right camera image, and x and y are the coordinates of the feature points in the left camera image. The specific values of the M matrix are calculated through four groups of mutually matching feature point coordinates. The target image is cut out from the left camera image according to the result of instance segmentation. All the target points of the target image are projected according to the perspective transformation matrix to obtain the registered left camera target image Image2. The Sobel operator is used to calculate the gradient of the left camera image and the right camera image respectively to obtain the fusion weight between the two. Then, the target position of the right camera image is weighted and fused with the registered left camera target image to obtain the visible light and infrared target fusion image image3.
[0050] The present invention is further set as the collaborative registration dual energy function calculation logic S(d)=Σ x,y |rectified_mask_left(x - d,y)-rectified_mask_right(x,y)|, where d is the pixel distance of translation, and S(d) is the collaborative registration dual energy value; the best binocular disparity calculation logic: Among them, disp is the best binocular disparity. Specifically, for each translation distance d, we calculate the pixel difference between the translated left image and the right image at the same position. The smaller the difference, the better the matching of the two images, and the smaller the collaborative registration dual energy value d. By calculating the energy for different translation distances, finally, the d that makes the collaborative registration dual energy value the smallest is selected as the best disparity; S(d) is the collaborative registration dual energy value, which reflects the total loss when the left-eye camera mask image is translated d distances to the left. rectified_mask_left(x - d,y) is the pixel value after translating d pixels in the corrected left image mask image, and rectified_mask_right(x,y) is the pixel value in the corrected right image mask image.
[0051] The present invention is further set as completing target positioning according to the principle of triangulation. The target positioning calculation logic: Among them, Location(X, Y, Z) is the target position, cx and cy are the coordinates of the center point in the pixel coordinate system, B is the actual physical distance between the optical centers of the binocular cameras, and f is the focal length of the cameras; the crane is used to carry the cameras to collect on-site images to form a binocular camera fusion image Image3, and Colmap is used for multi-view stereo modeling to generate a complete three-dimensional point cloud model of the lifting operation site of the lifting machinery; for the part with non-zero pixel values in the rectified_mask_left of the rectified left camera mask image, the best binocular disparity is used to fill it to obtain a disparity image. The binocular stereo vision principle is used to restore the disparity image into a three-dimensional point cloud image in space, and a thickness δ is given according to the empirical value. The point clouds in the generated planar point cloud image are randomly dispersed within the range space of δ before and after the original plane to simulate the thickness of the personnel to obtain the personnel target point cloud. Specifically, the target positioning method is based on the triangulation principle. The disparity information of the left and right images is obtained through binocular vision technology, and then the coordinate angle measurement principle of the target object in the three-dimensional space is calculated by the distance between two known observation points and the angle or disparity information obtained from these two observation points to calculate the position of the object in space. In binocular stereo vision, two cameras acquire images of the same scene, and the different disparities of the object can be used to inversely deduce its three-dimensional coordinates according to the geometric relationship of the triangle.
[0052] The present invention is further configured to use the image registration function of Colmap to match the currently captured image with the background image in the database, set the crane as the target, obtain the real-time spatial position Location(X, Y, Z) of the crane according to the target positioning method, and map the personnel target point cloud image to the background point cloud based on the spatial position of the crane to achieve real-time monitoring and tracking of the personnel. Specifically, the image registration function of Colmap is further used to compare the on-site real-time image captured by the current camera with the background image in the database, so as to realize the real-time position positioning of the crane at the lifting operation site, and map the personnel target point cloud to the background point cloud to achieve real-time monitoring and tracking of the personnel.
[0053] Embodiment 2
[0054] Please refer to Figure 2 , the exemplary position monitoring system based on heterologous image registration and binocular positioning includes:
[0055] Image acquisition and target detection module: The left and right binocular cameras are set up by an infrared camera and a visible light camera to obtain image data of the dangerous area of the lifting operation, and the YOLOv8 algorithm is used for target detection and instance segmentation of the image data to construct a mask image set;
[0056] Target pairing and calibration module: By constructing an energy loss function, pair and process the mask images of each target of the left and right cameras to obtain the target pairing and calibration images of the left camera target mask image and the right camera target mask image;
[0057] Image registration and fusion module: Extract the image feature points of the target pairing and introduce an angle fusion score, screen the extracted feature points to obtain the best feature points, perform image registration on the best feature points according to the bit difference degree, and perform fusion display on the registered left and right camera images;
[0058] Three-dimensional space reconstruction and point cloud generation module: Extract the image feature points of the target pairing and introduce an angle fusion score, screen the extracted feature points to obtain the best feature points, perform image registration on the best feature points according to the bit difference degree, and perform fusion display on the registered left and right camera images;
[0059] Background integration and real-time monitoring module: Obtain the spatial position of the crane through image matching, move the personnel target point cloud image to the background point cloud according to the spatial position coordinates of the crane to form a complete real-time personnel monitoring point cloud image.
[0060] It should be noted that a position monitoring system based on heterologous image registration and binocular positioning provided by the above embodiments and a position monitoring method based on heterologous image registration and binocular positioning provided by the above embodiments belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiments and will not be repeated here. A position monitoring system based on heterologous image registration and binocular positioning provided by the above embodiments can, in practical applications, allocate the above functions to different functional modules as needed, that is, divide the internal structure of the system into different functional modules to complete all or part of the functions described above. This is not limited here either.
[0061] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0062] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0063] In this application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0064] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0065] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0066] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0067] In several embodiments provided in this application, it should be understood that the disclosed systems can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be electrical, mechanical, or other forms.
[0068] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0069] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0070] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0071] As described above, the above are only specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A position monitoring method based on heterogeneous image registration and binocular positioning, characterized in that: include: By setting up an infrared camera and a visible light camera to form left and right binocular cameras, we can obtain image data of the dangerous area of the lifting operation. We use the YOLOv8 algorithm to perform target detection and instance segmentation on the image data to construct a mask image set. By constructing an energy loss function, the mask images of each target of the left and right cameras are paired and processed to obtain the target pairing and correction image of the left camera target mask image and the right camera target mask image; Extract the image feature points of the target pair and introduce the angle fusion score, screen the extracted feature points to obtain the best feature points, perform image registration on the best feature points according to the bit difference, and fuse the registered left and right camera images for display; The optimal binocular disparity is calculated by using the dual energy function of collaborative registration. The three-dimensional coordinates of the target image in the actual space are calculated by combining the optimal binocular disparity with the principle of triangulation, and a three-dimensional point cloud model of the lifting machinery hoisting operation site and the target point cloud of personnel are constructed. The spatial position of the crane is obtained through image matching, and the target point cloud image of the personnel is moved to the background point cloud according to the spatial position coordinates of the crane to form a complete real-time personnel monitoring point cloud image.
2. A position monitoring method based on heterogeneous image registration and binocular positioning according to claim 1, characterized in that: The YOLOv8 algorithm is used to perform target detection and instance segmentation on the acquired left camera infrared image Image_left and right camera visible light image Image_right to create a target mask contour image. After training mapping, the target mask binary image is obtained, and the mask image sets mask_left and mask_right are constructed, where mask_left is the left camera mask image set, and mask_right is the right camera mask image set.
3. The position monitoring method based on heterogeneous image registration and binocular positioning according to claim 2 is characterized in that: Energy loss function calculation logic: Loss match =∑(mask_left[a] i,j -mask_right[b] i,j ), Among them, Loss match is the energy loss function. The energy loss is calculated for each mask image of the left and right cameras. The right camera target mask image mask_right[b] paired with the left camera target mask image mask_left[a] corresponding to the minimum energy loss value is taken to form a target pair. The paired left and right mask images are corrected using the binocular image stereo correction technology to obtain the corrected left mask image rectified_mask_left and the corrected right mask image rectified_mask_right.
4. The position monitoring method based on heterogeneous image registration and binocular positioning according to claim 1 is characterized in that: Use the ORB operator to extract feature points from the target pairing and obtain candidate feature points f. The best feature point screening logic is: f best =max(S match (f)), where f best is the best feature point; S match (f) is the angle fusion score of the candidate feature points, and the calculation logic is: Where N is the number of feature points; d i is the distance of the i-th matching point; d max is the maximum allowed matching distance; w i is the weight of the i-th matching point; e angle,i is the angle error, calculation logic: in, and is the direction vector of the left and right camera target mask images under the i-th candidate feature point; S local geo,i Score local geometric consistency, calculation logic: Among them, K is the number of neighboring points of the i-th candidate feature point; d ij is the distance between the i-th matching point and the j-th matching point; w j is the weight of the jth neighborhood point.
5. The position monitoring method based on heterogeneous image registration and binocular positioning according to claim 1 is characterized in that: The Brief operator is used to calculate the descriptor of the optimal feature point. The Brute-Force matcher is used for the obtained feature points and descriptors. The matching condition is set to match based on the bit difference to obtain the matching result. The bit difference calculation logic is: Among them, D(A,B) is the bit difference, N is the number of bits of the descriptor, A i and B i are the values of the two descriptors at the i-th position, Is the exclusive OR operator.
6. The position monitoring method based on heterogeneous image registration and binocular positioning according to claim 1 is characterized in that: According to the matching results, the left camera image is cut out into the target image according to the instance segmentation result, and all target points are projected according to the perspective transformation matrix to obtain the registered left camera target image Image2; Image fusion principle: Image3(x,y)=αImage2(x,y)+(1-α)Image_right, where Image3 is the fused image, x and y are the horizontal and vertical coordinates of the image respectively, α is the fusion weight, and the calculation logic is: Among them, G red is the gradient amplitude of the infrared image in the target area, G rgb is the gradient amplitude of the visible light image in the target area, and ∈ is the minimum value to prevent the denominator from being zero.
7. The position monitoring method based on heterogeneous image registration and binocular positioning according to claim 6 is characterized in that: Co-registration dual energy function calculation logic S(d) = ∑ x,y |rectified_mask_left(xd,y)-rectified_mask_right(x,y)|, where d is the pixel distance of translation and S(d) is the dual energy value of collaborative registration; optimal binocular disparity calculation logic: Among them, disp is the optimal binocular disparity.
8. The position monitoring method based on heterogeneous image registration and binocular positioning according to claim 7 is characterized in that: The target positioning is completed based on the principle of triangulation. The target positioning calculation logic is: Among them, Location (X, Y, Z) is the target position, cx and cy are the coordinates of the center point in the pixel coordinate system, B is the actual physical distance between the optical centers of the binocular cameras, and f is the focal length of the camera; the on-site image is collected by carrying a camera on a crane to form a binocular camera fusion image Image3, and Colmap is used for multi-view stereo modeling to generate a complete three-dimensional point cloud model of the lifting operation site of the crane; the part of the pixel value in the rectified left camera mask image rectified_mask_left that is not 0 is filled with the best binocular disparity to obtain a disparity image, and the binocular stereo vision principle is used to restore the disparity image to a three-dimensional point cloud image in space, and the thickness δ is given according to the empirical value, and the point cloud in the generated plane point cloud image is randomly dispersed, and the dispersion space is within the range of δ before and after the original plane to simulate the thickness of the personnel to obtain the personnel target point cloud.
9. The position monitoring method based on heterogeneous image registration and binocular positioning according to claim 8, characterized in that: The image registration function of Colmap is used to match the currently captured image with the background image in the database. The crane is set as the target, and the real-time spatial position Location (X, Y, Z) of the crane is obtained according to the target positioning method. According to the spatial position of the crane, the target point cloud image of the personnel is mapped to the background point cloud to achieve real-time monitoring and tracking of the personnel.
10. A position monitoring system based on heterogeneous image registration and binocular positioning, used to implement a position monitoring method based on heterogeneous image registration and binocular positioning as described in any one of claims 1 to 9, characterized in that: include: Image acquisition and target detection module: By setting up an infrared camera and a visible light camera to form left and right binocular cameras, image data of the dangerous area of the lifting operation is obtained, and the YOLOv8 algorithm is used to perform target detection and instance segmentation on the image data to construct a mask image set; Target pairing and correction module: by constructing an energy loss function, the mask images of each target of the left and right cameras are paired and processed to obtain the target pairing and correction images of the left camera target mask image and the right camera target mask image; Image registration and fusion module: extracts the image feature points of the target pair and introduces the angle fusion score, screens the extracted feature points to obtain the best feature points, performs image registration on the best feature points based on the bit difference, and fuses and displays the registered left and right camera images; 3D space reconstruction and point cloud generation module: extracts the image feature points of the target pair and introduces the angle fusion score, screens the extracted feature points to obtain the best feature points, performs image registration on the best feature points based on the bit difference, and fuses and displays the registered left and right camera images; Background integration and real-time monitoring module: The spatial position of the crane is obtained through image matching, and the personnel target point cloud image is moved to the background point cloud according to the spatial position coordinates of the crane to form a complete real-time personnel monitoring point cloud image.
Citation Information
Patent Citations
Personnel identification and positioning method fusing target edge detection and scale invariant feature transformation
CN116883945A
Different-source image registration and binocular positioning method for visible light image and infrared image
CN118587283A
Abdominal cavity reconstruction and focus positioning method, system and equipment based on binocular endoscope
CN119313824A
Binocular inertial navigation SLAM positioning method fused with deep learning
CN119417899A
Plan-view projections of depth image data for object tracking
US7003136B1
Cited By
Unmanned aerial vehicle self-adaptive intelligent navigation and obstacle avoidance system based on multi-mode perception
CN120973055A
Hard trigger asymmetric alternating illumination three-dimensional reconstruction method based on binocular vision
CN122244337A