Target coordinate positioning system and method based on multi-view panoramic camera
By constructing the basic matrix and feature matching of a multi-view panoramic camera and combining it with the Hungarian algorithm to optimize target matching, the misjudgment problem of multi-view panoramic cameras in locating fast-moving targets is solved, and accurate target coordinate positioning and detection are achieved.
Patent Information
- Application Number
- CN202511269534.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-08
AI Technical Summary
When multi-view panoramic cameras locate fast-moving and small targets, there are problems with inaccurate target detection, high misjudgment rate and false alarms, especially in multi-target scenarios, which can easily lead to target loss and incorrect coordinates.
By constructing the basic matrix between the camera units of the multi-view panoramic camera, evaluating the spatial geometric relationship, combining target detection and feature matching, using the Hungarian algorithm to optimize target matching, and performing image feature fusion to obtain the three-dimensional coordinates of the target object.
It achieves precise positioning of fast-moving targets, reduces the misjudgment rate, improves the accuracy of target object coordinates and detection range, and reduces false alarms.
Smart Images

Figure CN120747237A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target coordinate positioning, and in particular to a target coordinate positioning system and method based on a multi-view panoramic camera. Background Art
[0002] A multi-eye panoramic camera is composed of multiple camera units combined in a specific way, and uses software and hardware to stitch together the images taken by multiple camera units to obtain a panoramic image with an ultra-large field of view. For fast-moving and small targets, the field of view of a traditional single camera is limited, and the target can easily move out of the field of view of a single camera, resulting in target tracking failure. Using a pan-tilt camera to track targets cannot grasp the positions of multiple targets at the same time, and it is easy to cause the target to be lost when there are multiple targets. However, a multi-eye panoramic camera can achieve horizontal panoramic coverage by arranging multiple cameras in a surround view, completely eliminating blind spots. In addition, a multi-eye panoramic camera can directly calculate the specific indicators of the target in three-dimensional space by using the visual difference of multiple cameras viewing the same scene from different perspectives to directly calculate the specific indicators of the target in three-dimensional space.
[0003] Although multi-eye panoramic cameras play an irreplaceable role in obtaining the positions of some fast-moving and small targets, there are still many problems in the actual target positioning process. For example, the target is far away from the panoramic multi-eye camera in actual distance and its movement posture changes are complex, making it difficult to detect it accurately and stably. In addition, since different types of targets have similar appearances and sizes, it is easy to cause different targets in images taken by multiple camera units to be judged as the same target in practice, making it impossible to accurately match the targets. This will not only lead to the loss of the target, but may even lead to the generation of incorrect target coordinates, causing false alarms and waste of resources. Summary of the Invention
[0004] The purpose of the present invention is to provide a target coordinate positioning system and method based on a multi-eye panoramic camera to solve the problems raised in the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a target coordinate positioning method based on a multi-view panoramic camera, the method comprising: Step S1: Obtain camera parameter data of different camera units in the multi-camera panoramic camera, construct the basic matrix between different camera units, evaluate the spatial geometric relationship between different camera units, and obtain spatial epipolar line data; Step S2: Obtain an image captured by a camera unit in a multi-camera panoramic camera, obtain historical images of the area where the image is located, obtain a target database in the platform, perform target object detection on the image, obtain pending target objects, and screen the pending target objects to obtain target object data; Step S3: Acquire target object data of other images captured at the same time as the image, obtain spatial epipolar line data between the camera unit of the image and the camera unit of the other images, and obtain candidate matching objects of the target object in the image in the other images to obtain target candidate matching data; Step S4: Obtain target candidate matching data of the image and other images, evaluate the comprehensive cost between the target object in the image and the candidate matching objects in other images, construct and solve the target cost matrix between the image and other images, and obtain target object matching data; Step S5: Acquire images taken at the same time, obtain target object matching data between the images, perform image feature fusion to obtain a fused image, obtain the three-dimensional coordinates of the target object in the fused image and aggregate them to obtain target coordinate data.
[0006] Furthermore, step S1 includes: Step S11: Obtaining camera parameter data of each camera unit in the multi-eye panoramic camera, wherein the camera parameter data includes the coordinates of the principal points and the values corresponding to the focal lengths of the camera units in the multi-eye panoramic camera; According to the camera parameter data, the intrinsic parameter matrix K of the camera unit is constructed, and the intrinsic parameter matrix of each camera unit is constructed; Step S12: constructing a basic matrix between each camera unit, wherein the process of constructing the basic matrix between the a-th camera unit and the b-th camera unit is: Get the three-dimensional coordinates P of the physical point P in the coordinate system of the ath camera unit and the bth camera unit respectively a and P b ; Calculate the rotation matrix R and translation vector t between the a-th camera unit and the b-th camera unit. The specific calculation formula is: P a =P b ·R+t; Get the antisymmetric matrix t' of the translation vector t and calculate the basic matrix F between the a-th camera unit and the b-th camera unit (a,b) =K b -T ·R·K a -1 , where K b -T is the transpose of the inverse matrix of the intrinsic parameter matrix K of the b-th camera unit; K a -1 is the intrinsic parameter matrix K of the a-th camera unit a The inverse matrix of Step S13: Evaluate the spatial geometric relationship between the a-th camera unit and the b-th camera unit. The specific evaluation process is as follows: Get the coordinates (x a α ,y a α ), get the homogeneous coordinates α´ of the image point α a =[x a α ,y a α ,1] T , calculate the epipolar line L of the image point α in the bth camera unit b α =F (a,b) ·α´ a ; Get the coordinates (x b β ,y b β ), get the homogeneous coordinates β´ of the image point β b =[x b β ,y b β ,1] T , calculate the epipolar line L of the image point β in the a-th camera unit a β =F (a,b) T β´ b , where F (a,b) T is the basic matrix F (a,b) The transpose of Obtain the epipolar lines of each image point of the a-th camera unit in the b-th camera unit, obtain the epipolar lines of each image point of the b-th camera unit in the a-th camera unit, and combine them to obtain the spatial epipolar line data between the a-th camera unit and the b-th camera unit; Step S14: Evaluate the spatial geometric relationship between different camera units to obtain spatial epipolar data between different camera units.
[0007] Furthermore, step S2 includes: Step S21: Acquire images captured by the camera units in the multi-eye panoramic camera in the current cycle, acquire historical images of the area where the images are located, construct a two-dimensional coordinate system for the images, pre-process the images, and obtain pixel values of each pixel in the images; Step S22: obtaining a preset target database in the platform, wherein the target database includes a preset image set corresponding to each target object; Detecting target objects in the image using a preset target detection model based on the target database to obtain object region data for each target object to be determined in the image, wherein the object region data is the minimum square region where the target object to be determined is located in the image; Obtain the side length of the minimum square area and record it as the side length of the target area of the target object to be determined; obtain the average value of the pixels in the minimum square area where the target object to be determined is located in the image and record it as the pixel value ζ of the marked area of the target object to be determined in the image; Step S23: Obtain the three-dimensional coordinates of the target object to be determined in the image, obtain various historical images of the area where the image is located, obtain the coordinates of the center point of the target object to be determined in the historical images based on the three-dimensional coordinates of the target object to be determined in the image, construct a square area with the side length of the target area of the target object to be determined as the side length and the coordinates of the center point as the center coordinates, obtain the average value of the pixels in the square area in the historical images, and obtain the pixel value of the characteristic area of the target object to be determined in the historical images; Step S24: Screening the target objects. The specific screening process is to obtain the mean value ζ' of the pixel values of the feature areas of the target objects in each historical image. When the absolute value of the pixel value ζ of the marked area minus the mean value ζ' is greater than the preset absolute value threshold, the target object is determined to be the target object in the target database, and the target object is retained and recorded as the target object in the image. Otherwise, the target object is determined not to be the target object in the target database, and the target object is eliminated. The coordinates of the center points of each target object in the image are obtained and aggregated to obtain the target object data in the image.
[0008] Furthermore, step S3 includes: Step S31: acquiring other images from the multi-view panoramic camera at the same time as the image capture, acquiring target object data from the other images, acquiring target object data from the image, and acquiring the coordinates of the center point of each target object in the image from the target object data; Step S32: obtaining spatial epipolar data between the camera unit of the image and the camera units of other images, obtaining the coordinates of the center point of a target object in the image, and obtaining the coordinates of the center points of several target objects in other images; Step S33: Obtain the epipolar line L of the coordinates of the center point of a target object in other images based on the spatial epipolar line data between the camera unit of the image and the camera units of other images. △ =[γ,η,λ] T , get the coordinates of the center point of the e-th target object in other images (x ▽ e ,y ▽ e ), calculate the epipolar line L△ The Euclidean distance d to the center coordinate of the e-th target object (▽,e) : , When the Euclidean distance d (▽,e) If the distance is less than the preset threshold, the e-th target object is recorded as a candidate matching object of a certain target object; otherwise, the e-th target object is not processed; Obtain candidate matching objects of each target object in the image in other images, and aggregate them to obtain target candidate matching data between the image and other images.
[0009] Furthermore, step S4 includes: Step S41: Obtain target candidate matching data between the image and other images, obtain candidate matching objects of the target object in the image in the other images from the target candidate matching data, obtain the Euclidean distance d' between the epipolar line of the target object in the other image and the center point coordinates of the candidate matching objects in the other image, obtain the normalized Euclidean distance d', and record it as the geometric cost value between the target object in the image and the candidate matching objects in the other images; Step S42: using a preset object detection model to obtain bounding boxes of each target object in the image and other images, cropping and resizing the bounding boxes in the image and other images, and inputting them into a preset feature extraction model to obtain feature vectors of each target object in the image and other images, and normalizing the feature vectors of each target object in the image and other images; Step S43: Obtain the feature vectors G of the w-th target object in the image and the v-th target objects in other images respectively w and G ▽ v , calculate the feature cost value Q between the w-th target object and the v-th target object (w,v) : , Step S44: when the vth target object is not a candidate matching object for the wth target object, the comprehensive cost value of the wth target object and the vth target object is set to a preset maximum value; When the vth target object is a candidate matching object of the wth target object, calculate the comprehensive cost value U of the wth target object and the vth target object (w,v) : , Among them, D (w,v) is the geometric cost between the vth target object and the wth target object; ψ Q is the preset feature cost weight coefficient; ψ Dis a preset geometric cost weight coefficient; ψ Q + ψ D = 1, ψ Q > 0, ψ D > 0; Step S45: Obtain the comprehensive cost values between each target object in the image and several target objects in other images, construct the target cost matrix between the image and other images, obtain the total number m of all target objects in the image, and obtain the total number n of several target objects in other images; When m = n, do not process the target cost matrix. When m ≠ n, fill the target cost matrix. The specific filling process is as follows: When m > n, fill m - n virtual columns in the last column of the target cost matrix. Among them, the comprehensive cost values of each element in the m - n virtual columns are all preset maximum values; When m < n, fill n - m virtual rows in the last row of the target cost matrix. Among them, the comprehensive cost values of each element in the n - m virtual columns are all preset maximum values; Step S46: Use the Hungarian algorithm to solve the target cost matrix, obtain the zero elements in each row and column of the characteristic target cost matrix, and obtain the matching pairs of the positions corresponding to the zero elements in the characteristic target cost matrix. Among them, a matching pair includes the target object in the image that matches the target object in other images; Obtain the target objects in other images that match each target object in the image, and gather them to obtain the target object matching data between the image and other images; The reason for obtaining the target objects in other images that match each target object in the image in the above steps is that in practice, a certain target object in an image usually has only one matching target object in other images. However, due to the different shooting angles and parameters of different camera units, and for fast - moving target objects with high speed and small size, traditional methods are prone to incorrect matching. The above steps analyze the comprehensive cost of the matching between a certain target object in the image and another target object in other images from two aspects of spatial geometry and appearance features, and use the Hungarian algorithm to solve the target cost matrix between the image and other images, so as to accurately obtain the matching relationship between the target objects between the image and other images. This not only reduces the probability of target object loss but also avoids the generation of virtual coordinates of target objects, greatly improving the accuracy of target object coordinate positioning.
[0010] Furthermore, Step S5 includes: Step S51: Acquire images taken at the same time by a multi-view panoramic camera, obtain target object matching data between the images, and obtain the target object in the other image that matches the target object in the one image from the target object matching data between the one image and the other image; Step S52: acquiring various camera parameters of the camera units corresponding to the respective images, and combining the target object matching data between the respective images to perform image feature fusion on the respective images to obtain a fused image; For example, the specific process of fusing image features of each image to obtain a fused image is as follows: Obtain the target object matching data between one image and another image, obtain the target object that matches one image and another image, obtain the projection matrix of the camera unit corresponding to one image and another image, and calculate the three-dimensional coordinates corresponding to the target object by solving the linear equation system; Use one image and another image to perform image reconstruction, and add the 3D points in several images to obtain the 3D coordinates of the target object in each image; Use the semi-global matching (SGM) method to find corresponding points in the other image for each pixel point in different images (not just the feature points corresponding to the target object); And according to the matched pixel points, the visual difference corresponding to the pixel points is obtained to obtain the depth maps corresponding to different perspectives. The depth maps corresponding to different perspectives are all converted into a preset unified world coordinate system and fused into a fused image corresponding to each image; Step S53: Locate each target object in the fused image to obtain the three-dimensional coordinates of each target object in the fused image, wherein the coordinate system of the three-dimensional coordinates is in a preset world coordinate system. The three-dimensional coordinates of each target object in the fused image are aggregated to obtain the target coordinate data in the fused image.
[0011] In order to better implement the above method, a target coordinate positioning system based on a multi-camera panoramic camera is also proposed. The system includes a camera spatial relationship evaluation module, a target matching module and a target positioning module; A camera spatial relationship evaluation module is used to evaluate the spatial geometric relationship between different camera units and obtain spatial epipolar data between different camera units; The target matching module is used to obtain target object data of the image and other images, obtain candidate matching objects of the target object of the image in other images based on the target object data, obtain target candidate matching data, construct and solve the target cost matrix between the image and other images, and obtain target object matching data between the image and other images; The target positioning module is used to obtain the target object matching data between the images taken at the same time in the multi-eye panoramic camera, fuse the image features of each image to obtain a fused image, obtain the three-dimensional coordinates of the target object in the fused image, and obtain the target coordinate data in the fused image.
[0012] Furthermore, the camera space relationship evaluation module includes a basic matrix construction unit and a camera space relationship evaluation unit; A basic matrix construction unit is used to obtain camera parameter data of different camera units in a multi-camera panoramic camera and construct a basic matrix between different camera units; The camera spatial relationship evaluation unit is used to evaluate the spatial geometric relationship between different camera units based on the basic matrix between different camera units, and obtain the spatial epipolar line data between different camera units.
[0013] Furthermore, the target matching module includes a target detection unit, a candidate matching acquisition unit, and a target matching unit; The target detection unit is used to detect target objects in the image and screen the target objects in the image in combination with historical images to obtain the target object data in the image; a candidate matching acquisition unit, configured to acquire target object data from other images captured at the same time as the image, acquire spatial epipolar line data between the camera unit of the image and the camera unit of the other images, acquire candidate matching objects of the target object of the image in the other images, and obtain target candidate matching data; The target matching unit is used to construct and solve the target cost matrix between the image and other images based on the target candidate matching data between the image and other images, and obtain the target object matching data between the image and other images.
[0014] Furthermore, the target positioning module includes an image fusion unit and a target positioning unit; An image fusion unit is used to acquire images taken at the same time by a multi-lens panoramic camera, obtain target object matching data between the images, and fuse image features of the images to obtain a fused image; The target positioning unit is used to obtain and aggregate the three-dimensional coordinates of the target object in the fused image to obtain the target coordinate data in the fused image.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention realizes the precise positioning of target coordinates based on a multi-eye panoramic camera, and considers that the multi-eye panoramic camera actually needs to detect fast and fast-moving target objects. Therefore, not only the preset model is used to detect the target object, but also the authenticity of the target object is further verified by analyzing the difference between the pixel values of the target object area in the historical image and the pixel values of the target object area in the image, and the target objects in the images taken by each camera unit in the multi-eye panoramic camera are accurately matched, thereby greatly reducing the probability of misjudgment of the target object, so that the position of the target object can be accurately grasped, and finally, according to the matching relationship of the target objects in the images taken by each camera unit, the images taken by each camera unit are fused, and then the target coordinates of the target object in the fused image are accurately positioned, which not only makes the captured image viewing angle range large so that more target objects can be detected, but also can achieve precise positioning of the target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a method flow chart of a target coordinate positioning method based on a multi-view panoramic camera of the present invention; Figure 2 This is a module flow chart of a target coordinate positioning system based on a multi-eye panoramic camera of the present invention. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0018] Example: Figure 1-Figure 2 As shown, the present invention provides a technical solution, a target coordinate positioning method based on a multi-view panoramic camera, the method comprising: Step S1: Obtain camera parameter data of different camera units in the multi-camera panoramic camera, construct the basic matrix between different camera units, evaluate the spatial geometric relationship between different camera units, and obtain spatial epipolar line data; For example, a camera unit refers to a camera head of a multi-view panoramic camera; Wherein, step S1 includes: Step S11: Obtaining camera parameter data of each camera unit in the multi-eye panoramic camera, wherein the camera parameter data includes the coordinates of the principal points and the values corresponding to the focal lengths of the camera units in the multi-eye panoramic camera; According to the camera parameter data, the intrinsic parameter matrix K of the camera unit is constructed, and the intrinsic parameter matrix of each camera unit is constructed; For example, the intrinsic parameter matrix K of the camera unit is specifically: , Among them, f x and f y are the focal length f of the camera unit converted into pixel units in the horizontal and vertical directions of the image; (c x ,c y ) is the principal point coordinate of the camera unit; c x ,c y They represent the x-axis and y-axis coordinates of the optical center in the image coordinate system, respectively. The optical center is the pixel coordinate of the point where the optical axis in the camera unit passes through the image sensor; For example, f x =f / dx, dx is the physical width of each pixel in the horizontal direction of the image; Step S12: constructing a basic matrix between each camera unit, wherein the process of constructing the basic matrix between the a-th camera unit and the b-th camera unit is: Get the three-dimensional coordinates P of the physical point P in the coordinate system of the ath camera unit and the bth camera unit respectively a and P b ; For example, a physical point P is a point that exists in reality; Calculate the rotation matrix R and translation vector t between the a-th camera unit and the b-th camera unit. The specific calculation formula is: P a =P b ·R+t; Get the antisymmetric matrix t' of the translation vector t and calculate the basic matrix F between the a-th camera unit and the b-th camera unit (a,b) =K b -T ·R·K a -1 , where K b -T is the transpose of the inverse matrix of the intrinsic parameter matrix K of the b-th camera unit; K a -1 is the intrinsic parameter matrix K of the a-th camera unit a The inverse matrix of For example, the antisymmetric matrix t´ is: , Where t=[t x ,t y ,t z ] T ; Step S13: Evaluate the spatial geometric relationship between the a-th camera unit and the b-th camera unit. The specific evaluation process is as follows: Get the coordinates (x a α ,y a α ), get the homogeneous coordinates α´ of the image point α a =[x a α ,y a α ,1] T , calculate the epipolar line L of the image point α in the bth camera unit b α =F (a,b) ·α´ a ; Get the coordinates (x b β ,y b β ), get the homogeneous coordinates β´ of the image point β b =[x b β ,y b β ,1] T , calculate the epipolar line L of the image point β in the a-th camera unit a β =F (a,b) T β´ b , where F (a,b) T is the basic matrix F (a,b) The transpose of Obtain the epipolar lines of each image point of the a-th camera unit in the b-th camera unit, obtain the epipolar lines of each image point of the b-th camera unit in the a-th camera unit, and combine them to obtain the spatial epipolar line data between the a-th camera unit and the b-th camera unit; Step S14: Evaluate the spatial geometric relationship between different camera units to obtain spatial epipolar data between different camera units; Step S2: Obtain an image captured by a camera unit in a multi-camera panoramic camera, obtain historical images of the area where the image is located, obtain a target database in the platform, perform target object detection on the image, obtain pending target objects, and screen the pending target objects to obtain target object data; Wherein, step S2 includes: Step S21: Acquire images captured by the camera units in the multi-eye panoramic camera in the current cycle, acquire historical images of the area where the images are located, construct a two-dimensional coordinate system for the images, pre-process the images, and obtain pixel values of each pixel in the images; For example, preprocessing includes,color space conversion, normalization and image resizing; Step S22: obtaining a preset target database in the platform, wherein the target database includes a preset image set corresponding to each target object; For example, the preset target objects include sparrows, pigeons, etc. Detecting target objects in the image using a preset target detection model based on the target database to obtain object region data for each target object to be determined in the image, wherein the object region data is the minimum square region where the target object to be determined is located in the image; Obtain the side length of the minimum square area and record it as the side length of the target area of the target object to be determined; obtain the average value of the pixels in the minimum square area where the target object to be determined is located in the image and record it as the pixel value ζ of the marked area of the target object to be determined in the image; For example, the preset target detection models include YOLO and Faster, which are trained using images corresponding to each target object in the target database, and each training image is labeled with the target object; Step S23: Obtain the three-dimensional coordinates of the target object to be determined in the image, obtain various historical images of the area where the image is located, obtain the coordinates of the center point of the target object to be determined in the historical images based on the three-dimensional coordinates of the target object to be determined in the image, construct a square area with the side length of the target area of the target object to be determined as the side length and the coordinates of the center point as the center coordinates, obtain the average value of the pixels in the square area in the historical images, and obtain the pixel value of the characteristic area of the target object to be determined in the historical images; Step S24: Screening the target objects. The specific screening process is to obtain the mean value ζ' of the pixel values of the feature areas of the target objects in each historical image. When the absolute value of the pixel value ζ in the marked area minus the mean value ζ' is greater than a preset absolute value threshold, the target object is determined to be a target object in the target database, and the target object is retained and recorded as the target object in the image. Otherwise, the target object is determined not to be a target object in the target database, and the target object is eliminated. The coordinates of the center points of each target object in the image are obtained and aggregated to obtain the target object data in the image. Step S3: Acquire target object data of other images captured at the same time as the image, obtain spatial epipolar line data between the camera unit of the image and the camera unit of the other images, and obtain candidate matching objects of the target object in the image in the other images to obtain target candidate matching data; Wherein, step S3 includes: Step S31: acquiring other images from the multi-view panoramic camera at the same time as the image capture, acquiring target object data from the other images, acquiring target object data from the image, and acquiring the coordinates of the center point of each target object in the image from the target object data; For example, the coordinates of the image and other images here are based on a two-dimensional coordinate system established based on the image and other images; Step S32: obtaining spatial epipolar data between the camera unit of the image and the camera units of other images, obtaining the coordinates of the center point of a target object in the image, and obtaining the coordinates of the center points of several target objects in other images; Step S33: Obtain the epipolar line L of the coordinates of the center point of a target object in other images based on the spatial epipolar line data between the camera unit of the image and the camera units of other images. △ =[γ,η,λ] T , get the coordinates of the center point of the e-th target object in other images (x ▽ e ,y ▽ e ), calculate the epipolar line L △ Euclidean distance d to the center coordinate of the e-th target object (▽,e) : , When the Euclidean distance d (▽,e) If the distance is less than the preset threshold, the e-th target object is recorded as a candidate matching object of a certain target object; otherwise, the e-th target object is not processed; Obtain candidate matching objects of each target object in the image in other images, and aggregate them to obtain target candidate matching data between the image and other images; Step S4: Obtain target candidate matching data of the image and other images, evaluate the comprehensive cost between the target object in the image and the candidate matching objects in other images, construct and solve the target cost matrix between the image and other images, and obtain target object matching data; Wherein, step S4 includes: Step S41: Obtain target candidate matching data between the image and other images, obtain candidate matching objects of the target object in the image in the other images from the target candidate matching data, obtain the Euclidean distance d' between the epipolar line of the target object in the other image and the center point coordinates of the candidate matching objects in the other image, obtain the normalized Euclidean distance d', and record it as the geometric cost value between the target object in the image and the candidate matching objects in the other images; For example, the value of the normalized Euclidean distance d´ is limited to between 0 and 1; Step S42: using a preset object detection model to obtain bounding boxes of each target object in the image and other images, cropping and resizing the bounding boxes in the image and other images, and inputting them into a preset feature extraction model to obtain feature vectors of each target object in the image and other images, and normalizing the feature vectors of each target object in the image and other images; For example, the feature extraction model is a preset model, a dedicated re-identification model, which is a model trained using a dataset and trained through metric learning techniques such as triplet loss, and can extract highly discriminative features; For example, the feature vector G´ of the ξ-th target object in the image ξ After normalization, we get the eigenvector G ξ, Eigenvector G ξ The specific calculation formula is: , Among them, ||G´ ξ ||2 is the eigenvector G´ ξ The Euclidean paradigm; Step S43: Obtain the feature vectors G of the w-th target object in the image and the v-th target objects in other images respectively w and G ▽ v , calculate the feature cost value Q between the w-th target object and the v-th target object (w,v) : , Step S44: when the vth target object is not a candidate matching object for the wth target object, the comprehensive cost value of the wth target object and the vth target object is set to a preset maximum value; When the vth target object is a candidate matching object of the wth target object, calculate the comprehensive cost value U of the wth target object and the vth target object (w,v) : , Among them, D (w,v)is the geometric cost value between the v-th target object and the w-th target object; ψ Q is the preset feature cost weight coefficient; ψ D is the preset geometric cost weight coefficient; ψ Q +ψ D = 1, ψ Q > 0, ψ D > 0; Step S45: Obtain the comprehensive cost values between each target object in the image and several target objects in other images, construct the target cost matrix between the image and other images, obtain the total number m of each target object in the image, and obtain the total number n of several target objects in other images; When m = n, no processing is performed on the target cost matrix. When m ≠ n, the target cost matrix is filled. The specific filling process is as follows: When m > n, fill m - n virtual columns in the last column of the target cost matrix. Among them, the comprehensive cost values of each element in the m - n virtual columns are all preset maximum values; When m < n, fill n - m virtual rows in the last row of the target cost matrix. Among them, the comprehensive cost values of each element in the n - m virtual columns are all preset maximum values; Step S46: Use the Hungarian algorithm to solve the target cost matrix, obtain the zero elements in each row and each column of the feature target cost matrix, and obtain the matching pairs of the positions corresponding to the zero elements in the feature target cost matrix. Among them, a matching pair includes the target object in the image that matches the target object in other images; For example, use the Hungarian algorithm to solve the target cost matrix to obtain the zero elements in each row and each column of the feature target cost matrix. The specific solution process is as follows: Set the objective function of the target cost matrix: , where U (i,j) is the value of the element corresponding to the i-th row and i-th column of the target cost matrix; Set the constraint conditions of the target cost matrix: 1. The target object in the image can only match one target object in other images: , The target object in other images can only match one target object in the image: , X(i, j) is a binary decision variable, and its value is only 0 or 1, indicating whether to assign i in the image to j in other images; The specific steps of the Hungarian algorithm are as follows: 1. Row subtraction: subtract the minimum value of each row from each row in the target cost matrix; 2. Column subtraction: subtract the minimum value of each column from each column in the target cost matrix; 3. Use the minimum number of horizontal and vertical lines to cover all zero elements in the target cost matrix; 4. Adjust the target cost matrix to find the value of the smallest element that is not covered. Subtract the value of the smallest element from the uncovered element, and add the value of the smallest element to all elements in the target cost matrix that are covered by horizontal and vertical lines. 5. Repeat steps 3-4 until the objective function of the target cost matrix is satisfied, thereby obtaining the characteristic target cost matrix; 6. Obtain a set of zero elements in the feature target cost matrix so that each row and each column of the feature target cost matrix has a corresponding zero element; Obtaining the target objects that match each target object in the image in other images, and aggregating them to obtain target object matching data between the image and other images; Step S5: Acquire images taken at the same time, obtain target object matching data between the images, perform image feature fusion to obtain a fused image, obtain the three-dimensional coordinates of the target object in the fused image and aggregate them to obtain target coordinate data; Wherein, step S5 includes: Step S51: Acquire images taken at the same time by a multi-view panoramic camera, obtain target object matching data between the images, and obtain the target object in the other image that matches the target object in the one image from the target object matching data between the one image and the other image; Step S52: acquiring various camera parameters of the camera units corresponding to the respective images, and combining the target object matching data between the respective images to perform image feature fusion on the respective images to obtain a fused image; For example, various camera parameters include the intrinsic matrix, distortion coefficient, rotation matrix and translation vector of the camera unit; Step S53: Positioning each target object in the fused image to obtain the three-dimensional coordinates of each target object in the fused image, wherein the coordinate system of the three-dimensional coordinates is in a preset world coordinate system. The three-dimensional coordinates of each target object in the fused image are aggregated to obtain the target coordinate data in the fused image. In order to better implement the above method, a target coordinate positioning system based on a multi-camera panoramic camera is also proposed. The system includes a camera spatial relationship evaluation module, a target matching module and a target positioning module; A camera spatial relationship evaluation module is used to evaluate the spatial geometric relationship between different camera units and obtain spatial epipolar data between different camera units; The target matching module is used to obtain target object data of the image and other images, obtain candidate matching objects of the target object of the image in other images based on the target object data, obtain target candidate matching data, construct and solve the target cost matrix between the image and other images, and obtain target object matching data between the image and other images; The target positioning module is used to obtain the target object matching data between the images taken at the same time in the multi-view panoramic camera, fuse the image features of each image to obtain a fused image, obtain the three-dimensional coordinates of the target object in the fused image, and obtain the target coordinate data in the fused image; Among them, the camera space relationship evaluation module includes a basic matrix construction unit and a camera space relationship evaluation unit; A basic matrix construction unit is used to obtain camera parameter data of different camera units in a multi-camera panoramic camera and construct a basic matrix between different camera units; A camera spatial relationship evaluation unit is used to evaluate the spatial geometric relationship between different camera units based on the basic matrix between different camera units to obtain spatial epipolar line data between different camera units; Among them, the target matching module includes a target detection unit, a candidate matching acquisition unit and a target matching unit; The target detection unit is used to detect target objects in the image and screen the target objects in the image in combination with historical images to obtain the target object data in the image; a candidate matching acquisition unit, configured to acquire target object data from other images captured at the same time as the image, acquire spatial epipolar line data between the camera unit of the image and the camera unit of the other images, acquire candidate matching objects of the target object of the image in the other images, and obtain target candidate matching data; A target matching unit is used to construct and solve a target cost matrix between the image and the other images based on target candidate matching data between the image and the other images, so as to obtain target object matching data between the image and the other images; Among them, the target positioning module includes an image fusion unit and a target positioning unit; An image fusion unit is used to acquire images taken at the same time by a multi-lens panoramic camera, obtain target object matching data between the images, and fuse image features of the images to obtain a fused image; The target positioning unit is used to obtain and aggregate the three-dimensional coordinates of the target object in the fused image to obtain the target coordinate data in the fused image.
[0019] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
Claims
1. A target coordinate positioning method based on a multi-view panoramic camera, characterized in that: The method comprises: Step S1: Obtain camera parameter data of different camera units in the multi-camera panoramic camera, construct the basic matrix between different camera units, evaluate the spatial geometric relationship between different camera units, and obtain spatial epipolar line data; Step S2: Obtain an image captured by a camera unit in a multi-camera panoramic camera, obtain historical images of the area where the image is located, obtain a target database in the platform, perform target object detection on the image, obtain pending target objects, and screen the pending target objects to obtain target object data; Step S3: Acquire target object data of other images captured at the same time as the image, acquire spatial epipolar line data between the camera unit of the image and the camera unit of the other images, acquire candidate matching objects of the target object of the image in the other images, and obtain target candidate matching data; Step S4: Obtain target candidate matching data of the image and other images, evaluate the comprehensive cost between the target object in the image and the candidate matching objects in other images, construct and solve the target cost matrix between the image and other images, and obtain target object matching data; Step S5: Acquire images taken at the same time, obtain target object matching data between the images, perform image feature fusion to obtain a fused image, obtain the three-dimensional coordinates of the target object in the fused image and aggregate them to obtain target coordinate data.
2. The target coordinate positioning method based on a multi-view panoramic camera according to claim 1, characterized in that: The step S1 comprises: Step S11: Obtaining camera parameter data of each camera unit in the multi-eye panoramic camera, wherein the camera parameter data includes the coordinates of the principal points and the values corresponding to the focal lengths of the camera units in the multi-eye panoramic camera; According to the camera parameter data, the intrinsic parameter matrix K of the camera unit is constructed, and the intrinsic parameter matrix of each camera unit is constructed; Step S12: constructing a basic matrix between each camera unit, wherein the process of constructing the basic matrix between the a-th camera unit and the b-th camera unit is: Get the three-dimensional coordinates P of the physical point P in the coordinate system of the ath camera unit and the bth camera unit respectively a and P b ; Calculate the rotation matrix R and translation vector t between the a-th camera unit and the b-th camera unit. The specific calculation formula is: P a =P b ·R+t; Get the antisymmetric matrix t' of the translation vector t and calculate the basic matrix F between the a-th camera unit and the b-th camera unit (a,b) =K b -T ·R·K a -1 , where K b -T is the transpose of the inverse matrix of the intrinsic parameter matrix K of the b-th camera unit; K a -1 is the intrinsic parameter matrix K of the a-th camera unit a The inverse matrix of Step S13: Evaluate the spatial geometric relationship between the a-th camera unit and the b-th camera unit. The specific evaluation process is as follows: Get the coordinates (x a α ,y a α ), get the homogeneous coordinates α´ of the image point α a =[x a α ,y a α ,1] T , calculate the epipolar line L of the image point α in the bth camera unit b α =F (a,b) ·α´ a ; Get the coordinates (x b β ,y b β ), get the homogeneous coordinates β´ of the image point β b =[x b β ,y b β ,1] T , calculate the epipolar line L of the image point β in the a-th camera unit a β =F (a,b) T β´ b , where F (a,b) T is the basic matrix F (a,b) The transpose of Obtain the epipolar lines of each image point of the a-th camera unit in the b-th camera unit, obtain the epipolar lines of each image point of the b-th camera unit in the a-th camera unit, and combine them to obtain the spatial epipolar line data between the a-th camera unit and the b-th camera unit; Step S14: Evaluate the spatial geometric relationship between different camera units to obtain spatial epipolar data between different camera units.
3. The target coordinate positioning method based on a multi-view panoramic camera according to claim 2, characterized in that: The step S2 comprises: Step S21: obtaining images captured by camera units in a multi-eye panoramic camera in a current cycle, obtaining historical images of the area where the images are located, constructing a two-dimensional coordinate system for the images, preprocessing the images, and obtaining pixel values of each pixel in the images; Step S22: obtaining a preset target database in the platform, wherein the target database includes a preset image set corresponding to each target object; Detecting target objects in the image based on the target database and using a preset target detection model to obtain object region data of each target object to be determined in the image, wherein the object region data is a minimum square region where the target object to be determined in the image is located; Obtain the side length of the minimum square area and record it as the side length of the target area of the target object to be determined; obtain the average value of the pixels in the minimum square area where the target object to be determined is located in the image and record it as the pixel value ζ of the marked area of the target object to be determined in the image; Step S23: obtaining the three-dimensional coordinates of the target object to be determined in the image, obtaining various historical images of the area where the image is located, obtaining the coordinates of the center point of the target object to be determined in the historical images based on the three-dimensional coordinates of the target object to be determined in the image, constructing a square area with the side length of the target area of the target object to be determined as the side length and the coordinates of the center point as the center coordinates, obtaining the average value of the pixels in the square area in the historical images, and obtaining the pixel value of the characteristic area of the target object to be determined in the historical images; Step S24: Screening the target objects. The specific screening process is to obtain the mean value ζ' of the pixel values of the feature areas of the target objects in each historical image. When the absolute value of the pixel value ζ of the marked area minus the mean value ζ' is greater than a preset absolute value threshold, the target object is determined to be a target object in the target database, and the target object is retained and recorded as the target object in the image. Otherwise, the target object is determined not to be a target object in the target database, and the target object is eliminated. The coordinates of the center points of each target object in the image are obtained and aggregated to obtain the target object data in the image.
4. The target coordinate positioning method based on a multi-view panoramic camera according to claim 3, characterized in that: The step S3 comprises: Step S31: acquiring other images from a multi-view panoramic camera that are taken at the same time as the image, acquiring target object data from the other images, acquiring the target object data from the image, and acquiring the coordinates of the center points of each target object in the image from the target object data; Step S32: obtaining spatial epipolar data between the camera unit of the image and the camera units of other images, obtaining the coordinates of the center point of a target object in the image, and obtaining the coordinates of the center points of several target objects in other images; Step S33: Obtain the epipolar line L of the coordinates of the center point of a target object in other images based on the spatial epipolar line data between the camera unit of the image and the camera units of other images. △ =[γ,η,λ] T , get the coordinates of the center point of the e-th target object in other images (x ▽ e ,y ▽ e ), calculate the epipolar line L △ Euclidean distance d to the center coordinate of the e-th target object (▽,e) : , When the Euclidean distance d (▽,e) If the distance is less than a preset threshold, the e-th target object is recorded as a candidate matching object of a certain target object; otherwise, the e-th target object is not processed; Candidate matching objects of each target object in the image in other images are obtained and aggregated to obtain target candidate matching data between the image and other images.
5. The target coordinate positioning method based on a multi-view panoramic camera according to claim 4, characterized in that: The step S4 comprises: Step S41: Obtain target candidate matching data between the image and other images, obtain candidate matching objects of the target object in the image in the other images from the target candidate matching data, obtain the Euclidean distance d' between the epipolar line of the target object in the other image and the center point coordinates of the candidate matching objects in the other images, obtain the normalized Euclidean distance d', and record it as the geometric cost value between the target object in the image and the candidate matching objects in the other images; Step S42: Use a preset object detection model to obtain the bounding boxes of each target object in the said image and other images, crop and resize the bounding boxes in the said image and other images, and input them into a preset feature extraction model to obtain the feature vectors of each target object in the said image and other images, and normalize the feature vectors of each target object in the said image and other images; Step S43: Obtain the feature vectors G of the w-th target object in the image and the v-th target objects in other images respectively w and G ▽ v , calculate the feature cost value Q between the w-th target object and the v-th target object (w,v) : , Step S44: When the v-th target object is not a candidate matching object of the w-th target object, set the comprehensive cost value of the w-th target object and the v-th target object to a preset maximum value; When the vth target object is a candidate matching object of the wth target object, calculate the comprehensive cost value U of the wth target object and the vth target object (w,v) : , Among them, D (w,v) is the geometric cost between the vth target object and the wth target object; ψ Q is the preset feature cost weight coefficient; ψ D is the preset geometric cost weight coefficient; ψ Q +ψ D =1,ψ Q >0,ψ D >0; Step S45: Obtain the comprehensive cost values between each target object in the said image and several target objects in other images, construct a target cost matrix between the said image and other images, obtain the total number m of each target object in the said image, and obtain the total number n of several target objects in other images; When m = n, do not process the target cost matrix. When m ≠ n, fill the target cost matrix. The specific filling process is as follows: When m > n, fill m - n virtual columns in the last column of the target cost matrix, where the comprehensive cost values of each element in the m - n virtual columns are all preset maximum values; When m < n, fill n - m virtual rows in the last row of the target cost matrix, where the comprehensive cost values of each element in the n - m virtual columns are all preset maximum values; Step S46: Use the Hungarian algorithm to solve the target cost matrix, obtain the zero elements in each row and column of the feature target cost matrix, and obtain the matching pairs of the positions corresponding to the zero elements in the feature target cost matrix. Among them, a matching pair includes the target object in the said image that matches the target object in other images; Obtain the target objects in other images that match each target object in the said image, and gather them to obtain the target object matching data between the said image and other images.
6. The target coordinate positioning method based on a multi-view panoramic camera according to claim 5, characterized in that: The said step S5 includes: Step S51: Obtain each image with the same shooting time in the multi-view panoramic camera, obtain the target object matching data between each image, and obtain the target object in one image that matches the target object in another image from the target object matching data between one image and another image; Step S52: Obtain the camera parameters of each camera unit corresponding to the said each image, and combine the target object matching data between each image to perform image feature fusion on each image to obtain a fused image; Step S53: Locate each target object in the fused image to obtain the three-dimensional coordinates of each target object in the fused image. Among them, the coordinate systems of the three-dimensional coordinates are all in a preset world coordinate system, and gather the three-dimensional coordinates of each target object in the fused image to obtain the target coordinate data in the fused image.
7. A target coordinate positioning system based on a multi-view panoramic camera, used to execute the target coordinate positioning method based on a multi-view panoramic camera according to any one of claims 1 to 6, characterized in that: The said system includes a camera spatial relationship evaluation module, a target matching module, and a target positioning module; The camera spatial relationship evaluation module is used to evaluate the spatial geometric relationship between different camera units to obtain the spatial epipolar line data between different camera units; The target matching module is used to obtain target object data of the image and other images, obtain candidate matching objects of the target object of the image in other images based on the target object data, obtain target candidate matching data, construct and solve the target cost matrix between the image and other images, and obtain target object matching data between the image and other images; The target positioning module is used to obtain target object matching data between images taken at the same time in a multi-eye panoramic camera, perform image feature fusion on the images to obtain a fused image, obtain the three-dimensional coordinates of the target object in the fused image, and obtain the target coordinate data in the fused image.
8. The target coordinate positioning system based on a multi-view panoramic camera according to claim 7, characterized in that: The camera space relationship evaluation module includes a basic matrix construction unit and a camera space relationship evaluation unit; The basic matrix construction unit is used to obtain camera parameter data of different camera units in the multi-eye panoramic camera and construct a basic matrix between different camera units; The camera spatial relationship evaluation unit is used to evaluate the spatial geometric relationship between different camera units based on the basic matrix between different camera units, and obtain spatial epipolar line data between different camera units.
9. The target coordinate positioning system based on a multi-view panoramic camera according to claim 7, characterized in that: The target matching module includes a target detection unit, a candidate matching acquisition unit and a target matching unit; The target detection unit is used to detect target objects in the image and screen the target objects to be determined in the image in combination with historical images to obtain target object data in the image; The candidate matching acquisition unit is configured to acquire target object data from other images captured at the same time as the image, acquire spatial epipolar line data between the camera unit of the image and the camera unit of the other images, acquire candidate matching objects of the target object of the image in the other images, and obtain target candidate matching data; The target matching unit is used to construct and solve the target cost matrix between the image and other images based on the target candidate matching data between the image and other images, so as to obtain the target object matching data between the image and other images.
10. The target coordinate positioning system based on multi-view panoramic cameras according to claim 7, characterized in that: The target positioning module includes an image fusion unit and a target positioning unit; The image fusion unit is used to acquire images taken at the same time by the multi-eye panoramic camera, obtain target object matching data between the images, and fuse the image features of the images to obtain a fused image; The target positioning unit is used to acquire and collect the three-dimensional coordinates of the target object in the fused image to obtain the target coordinate data in the fused image.
Citation Information
Patent Citations
Multi-camera collaboration-based method for detecting, positioning and tracking unmanned aerial vehicle
CN104197928A
Cross-camera multi-view scene target continuous tracking and re-identification positioning method
CN116245919A
Multi-target tracking method and system in panoramic imaging
CN116612147A
Multi-target visual tracking method and device based on camera parameters, equipment and medium
CN116777950A
Multi-modal ship target association method based on multi-feature fusion
CN118675022A
Cited By
Security target panoramic tracking method and system based on multi-source video data fusion
CN121837315A
Security target panoramic tracking method and system based on multi-source video data fusion
CN121837315B