Wharf cross-camera multi-target tracking method and system based on three-dimensional map

By constructing a 3D model map in the dock monitoring system and performing voxelization processing, combined with target detection and cross-camera trajectory matching, the problem of target localization and tracking in complex scenes by a monocular camera was solved, achieving accurate 3D target localization and robust cross-camera tracking, thus improving the reliability and robustness of the system.

CN120876543AActive Publication Date: 2025-10-31江苏省港口集团信息科技有限公司
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511376163.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2025-10-31
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

Existing monocular camera-based dock monitoring systems struggle to provide accurate 3D target localization and stable cross-camera tracking in complex scenarios, resulting in frequent breakage and incorrect correlation of target trajectories, and failing to effectively utilize the 3D spatial structure of the dock.

Method used

By constructing a 3D model map of the dock and performing voxelization, and combining target detection, monocular tracking, and cross-camera trajectory matching, a multi-dimensional cost function is established using 3D spatial information and re-identification features to achieve accurate positioning and robust tracking of the target in 3D space.

Benefits of technology

It significantly improves target positioning accuracy and tracking system reliability, ensures the physical authenticity of generated trajectories, reduces reliance on single visual features, and enhances robustness in low-quality video and complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876543A_ABST
    Figure CN120876543A_ABST
Patent Text Reader

Abstract

The invention discloses a wharf cross-camera multi-target tracking method and system based on a three-dimensional map, and belongs to the technical field of intelligent monitoring. According to the method, voxelization processing is carried out on a wharf three-dimensional model, color and texture feature matching is carried out on the wharf three-dimensional model and a two-dimensional image in a real-time monitoring video, and a detected target is accurately projected to a three-dimensional space. In a single camera, an optimized ByteTrack algorithm is combined with Kalman filtering to carry out continuous tracking. When a target crosses different cameras, high-robustness identity matching is carried out through a cost function which fuses a re-identification (ReID) feature, a three-dimensional space distance, a speed constraint and a target occurrence frequency weight. According to the invention, through deep fusion of three-dimensional geographic information and visual analysis, the target positioning precision and the stability of cross-mirror tracking are significantly improved, the physical authenticity of the tracking trajectory is ensured, and the method is especially suitable for complex scenes such as wharfs and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent monitoring technology, specifically to a method and system for multi-target tracking across cameras at a dock based on a 3D map. Background Technology

[0002] With the continuous advancement of smart port construction, real-time, accurate, and continuous monitoring of various moving targets such as ships, vehicles, and personnel within the terminal area has become a core requirement for ensuring operational safety and efficiency. Among existing monitoring technologies, visual solutions based on monocular cameras have gained widespread application due to their flexible deployment and relatively low cost. However, when directly applying such solutions to the open and complex environment of a terminal, their inherent limitations are becoming increasingly apparent.

[0003] First, traditional monocular vision, essentially analyzing two-dimensional images, inherently lacks depth information. This results in a severe deficiency in target localization accuracy, making it difficult to meet the requirements of refined port management for accurate physical spatial location. Second, in large-scale scenarios covered by multiple cameras, when a target disappears from the field of view of one camera and enters another, cross-camera target re-identification (ReID) technology is needed to maintain the continuity of its identity. Existing ReID methods heavily rely on clear and stable visual features, but in common situations at ports such as long distances, changing lighting, and target occlusion, visual features are easily blurred or rendered ineffective, leading to tracking and matching failures and frequent breaks in the target trajectory. More importantly, these purely visual information-based methods fail to effectively utilize the inherent three-dimensional spatial structure of ports. The generated motion trajectories lack constraints from the physical world, often resulting in illogical associations such as vehicles passing through walls or ships landing, which seriously affects the reliability of the entire monitoring system.

[0004] To address the challenges of cross-camera tracking, the industry has explored various approaches. For example, Chinese patent CN107689054B discloses a method for constructing a multi-camera topological connectivity graph. This method utilizes clustering algorithms to determine the camera's entry and exit regions and estimate transition times, thereby assisting in cross-camera tracking. While this approach provides a solution, it fundamentally remains reliant on 2D visual information. Its depth estimation still relies on changes in the size of 2D detection boxes, resulting in limited accuracy. Furthermore, the robustness of its feature matching significantly decreases when the target is far from the camera or when the feature pixel count is too low, indicating that problems persist.

[0005] Therefore, how to get rid of the over-reliance on single visual features, deeply integrate the existing three-dimensional geographic information of the dock with two-dimensional video data, and design a cross-camera tracking mechanism that remains robust in low-definition and complex environments is a key technical problem that urgently needs to be solved in the field of intelligent monitoring. Summary of the Invention

[0006] To overcome the existing problems and shortcomings, this invention proposes a multi-target tracking method for docks across cameras based on 3D maps, comprising the following steps:

[0007] S1. 3D map processing: Obtain a 3D model map of the dock, and perform voxelization processing on the 3D model map to generate a voxel cloud containing 3D coordinates and color or texture information;

[0008] S2. Target Detection: Perform real-time target detection on surveillance video acquired by at least one camera, and generate a detection result for each detected target, including its location information and confidence level in a two-dimensional image;

[0009] S3. Three-dimensional spatial mapping: Based on the pre-calibrated camera intrinsic and extrinsic parameters, by matching the color or texture features of the two-dimensional image region with the voxel cloud, a mapping relationship between the surveillance video and the three-dimensional map is established, and the target in the detection result is projected into the three-dimensional map to obtain the three-dimensional spatial coordinates of the target;

[0010] S4. Monocular tracking: Within the field of view of a single camera, the motion trajectory of the target is predicted using Kalman filtering, and the same target in consecutive video frames is continuously tracked and its trajectory status is managed through data association algorithms.

[0011] S5. Cross-camera trajectory matching: When the target switches between different camera views, a matching cost function is constructed that integrates target re-identification (ReID) feature distance, 3D spatial distance and 3D spatial velocity constraints to calculate the trajectory correlation under different cameras and achieve cross-camera target identity matching.

[0012] Furthermore, the process of projecting the target onto the 3D map specifically includes:

[0013] Extract voxel data within the camera's field of view and statistically analyze its color histogram and texture features;

[0014] A sliding window is used on the two-dimensional image to calculate the color histogram and texture features of each window sub-region;

[0015] By calculating the color and texture similarity between voxel regions and image sub-regions, the best matching region is found, thereby establishing the correspondence between voxel clouds and image regions.

[0016] Furthermore, the target detection generates a three-dimensional detection box that includes center coordinates, length, width, and height.

[0017] Furthermore, the continuous tracking within the field of view of a single camera is optimized based on the ByteTrack algorithm, which associates high-confidence detection boxes through the first matching and low-confidence detection boxes through the second matching to enhance the tracking stability of occluded targets.

[0018] Furthermore, the cost function for cross-camera trajectory matching also includes a global weight calculated based on the target's historical trajectory data, which is used to adjust the probability of different targets appearing under a specific camera.

[0019] Furthermore, the ReID features are extracted through a lightweight backbone network, and the average feature value of the target in the most recent few frames is used as its template feature.

[0020] A multi-target tracking system for a dock based on a 3D map, characterized in that it includes:

[0021] The 3D coordinate mapping module is used to voxelize the 3D model map and establish its matching relationship with the 2D monitoring image, projecting the detected target into 3D space;

[0022] The target detection module is used to detect targets from surveillance video in real time.

[0023] A monocular tracking module is used for continuous trajectory tracking of a target within a single camera;

[0024] The cross-camera matching module is used to match the same target across different cameras based on a cost function that integrates re-identification features, 3D spatial distance, and velocity constraints.

[0025] The trajectory management and visualization module is used to maintain and update the trajectory status of all targets and to visualize it on a 3D map.

[0026] Furthermore, the target detection module includes a feature generation module, a feature fusion module that integrates a path aggregation network (PANet) and a bidirectional feature pyramid (BiFPN) structure, and a feature detection module that outputs a 3D detection box.

[0027] Beneficial effects:

[0028] This invention significantly improves the accuracy of target localization and tracking. By voxelizing the 3D model of the dock and establishing precise matching relationships between the voxel cloud and the 2D image in multiple dimensions such as color and texture, this method can directly and accurately project targets detected in videos into the real 3D physical space. This transformation fundamentally solves the problem of traditional monocular vision lacking depth information, making target localization and trajectory tracking no longer a simulation of a 2D plane, but a true reproduction in 3D space, greatly improving accuracy.

[0029] In the cross-camera target matching process, this invention significantly enhances tracking robustness in low-quality video and complex environments by constructing a comprehensive cost function that integrates multi-dimensional information. This function innovatively combines factors such as 3D spatial distance, 3D spatial velocity constraints, and even the weight of high-frequency target occurrence areas based on dock operation logic, with traditional ReID visual features in a weighted combination. This multimodal fusion mechanism greatly reduces the dependence on a single high-quality visual feature. Even in extreme cases where target features are blurred, the system can achieve reliable matching through more robust spatial and logical constraints, effectively avoiding trajectory breakage.

[0030] This invention ensures the physical authenticity of the final generated trajectory by placing all tracking activities within a unified framework of a 3D map. The 3D map not only provides precise coordinates but also rich spatial semantic information, such as road, building, and water boundaries. Utilizing this information, this method can automatically filter out a large number of abnormal trajectories that violate physical principles due to mismatches, ensuring that the output tracking results fully conform to the logic of the actual scene and significantly improving the credibility and reliability of the entire system. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a schematic diagram of the target detection framework in an embodiment of the present invention.

[0033] Figure 2 This is an overall flowchart of the multi-target tracking method for a dock based on a 3D map in an embodiment of the present invention.

[0034] Figure 3 This is a flowchart illustrating the specific process of the monocular tracking module in an embodiment of the present invention. Detailed Implementation

[0035] The present application will be described below with reference to specific embodiments:

[0036] Example 1:

[0037] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and a specific embodiment. It should be understood that the specific embodiment described herein is merely illustrative and not intended to limit the scope of protection of this invention.

[0038] The following description, in conjunction with the accompanying drawings and embodiments, further illustrates the present invention. A specific flowchart is shown below. Figure 2 As shown.

[0039] A method and system for multi-target tracking across cameras at a dock based on a 3D map, the system comprising five core modules:

[0040] Target detection module: Real-time multi-target detection based on single-stage network. Targeting the small target features in the monitoring perspective, it fully considers the structural advantages of Path Aggregation Network (PANet) and Bidirectional Feature Pyramid (BiFPN) to construct a new cross-scale up-down fusion module, which outputs two-dimensional detection boxes for different targets.

[0041] Monocular tracking module: Optimizes the ByteTrack algorithm to achieve multi-target tracking of 3D targets within a single camera;

[0042] 3D coordinate mapping module: Performs voxelization on the 3D model map, uses feature matching algorithm to construct the matching relationship between the two, and generates the 3D coordinate information of the target in the 3D map based on the preset target size;

[0043] Cross-camera matching module: Integrates target re-identification algorithm (ReID) feature distance to establish cross-camera trajectory association with three-dimensional spatial constraints based on dock operation logic;

[0044] The trajectory management and visualization module maintains the target trajectory status, performs update, deletion, and creation operations on trajectories, maintains trajectory matching effects in monocular and multi-view tracking, and outputs 3D map visualization results. The following flowchart describes the process step by step.

[0045] First, a corresponding object detection framework is established. Considering the need for subsequent matching with a 3D map, the designed detection framework generates 3D detection results during the regression phase. The object detection framework consists of three modules: a feature generation module (BackBone), a feature fusion module (Neck), and a feature detection module (Head), as shown in the specific structure below. Figure 1 As shown in the diagram, the feature generation module mainly consists of multiple convolutional structures that fully characterize the input image, generating feature layers of different sizes. Three feature layers of different scales are derived from the feature generation module to detect targets of different sizes. The feature fusion module leverages the advantages of the PANet (Path Aggregation Network) and Bidirectional Feature Pyramid (BiFPN) structures to construct a new cross-scale fusion module. The overall network structure consists of two sets of upsampling models and one set of downsampling models, including intermediate layers for feature matching and fusion between different levels. The overall structure is a top-down and bottom-up network, as detailed below. Figure 1 As shown in the Neck diagram, feature information from feature layers of different sizes is fully integrated. The feature detection module consists of detection heads of different scales, each including a classifier and a regressor, used to generate feature detection structures of different sizes. The classifier generates the corresponding type, and the regressor obtains the corresponding detection results. The final result structure is d = (cent_x, cent_y, length, width, height, confidence). The regressor directly generates the 3D information of the detection box, including the center coordinates of the detection box, its corresponding length, width, height, and confidence score. Non-maximum suppression is used to remove detection results that do not meet the requirements. The 3D coordinate detection results are used in subsequent monocular matching and 3D map matching processes.

[0046] Based on the acquisition of the target's 3D coordinates, a monocular tracking module is introduced. This module is responsible for the continuous tracking of the target within a single camera. The core process includes four main steps: tracker initialization, Kalman filter prediction, data association (initial matching + secondary matching), trajectory update, and state management. The specific process is as follows: Figure 3 As shown. Based on ByteTrack, the dimensional information is optimized to establish a target tracking system for 3D detection results. The specific steps are as follows:

[0047] First, initialize the tracker and obtain the detection results mentioned earlier. Each element contains the detection results of a structure (cent_x, cent_y, length, width, height, conf). For high-confidence detection results in the initial frame, a corresponding tracker is created, and the corresponding motion vector is initialized. Covariance matrix Define the state transition matrix. The motion vector is used to describe the changing trend of each parameter.

[0048] For each tracker in the previous step, Kalman filter prediction is performed. Specifically, for each active tracker (in "active" state), based on the state of the previous frame... To predict the detection results in the next frame

[0049]

[0050] This includes the current trajectory prediction result for the next frame.

[0051] Regarding the test results Select high-confidence detection boxes from the detection results and calculate the relationship between them and different prediction boxes. The optimal matrix is ​​matched using the Hungarian algorithm, where 1- Let be the cost matrix. If a match is found, the tracker is updated; otherwise, the tracker that failed to match is marked. Similarly, select detection boxes with lower confidence and use the Hungarian algorithm again to match them with the remaining detection boxes.

[0052] Based on the matching results from the previous step, the following status update options are selected: 1. For trackers that have a successful match, update the corresponding status; 2. For trackers that have failed to match twice, mark them as not inactive. For trackers that are continuously marked as not inactive, terminate the corresponding tracker. Trackers that have a successful match are marked as active.

[0053] Considering the uncertain lifespan of current surveillance cameras, the 2D coordinates of the cameras are calibrated to improve the accuracy of 3D map matching. Given the fixed camera positions and large number of cameras, Zhang's calibration method is chosen for parameter calculation. A chessboard calibration board larger than 40cm*40cm is generated, and more than 20 photos are taken from different angles using the cameras. OpenCV sub-pixel corner detection is used to obtain pixel coordinates. If we set the plane containing the chessboard squares to Z=0, then we can set the world coordinate system of the corner points of the chessboard to be... After obtaining the pixel coordinate system and world coordinate system, the chessboard is solved using the least squares method:

[0054]

[0055] Obtain the corresponding intrinsic parameter matrix And the corresponding distortion matrix P.

[0056] After calibrating the camera, the next step is to match the 2D target with the 3D map. In practice, 3D mapping is commonly used... Figure 1 Generally, 3D model maps are created using methods such as oblique photogrammetry. While creating 3D maps from point clouds offers higher accuracy, it also incurs higher costs and is therefore less common. Furthermore, 3D models generated through oblique photogrammetry and post-processing cannot be used directly. Therefore, to match them with 2D images, voxelization is performed on the 3D model.

[0057] Generally, 3D model maps are based on triangular meshes. Here, we choose to convert the model to a voxel cloud by mesh voxelization. Let the side length of the selected voxel be... Find the minimum coordinate value of the triangular mesh model in 3D space. and maximum coordinate value This determines the bounding box encompassing the entire model. A regular voxel mesh is constructed in 3D space based on the voxel resolution and voxelization range. The number of voxels in the x, y, and z directions of the voxel mesh is determined. Voxels are uniformly divided within the corresponding range, and each voxel has a unique 3D index (i, j, k) corresponding to its position in the voxel mesh.

[0058] Each triangular facet in the triangular mesh model is processed sequentially. Each triangle is defined by three vertices (v1, v2, v3), each containing 3D coordinate information. For each triangle, an intersection test is performed on each voxel in the voxel mesh. To simplify the calculation process, the separation axis theorem is used to determine the intersection, checking whether the projections of the triangle and voxel on each coordinate axis overlap. If the projections of both overlap on all possible separation axes, then the triangle and voxel intersect; if they do not overlap on any separation axis, then they do not intersect. Voxels that intersect with the triangular mesh are marked as "occupied," while those that do not intersect are marked as "empty." This state marking is the most basic attribute of the voxel, used to distinguish between the interior and exterior spaces of the model. All point clouds in the occupied state are further preserved. Furthermore, after determining the intersection state, the color channels (R, G, B) contained in the triangular mesh can be further fused into the voxel information.

[0059] When the approximate area captured by the camera in a specific scenario is known on the map, extracting voxel data of the camera's field of view can significantly reduce the time of subsequent algorithm matching processes and improve accuracy.

[0060] Based on the above data processing results, voxel data of the corresponding 3D model from the monitoring perspective can be obtained. The range of this voxel data should be slightly larger than the monitoring range. Subsequent data matching between the two is then performed. Since the selected target voxels overlap significantly with the monitored area, regions can be extracted to achieve matching. The matching criteria are color information and corresponding texture information.

[0061] For a selected voxel region, the color information of the voxels within it is statistically analyzed. The proportion of voxels in each color interval within the voxel region is calculated to form a three-dimensional color histogram.

[0062] The voxel region is converted to grayscale representation. Using the central voxel of the voxel region as a reference, different offsets are set, and the grayscale co-occurrence matrix is ​​calculated. The grayscale co-occurrence matrix records the frequency of different grayscale value pairs occurring simultaneously under specific offset conditions. Texture feature values ​​such as energy, contrast, correlation, and entropy are calculated from the grayscale co-occurrence matrix.

[0063] For the input image, the RGB color of each pixel is statistically analyzed in the same way as the voxel region color histogram calculation to construct the image's color histogram. The image color histogram is based on two-dimensional pixels and needs to be dynamically updated as the sliding window moves.

[0064] The image is converted to grayscale. For each sub-region within the sliding window, the gray-level co-occurrence matrix and corresponding texture feature values ​​are calculated. The size of the sliding window needs to be adjusted according to actual needs; generally, a window size similar to or slightly larger than the voxel region is chosen to ensure that possible matching regions are fully covered.

[0065] The similarity between the color histogram of a voxel region and the color histogram of a sub-region of the image is calculated using Bach's distance. A smaller Bach's distance indicates greater similarity between the two color histograms, meaning their color distributions are more similar. The specific expression is:

[0066]

[0067] Where P and Q are the color histograms of the voxel region and the image sub-region, respectively, and n is the number of color intervals.

[0068] For texture features, since there are multiple texture feature values, weighted Euclidean distance is used to calculate similarity. First, each texture feature value is normalized to ensure it falls within the same numerical range. The formula for calculating the weighted Euclidean distance is:

[0069]

[0070] in The weights of the corresponding features, and These are the texture feature values ​​for the voxel region and the image sub-region.

[0071] Color similarity and texture similarity are weighted and summed according to a certain ratio to obtain the comprehensive similarity S. The larger the comprehensive similarity S value, the higher the similarity between the voxel region and the image sub-region.

[0072] After the sliding window traverses the entire image, the overall similarity calculated for all window positions is compared. The window position with the highest overall similarity is found, and the image sub-region corresponding to this window is the region that best matches the voxel region. This establishes the correspondence between the voxel region and this image sub-region, completing the region-based matching process.

[0073] Combining the (cent_x,cent_y,length,width,height,conf) information calculated earlier, the target location information obtained in the image can be directly represented in the voxel cloud. Since the coordinates of the voxel cloud are consistent with those of the 3D map, this information can also be used to directly project the detected target onto the 3D map.

[0074] The matching results above allow the 3D detected target to be directly marked on a 3D map. Subsequent fusion of the target's 2D image and 3D coordinate information enables cross-camera trajectory matching. Cross-camera trajectory matching primarily relies on a 2D Re-identification algorithm (ReID) and 3D spatial window constraints. Considering the performance bottlenecks among multiple algorithms, a lightweight network, ResNet-34, is used as the backbone network for feature extraction in the Re-identification algorithm. The specific detection process is as follows:

[0075] In the historical tracker, the most recent 5 high-confidence target ROIs were extracted, and the range of interest for each was more than 20% larger than the detection box.

[0076] After normalizing the corresponding ROI by scaling it to 224*224 using bilinear interpolation, the result is output into the corresponding ReID network.

[0077] The average value of the features from the 5 frames is output as the target template feature.

[0078] In the previous step, we obtained the feature values ​​of the targets and calculated the matching cost by combining spatial information. The matching relationship between targets is determined by a cost function. It is determined by three parts:

[0079]

[0080] in , , The weights are the corresponding module losses, where As a feature loss, the ReID algorithm obtains the feature vector distance between different targets. This is the 3D spatial loss, used to calculate its relative position loss in the world coordinate system. For velocity loss, we calculate whether the target's maximum velocity in 3D space meets the displacement distance requirement. Simultaneously, considering that in the actual generation of the dock, there is generally a fixed path and different targets appear frequently in different scenarios, a global weight 't' is introduced to determine the probability weight of different targets appearing under different cameras, combining information on different target types. If a target type appears frequently under certain cameras in past trajectories, this value will be higher, and vice versa.

[0081] In the preceding text, we discussed the use of monocular and multi-view matching for trajectory tracking. Each time a target is detected by a monocular camera, multi-view target matching is performed, and the trajectory state transition follows a defined update rule. Firstly, a trajectory has three states: active, inactive, and terminated in monocular mode. A trajectory in a matching state in monocular mode is active, while a trajectory not in a matching state is inactive. If the number of inactive frames exceeds 30, it is terminated. Each target has an independent tracker in monocular mode, and each tracker has a unique ID, indicating that the target is the same under that tracker. Different trackers also have corresponding tracking IDs in multi-view mode, and these IDs are shared. That is, trackers are independent, but IDs are shared. For trackers that can be matched by multi-view, if not all trackers are in a terminated state, it is considered that the target has left the dock, and all trackers are discarded. Otherwise, all trackers that can be matched are retained.

[0082] Any person skilled in the art can make various corresponding changes or modifications without departing from the core ideas and principles of the present invention, and such changes or modifications should all fall within the protection scope of the appended claims. Therefore, the patent protection scope of the present invention should be determined by the appended claims.

Claims

1. A method for multi-target tracking across cameras at a dock based on a 3D map, characterized in that, Includes the following steps: S1. 3D map processing: Obtain a 3D model map of the dock, and perform voxelization processing on the 3D model map to generate a voxel cloud containing 3D coordinates and color or texture information; S2. Target Detection: Perform real-time target detection on surveillance video acquired by at least one camera, and generate a detection result for each detected target, including its location information and confidence level in a two-dimensional image; S3. Three-dimensional spatial mapping: Based on the pre-calibrated camera intrinsic and extrinsic parameters, by matching the color or texture features of the two-dimensional image region with the voxel cloud, a mapping relationship between the surveillance video and the three-dimensional map is established, and the target in the detection result is projected into the three-dimensional map to obtain the three-dimensional spatial coordinates of the target; S4. Monocular tracking: Within the field of view of a single camera, the motion trajectory of the target is predicted using Kalman filtering, and the same target in consecutive video frames is continuously tracked and its trajectory status is managed through data association algorithms. S5. Cross-camera trajectory matching: When the target switches between different camera views, a matching cost function is constructed that integrates target re-identification (ReID) feature distance, 3D spatial distance and 3D spatial velocity constraints to calculate the trajectory correlation under different cameras and achieve cross-camera target identity matching.

2. The method according to claim 1, characterized in that, Step S3, which involves projecting the target from the detection results onto the 3D map, specifically includes: Extract voxel data within the camera's field of view and statistically analyze its color histogram and texture features; A sliding window is used on the two-dimensional image to calculate the color histogram and texture features of each window sub-region; By calculating the color and texture similarity between voxel regions and image sub-regions, the best matching region is found, thereby establishing the correspondence between voxel clouds and image regions.

3. The method according to claim 1 or 2, characterized in that, In step S2, the target detection generates a three-dimensional detection box containing the center coordinates, length, width, and height.

4. The method according to claim 1, characterized in that, In step S4, continuous tracking within the field of view of a single camera is optimized based on the ByteTrack algorithm. This involves first-matching and associating high-confidence detection boxes, and second-matching and associating low-confidence detection boxes to enhance the tracking stability of occluded targets.

5. The method according to claim 1, characterized in that, The cost function for cross-camera trajectory matching in step S5 also includes a global weight calculated based on the target's historical trajectory data. This global weight is used to adjust the probability of different targets appearing under a specific camera.

6. The method according to claim 1, characterized in that, In step S5, the re-identified features are extracted through a lightweight backbone network, and the average feature value of the target in the most recent frames is used as its template feature.

7. A multi-target tracking system for a dock based on a 3D map, characterized in that, include: The 3D coordinate mapping module is used to voxelize the 3D model map and establish its matching relationship with the 2D monitoring image, projecting the detected target into 3D space; The target detection module is used to detect targets from surveillance video in real time. A monocular tracking module is used for continuous trajectory tracking of a target within a single camera; The cross-camera matching module is used to match the same target across different cameras based on a cost function that integrates re-identification features, 3D spatial distance, and velocity constraints. The trajectory management and visualization module is used to maintain and update the trajectory status of all targets and to visualize it on a 3D map.

8. The system according to claim 7, characterized in that, The target detection module includes a feature generation module, a feature fusion module that integrates a path aggregation network (PANet) and a bidirectional feature pyramid (BiFPN) structure, and a feature detection module that outputs a 3D detection box.

Citation Information

Patent Citations

  • A method for multi-camera topology connectivity graph construction and cross-camera target tracking

    CN107689054B

  • Moving target cross-lens tracking method based on three-dimensional calibration

    CN116402857A

  • Cross-view multi-target real-time trajectory tracking method and system

    CN116580107A

  • Multi-target tracking method for sea surface scene

    CN118941595A

  • Target tracking method and device based on multi-camera coordinate conversion, computer equipment and computer program product

    CN119295509A