A dynamic target tracking method based on multiple cameras and related equipment
By introducing timestamps and real-time updated parameters into a multi-camera system, combined with dynamic target prediction and camera orientation adjustment, the problems of prediction loss and environmental adaptability in dynamic target tracking are solved, achieving accurate and robust tracking of dynamic targets.
Patent Information
- Application Number
- CN202511177669.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing dynamic target tracking systems based on multiple cameras have shortcomings in dynamic target prediction and environmental adaptability, leading to problems such as target loss and trajectory breakage.
By introducing timestamped frame images and real-time updated camera parameters, combined with a dynamic target position prediction mechanism and camera orientation adjustment, accurate tracking of dynamic targets is achieved.
It improves the accuracy and robustness of dynamic target tracking, ensures continuous tracking in complex environments, and avoids target loss and trajectory breakage.
Smart Images

Figure CN120673089B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent transportation technology, and more specifically, to a dynamic target tracking method and related equipment based on multiple cameras. Background Technology
[0002] With the widespread application of intelligent transportation systems, dynamic target tracking technology based on multiple cameras has become a core requirement for scenarios such as vehicle monitoring and behavior analysis. Existing cross-camera target tracking systems typically rely on a three-dimensional spatial coordinate system for target localization, but in practical applications, they face the following technical bottlenecks:
[0003] Lack of dynamic target prediction: Traditional technologies primarily associate target positions with static coordinates, lacking effective modeling of the time dimension. When a target vehicle moves at high speed, existing systems struggle to accurately predict its future trajectory, leading to target loss or tracking interruptions during cross-camera tracking. For example, in scenarios like highways, the inability to predict the vehicle's position in the next moment can cause camera switching to lag, resulting in a break in the tracking chain.
[0004] Poor environmental adaptability: Existing technologies often rely on pre-calibrated and fixed camera extrinsic parameters. However, in complex real-world application environments, factors such as sudden changes in lighting conditions, weather interference like rain and fog, or temporary obstruction can cause the mapping relationship of these fixed parameters to become inaccurate. For example, when camera images are overexposed due to strong light, traditional calibration parameters may become invalid, increasing the coordinate mapping error of the target between adjacent cameras, which in turn causes trajectory breakage and affects the robustness and tracking accuracy of the system.
[0005] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0006] The purpose of this application is to provide a dynamic target tracking method and related equipment based on multiple cameras, which can effectively predict the position of dynamic targets and adaptively adjust camera parameters to improve environmental adaptability, thereby achieving accurate and robust tracking of dynamic targets.
[0007] In a first aspect, this application provides a dynamic target tracking method based on multiple cameras, which tracks vehicles using a multi-camera tracking system with multiple cameras having adjustable orientations, including the following steps:
[0008] A1. Real-time acquisition of timestamped frame images captured by the current camera, identification of the target vehicle from them, and acquisition of the identification features of the target vehicle; the current camera is a camera whose field of view covers the target vehicle; the identification features include pixel position coordinates;
[0009] A2. Based on the real-time intrinsic and extrinsic parameters of the current camera, the pixel position coordinates are converted into world coordinates to generate the target vehicle motion trajectory segment corresponding to the current camera; wherein, if the current orientation of the current camera is the initial orientation, the real-time intrinsic and extrinsic parameters are the intrinsic and extrinsic parameters pre-calibrated under the initial orientation;
[0010] A3. When the target vehicle enters the field of view boundary area of the current camera, predict the world coordinate position of the target vehicle after a preset buffer time based on the target vehicle's motion trajectory segment corresponding to the current camera, and record it as the predicted target position;
[0011] A4. If the predicted target position is outside the field of view of the multi-camera tracking system, proceed to step A6; otherwise, proceed to step A5.
[0012] A5. Determine a nearby camera based on the predicted target location, adjust the orientation of the nearby camera to align with the predicted target location, update the intrinsic and extrinsic parameters of the nearby camera, and use the nearby camera as the new current camera, then return to step A1;
[0013] A6. Merge all the target vehicle motion trajectory segments to obtain the total motion trajectory.
[0014] Secondly, this application provides an electronic device, including a processor and a memory, wherein the memory stores a computer program executable by the processor, and when the processor executes the computer program, it performs the steps of the dynamic target tracking method based on multiple cameras as described above.
[0015] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the multi-camera-based dynamic target tracking method described above.
[0016] Beneficial effects: The dynamic target tracking method and related equipment based on multiple cameras provided in this application effectively solve the problems of missing dynamic target prediction and poor environmental adaptability in the prior art by introducing a target position prediction mechanism based on the time dimension and combining it with a strategy of dynamically adjusting the camera orientation and updating intrinsic and extrinsic parameters in real time. This improves the accuracy and robustness of dynamic target tracking and has the advantages of being able to effectively predict the position of dynamic targets and adaptively adjust camera parameters to improve environmental adaptability, thereby achieving accurate and robust tracking of dynamic targets. Attached Figure Description
[0017] Figure 1 A flowchart of a dynamic target tracking method based on multiple cameras provided in this application.
[0018] Figure 2 This is a schematic diagram of the structure of an electronic device provided in this application.
[0019] Labeling explanations: 301, processor; 302, memory; 303, communication bus. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0021] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0022] Please refer to Figure 1 This application discloses a dynamic target tracking method based on multiple cameras, which tracks vehicles using a multi-camera tracking system with multiple cameras having adjustable orientations. The method includes the following steps:
[0023] A1. Real-time acquisition of timestamped frame images captured by the current camera, identification of the target vehicle, and acquisition of the target vehicle's identification features; the current camera is the camera whose field of view covers the target vehicle; the identification features include pixel position coordinates;
[0024] A2. Based on the real-time intrinsic and extrinsic parameters of the current camera, convert the pixel position coordinates into world coordinates to generate the target vehicle motion trajectory segment corresponding to the current camera; wherein, if the current orientation of the current camera is the initial orientation, the real-time intrinsic and extrinsic parameters are the intrinsic and extrinsic parameters pre-calibrated under the initial orientation;
[0025] A3. When the target vehicle enters the field of view boundary area of the current camera, predict the world coordinate position of the target vehicle after a preset buffer time based on the target vehicle's motion trajectory segment corresponding to the current camera, and record it as the predicted target position;
[0026] A4. If the predicted target location is outside the field of view of the multi-camera tracking system, proceed to step A6; otherwise, proceed to step A5.
[0027] A5. Determine a nearby camera based on the predicted target location, adjust the orientation of the nearby camera to align with the predicted target location, update the intrinsic and extrinsic parameters of the nearby camera, and use the nearby camera as the new current camera. Return to step A1.
[0028] A6. Merge all target vehicle trajectory segments to obtain the total trajectory.
[0029] Among them, a multi-camera tracking system refers to a monitoring network composed of multiple cameras, which can be implemented through distributed deployment or centralized management. For example, it can be deployed at urban traffic intersections or along highways. Its main purpose is to achieve continuous coverage and tracking of targets in a large area.
[0030] Among them, the adjustable-direction camera refers to a camera whose shooting direction can be adjusted according to instructions. It can be achieved by using an electric pan-tilt head or a robotic arm control, such as controlling its rotation and pitch angles through a stepper motor or servo motor. Its main purpose is to achieve flexible capture of dynamic targets and switching of field of view.
[0031] Among them, the frame image with timestamp means that each frame image is accompanied by precise time information. This can be achieved by using system clock synchronization or GPS timing module, such as embedding the acquisition time in the image data packet. Its main purpose is to realize the time dimension modeling of the target's motion and provide a basis for subsequent trajectory prediction.
[0032] Real-time intrinsic and extrinsic parameters refer to the internal and external parameters of the camera in its current operating state. These parameters can be dynamically acquired or updated according to actual conditions. This can be achieved using online calibration algorithms or vision-based self-calibration techniques, such as identifying known feature points in the scene or using deep learning models for parameter estimation. The main purpose is to ensure accurate coordinate transformation to adapt to environmental changes or camera attitude adjustments. Each camera has a preset default orientation, i.e., an initial orientation. The intrinsic and extrinsic parameters under this initial orientation are pre-calibrated (e.g., using a checkerboard pattern). When the system starts, each camera's orientation is this initial orientation. When the target vehicle enters the field of view of the first camera, the first camera's orientation remains the initial orientation. Therefore, its real-time intrinsic and extrinsic parameters are equal to the pre-calibrated intrinsic and extrinsic parameters.
[0033] Among them, world coordinates refer to the position representation of the target in a unified three-dimensional spatial coordinate system. It can be implemented using a Cartesian coordinate system or a geodetic coordinate system, such as the X, Y, and Z coordinates with a fixed point as the origin. Its main purpose is to achieve a unified description of the target position and trajectory fusion under the field of view of different cameras.
[0034] The target vehicle trajectory segment refers to the sequence of world coordinates of the target vehicle at different points in time within the field of view of a single camera. It can be generated using linear interpolation or curve fitting methods, such as connecting a series of spatiotemporal coordinate points. Its main purpose is to record and analyze the local motion patterns of the target, providing data support for subsequent motion prediction.
[0035] The field of view boundary region refers to a specific range at the edge of the camera's field of view. It can be defined by an image pixel distance threshold or a geometric projection area, such as a fixed pixel width region at the edge of the image. Its main purpose is to provide early warning of a target that is about to leave the current field of view, triggering subsequent prediction and switching mechanisms.
[0036] Among them, the predicted target position refers to the target's world coordinates at a future point in time inferred from the target's historical movement trajectory. It can be achieved using Kalman filtering or deep learning prediction models, such as calculating the future position based on the target's current speed and direction. Its main purpose is to achieve advance prediction of the target's movement and provide a forward-looking basis for camera switching.
[0037] The preset buffer time refers to the time interval from when the target enters the boundary area of the field of view to when the target position is predicted. It can be implemented by using a fixed time value or by dynamically adjusting it according to the target speed. Its main purpose is to reserve sufficient time for camera adjustment and switching to ensure the continuity of tracking.
[0038] Among them, the neighboring camera refers to the camera that is near the predicted target location and can take over the tracking of the target. It can be determined based on geographical distance or field of view overlap, such as the camera that is closest to the predicted location and whose field of view can cover it. Its main purpose is to achieve a smooth handover of the tracking task.
[0039] The core innovation of this application lies in introducing timestamped frame images and real-time updated camera parameters, combined with a dynamic target position prediction mechanism based on motion trajectory segments, and on this basis, realizing prediction-driven intelligent camera switching and adaptive parameter updates, thereby solving the problems of missing dynamic target prediction and poor environmental adaptability in the prior art, and achieving the effect of continuous and reliable tracking of dynamic targets.
[0040] By employing the above-described scheme, this application effectively addresses the problems of missing dynamic target prediction and poor environmental adaptability in existing multi-camera-based dynamic target tracking methods. Specifically, by introducing timestamped frame images and generating motion trajectory segments containing time information, this application can model the motion of the target vehicle in the time dimension, thereby achieving accurate prediction of the target's future position and avoiding tracking interruptions and target loss due to the inability to predict the target's movement. Simultaneously, by using real-time updated camera intrinsic and extrinsic parameters, combined with prediction-driven camera orientation adjustment and parameter adaptive update mechanisms, this application significantly enhances the system's adaptability in complex and changing environments, ensuring the accuracy of coordinate transformation and the continuity of tracking, effectively avoiding trajectory breakage problems caused by inaccurate fixed parameters. Overall, this application achieves continuous and reliable tracking of dynamic targets, improving the reliability and practicality of the tracking system.
[0041] The identification features may also include license plate number, vehicle model, and vehicle color.
[0042] In some implementations, step A1 includes:
[0043] A101. Real-time acquisition of timestamped frame images captured by the current camera;
[0044] A102. Use the YOLOv8 algorithm to perform target detection on the frame image to obtain the bounding box of the target vehicle, which is used to determine the pixel position coordinates of the target vehicle;
[0045] A103. Using optical character recognition technology to identify the image within the identification frame to obtain the license plate number of the target vehicle;
[0046] A104. Use a deep learning classifier to classify the image within the bounding box to obtain the vehicle model and color of the target vehicle.
[0047] The YOLOv8 algorithm is a real-time object detection algorithm based on deep learning. It can be implemented using a single-stage detection framework, achieving efficient and accurate object localization by directly predicting the bounding boxes (i.e., label boxes) and categories of objects in an image. Specifically, the pixel coordinates of the center point of the label box can be extracted as the pixel position coordinates of the target vehicle.
[0048] Optical character recognition (OCR) is a technology that converts text characters in an image into machine-editable text. It can be implemented using end-to-end models based on convolutional neural networks and recurrent neural networks, or by combining traditional image processing with machine learning, such as through steps like image preprocessing, character segmentation, feature extraction, and classifier recognition.
[0049] Deep learning classifiers are image classification models based on deep neural networks. They can be implemented using convolutional neural network architectures, such as ResNet, VGG, or EfficientNet, and classify objects in images by learning high-level semantic features of the images.
[0050] By introducing license plate number, vehicle model, and vehicle color as recognition features, and employing the YOLOv8 algorithm for target detection, optical character recognition technology to obtain the license plate number, and a deep learning classifier to obtain the vehicle model and color, this solution enhances the accuracy and stability of target vehicle recognition in multi-camera dynamic target tracking systems. This solves the problem of tracking interruption or misidentification that may occur in complex scenarios such as multi-camera switching, the presence of multiple similar vehicles, or temporary occlusion of the target vehicle when relying solely on pixel position coordinates for recognition. This solution provides more reliable identity verification information, enabling the system to stably and continuously track target vehicles.
[0051] Specifically, intrinsic parameters include intrinsic parameter matrices, and extrinsic parameters include translation and rotation matrices.
[0052] The intrinsic parameter matrix is a mathematical model describing the internal optical characteristics and imaging geometry of a camera. It maps two-dimensional pixel coordinates on the image plane to three-dimensional spatial coordinates in the camera coordinate system. The intrinsic parameter matrix typically includes parameters such as focal length, principal point coordinates, and distortion coefficients. The extrinsic parameters, namely the translation and rotation matrices, describe the camera's position and orientation in the world coordinate system, transforming the three-dimensional coordinates in the camera coordinate system to a unified world coordinate system. The translation matrix represents the displacement of the camera's optical center relative to the origin of the world coordinate system, while the rotation matrix represents the rotation of the camera coordinate system relative to the world coordinate system. By specifying the concrete forms of these parameters, a clear and standardized mathematical model can be provided for subsequent precise coordinate transformations, ensuring the accuracy and operability of the transformation.
[0053] In some implementations, step A2 includes:
[0054] A201. Based on the real-time intrinsic parameters of the current camera, convert the pixel position coordinates into camera coordinates in the current camera coordinate system;
[0055] A202. Based on the real-time extrinsic parameters of the current camera, convert the camera coordinates into three-dimensional world coordinates in the world coordinate system;
[0056] A203. Add the timestamp of the current frame image to the three-dimensional world coordinates to form four-dimensional spatiotemporal coordinates, which will serve as the final world coordinates;
[0057] A204. Based on the final world coordinates, generate the target vehicle motion trajectory segment with time information corresponding to the current camera.
[0058] The coordinate transformation methods in steps A201 and A202 are existing technologies and will not be described in detail here.
[0059] This process involves adding the timestamp of the current frame image to the three-dimensional world coordinates, forming four-dimensional spatiotemporal coordinates, which serve as the final world coordinates. This process expands static three-dimensional spatial coordinates into four-dimensional spatiotemporal coordinates that include temporal information. By binding temporal information with spatial location, each target location point carries the time information of its occurrence, thereby enabling more accurate capture of the target's motion state and providing crucial temporal dimension data for subsequent dynamic target motion analysis and prediction.
[0060] In this process, the four-dimensional spatiotemporal coordinates can be connected in chronological order to form a trajectory segment containing time information, which is the target vehicle's motion trajectory segment.
[0061] By clearly defining the specific composition of the camera's intrinsic and extrinsic parameters and innovatively integrating the timestamp of the current frame image into the three-dimensional world coordinates to form a four-dimensional spatiotemporal coordinate system, this method overcomes the limitations of traditional three-dimensional spatial coordinate systems in dynamic target modeling. This four-dimensional spatiotemporal coordinate system with time information precisely binds each target location point to the time of its occurrence, thus accurately reflecting the dynamic motion state of the target vehicle and effectively capturing its motion patterns even in high-speed moving scenes. Consequently, subsequent trajectory prediction can be based on more comprehensive and accurate spatiotemporal data, significantly improving prediction accuracy. This directly ensures the continuity and accuracy of cross-camera tracking, effectively avoiding target loss due to inaccurate predictions and improving the robustness of the entire tracking system.
[0062] In some implementations, step A3 includes:
[0063] A301. Based on the current pixel position coordinates of the target vehicle and the pixel size of the frame image, determine whether the target vehicle has entered the current camera's field of view boundary area;
[0064] A302. If the target vehicle enters the field of view boundary area of the current camera, calculate the current moving speed of the target vehicle based on the target vehicle's movement trajectory segment corresponding to the current camera.
[0065] A303. Calculate the predicted target position based on the preset buffer time and the current moving speed.
[0066] The view boundary region can be defined using a preset pixel distance threshold or a percentage of the image size. For example, it could be a certain pixel range inward from the image edge or a certain percentage range of the image width / height. The pixel size of a frame image refers to the width and height of the current frame image, in pixels.
[0067] The current moving speed refers to the instantaneous speed of the target vehicle at the current moment or its average speed over a recent period. It can be represented using a velocity vector or velocity components, for example, as... , These represent the velocity components along the x, y, and z axes, respectively, with T being the transpose. Linear regression analysis or the finite difference method can be used to calculate the velocity components of the target vehicle along the x, y, and z world coordinate axes, thus obtaining its current velocity vector.
[0068] The preset buffer time is actually a period of time reserved for camera switching and adjustment operations. It can be set as a fixed time value, a time value that is dynamically adjusted according to the system load, or a time value that is adaptively adjusted according to the target vehicle speed.
[0069] The predicted target location can be calculated using the following formula:
[0070] ;
[0071] in, , , These are the x, y, and z axis coordinates of the predicted target location, respectively. , , These are the x, y, and z coordinates of the target vehicle's current position. This is the preset buffer time.
[0072] This solution effectively addresses the insufficient prediction accuracy issue in existing technologies by refining the target location prediction process. By determining whether the target vehicle has entered the field of view boundary region based on its current pixel position coordinates and the pixel size of the frame image, the system can trigger the prediction mechanism in real time and accurately, avoiding prediction lag. Furthermore, by calculating the target vehicle's current speed using its trajectory segment, the prediction is no longer a simple static estimation but fully considers the actual dynamic motion state of the target vehicle, significantly improving prediction accuracy. Finally, by combining the preset buffer time and the accurately calculated current speed, the predicted target position is determined, providing a reliable and forward-looking basis for subsequent camera switching and adjustments. This enables the system to switch cameras more promptly and smoothly, ensuring continuous and stable target tracking even when the target vehicle is moving at high speed or in a complex motion state, effectively preventing target loss.
[0073] Preferably, step A301 may include:
[0074] The feature distance from the edge of the target vehicle in the current camera's field of view is obtained using the following formula:
[0075] L=min(u',1-u',v',1-v');
[0076] Where L is the feature distance from the edge (representing the closest distance between the target vehicle and the current camera's field of view boundary), u' is the normalized value of the horizontal coordinate of the target vehicle's current pixel position coordinates, and v' is the normalized value of the vertical coordinate of the target vehicle's current pixel position coordinates.
[0077] If the feature distance from the edge is less than or equal to the preset distance threshold, the target vehicle is determined to have entered the current camera's field of view boundary area; otherwise, the target vehicle is determined not to have entered the current camera's field of view boundary area.
[0078] The normalized value refers to mapping the original pixel coordinates to a floating-point value between 0 and 1, where 0 represents the left or top boundary of the image, and 1 represents the right or bottom boundary. This normalization process eliminates the influence of different camera resolutions, making subsequent judgments more universal.
[0079] The preset edge distance threshold is a critical value used to define whether a target vehicle has entered the boundary area of the field of view. It can be set through experiments or experience according to the needs of the actual application scenario to balance the timeliness and stability of tracking. For example, it can be 0.1, but it is not limited to this.
[0080] This solution introduces the concept of feature offset and uses the normalized value of the target vehicle's pixel position coordinates for calculation, providing a precise, standardized, and robust method for determining whether a target vehicle has entered the current camera's field of view boundary. This method overcomes the limitations of relying solely on raw pixel coordinates and frame image pixel dimensions, accurately identifying when the target vehicle is truly at the edge of the camera's effective tracking range. Therefore, this solution optimizes the triggering timing of target prediction and camera switching, ensuring that subsequent target position prediction steps are triggered consistently and reliably across different camera configurations and scenarios. This guarantees the continuity and smoothness of the multi-camera tracking system when the target vehicle crosses the camera's field of view.
[0081] Specifically, step A4 may include:
[0082] If the predicted target location is outside the optical center scanning range of all other cameras besides the current camera, then the predicted target location is determined to be outside the field of view of the multi-camera tracking system.
[0083] The optical center scanning range refers to the range of positions that the optical center can reach within the camera's orientation adjustment range. When the predicted target position is outside the optical center scanning range of all other cameras except the current camera, it means that the optical centers of other cameras cannot be aligned with the predicted target position, and therefore no camera will take over. At this time, step A6 is executed to fuse all target vehicle motion trajectory segments to obtain the total motion trajectory.
[0084] In some implementations, step A5 includes:
[0085] A501. Calculate the distance between the predicted target location and the field of view of all other cameras except the current camera, and denot it as the deviation distance;
[0086] A502. Select another camera whose deviation distance is less than a preset distance threshold as a neighboring camera;
[0087] A503. Based on the predicted target position, the current rotation angle of the nearby camera, the current pitch angle of the nearby camera, and the real-time extrinsic parameters of the nearby camera, calculate the amount of rotation angle adjustment and pitch angle adjustment required to align the optical center of the nearby camera with the predicted target position.
[0088] A504. Adjust the rotation and tilt angles of adjacent cameras based on the rotation and tilt angle adjustments.
[0089] A505. Acquire three consecutive frames of images from a nearby camera after adjusting the rotation and pitch angles. Input the three consecutive frames of images into a pre-trained deep learning calibration model to obtain the external parameter correction amount output by the deep learning calibration model.
[0090] A506. Update the extrinsic parameters of nearby cameras based on the extrinsic parameter correction amount;
[0091] A507. Calculate the zoom factor based on the pixel size of the identification frame obtained by the nearby camera and the actual size of the target vehicle, in order to adjust the focal length of the nearby camera;
[0092] A508. Update the intrinsic parameters of nearby cameras based on the adjusted focal length;
[0093] A509. After completing the adjustment of the intrinsic and extrinsic parameters of the neighboring camera, use the neighboring camera as the new current camera and return to step A1.
[0094] In step A501, the field of view of other cameras refers to the three-dimensional spatial area that each other camera can cover in its current pose (or orientation). Specifically, the field of view boundary can be defined and calculated using the intrinsic and extrinsic parameters of the cameras, as well as their field of view angle parameters. For example, a frustum model of each camera can be constructed to determine the position of the frustum boundary. The deviation distance can be calculated using various geometric distance metrics. For example, the Euclidean distance from the predicted target position to the center point of the field of view of each other camera can be calculated, or the shortest distance from the predicted target position to the field of view boundary can be calculated.
[0095] In step A502, the preset distance threshold is a configurable parameter used to limit the selection range of neighboring cameras. Specifically, it can be set according to the actual application scenario, camera deployment density, and tracking accuracy requirements. Various strategies can be employed to select neighboring cameras. For example, the camera with the smallest deviation distance can be selected, or the camera with the best image quality and lowest occlusion risk within the threshold range can be selected, or the camera with the smallest angle adjustment (e.g., the sum of rotation and pitch angle adjustments) required to align the optical center with the predicted target position can be selected as the neighboring camera.
[0096] In step A503, the current rotation and pitch angles of the nearby camera can be obtained in real time through the camera's own sensors (such as an encoder). The adjustment amounts of the rotation and pitch angles can be calculated using inverse kinematics algorithms. For example, the world coordinates of the predicted target position are converted into coordinates in the camera coordinate system of the nearby camera. Then, based on the vector relationship between the optical center of the nearby camera and the target position, combined with the current attitude (rotation and pitch angles) of the nearby camera, the rotation and translation transformations required to align the optical axis with the target are calculated, and then decomposed into the adjustment amounts of the rotation and pitch angles.
[0097] In step A504, adjusting the rotation and tilt angles of the adjacent camera can be achieved by controlling the stepper motor or servo motor inside the adjacent camera. For example, control commands can be sent to the PTZ (Pan-Tilt-Zoom) unit to make it rotate precisely according to the calculated adjustment amount. The adjustment process can use closed-loop control, ensuring the accuracy of the adjustment by feeding back the actual rotation angle of the camera.
[0098] In step A505, acquiring three consecutive frames of images is to provide temporal information, helping the deep learning calibration model better understand scene changes and motion information. The pre-trained deep learning calibration model can be a convolutional neural network (CNN) combined with a recurrent neural network (RNN) or a Transformer model. Training data can include images under different lighting, viewpoint, and occlusion conditions, along with their corresponding true extrinsic parameter deviations. Extrinsic parameter corrections can be small adjustments to rotation and translation matrices, such as rotation corrections in Euler angles or Lie algebra form, and 3D translation corrections. For example, the extrinsic parameter corrections output by the deep learning calibration model can be represented as an array: ΔP=[Δrx,Δry,Δrz,Δtx,Δty,Δtz], where ΔP represents the extrinsic parameter correction output by the deep learning calibration model, Δrx, Δry, and Δrz are rotation corrections along three axes, and Δtx, Δty, and Δtz are translation corrections along three axes.
[0099] In step A506, updating the extrinsic parameters of neighboring cameras can be achieved by superimposing the correction amount obtained in A505 onto the current real-time extrinsic parameters of the cameras. For example, if the extrinsic parameter correction amount is the increment of the rotation matrix and translation matrix, it can be updated using matrix multiplication and vector addition. Specifically, it can be updated using the following formula:
[0100] ;
[0101] ;
[0102] ;
[0103] ;
[0104] in, For the updated rotation matrix, The rotation matrix of the nearest camera. The updated translation matrix, This is the current translation matrix of the neighboring cameras. It is a vector consisting of rotational corrections along three axes. It is a vector consisting of translation corrections along three axes. Indicates according to The generated rotation matrix.
[0105] In step A507, the pixel size of the identification box obtained by the neighboring camera in identifying the target vehicle can be obtained through a target detection algorithm (such as YOLOv8). The actual size of the target vehicle can be obtained from a preset database based on its vehicle type information; for example, different vehicle types such as sedans, SUVs, and trucks have different average lengths, widths, and heights. The zoom factor can be calculated based on the principle of perspective projection. For example, by comparing the ratio of the pixel size of the target vehicle in the image to its actual size, combined with the current focal length, the required focal length adjustment can be calculated so that the target appears at the desired size in the image.
[0106] In step A508, updating the intrinsic parameters of the neighboring cameras refers to updating the focal length parameter in the intrinsic parameter matrix. The intrinsic parameter matrix typically includes the focal length parameter, principal point coordinates, and distortion coefficients. When the focal length is adjusted, the corresponding focal length parameter in the intrinsic parameter matrix is updated.
[0107] In step A509, using the neighboring camera as the new current camera means switching the dominant camera for the current tracking task to this neighboring camera within the system, and using its latest intrinsic and extrinsic parameters as real-time parameters. Returning to step A1 means that the system will start from the new current camera and continue to acquire images, identify targets, transform coordinates, and generate trajectory segments in real time, thereby achieving seamless cross-camera tracking.
[0108] This solution effectively addresses the problem of inaccurate updates to intrinsic and extrinsic parameters in multi-camera dynamic target tracking after camera orientation adjustments by introducing a series of refined steps. By calculating the deviation distance between the predicted target position and the field of view of other cameras and selecting neighboring cameras, it ensures that the most suitable camera for takeover tracking is chosen during switching, avoiding unnecessary switching or improper selection. By accurately calculating and adjusting the rotation and pitch angles of neighboring cameras based on the predicted target position, it ensures that the cameras can accurately align with the future position of the target vehicle, improving the efficiency and accuracy of target acquisition. Crucially, by acquiring continuous images after orientation adjustments and inputting them into a pre-trained deep learning calibration model to obtain extrinsic parameter corrections, and updating the extrinsic parameters accordingly, this solution adaptively and robustly corrects extrinsic parameter deviations caused by changes in camera orientation or environmental factors, significantly improving the accuracy of extrinsic parameters and thus ensuring the accuracy of pixel position coordinate to world coordinate transformation. Furthermore, by calculating the zoom factor and adjusting the focal length based on the pixel size and actual size of the target vehicle's identification frame, and then updating the intrinsic parameters, this solution can dynamically optimize image quality, ensuring that the target vehicle is presented at its optimal size in the new camera's field of view, thus improving the clarity and accuracy of target recognition. These improvements work together to enable the system to achieve seamless, high-precision tracking of the target vehicle across different camera fields of view, significantly reducing the risk of target loss and improving the continuity and reliability of overall tracking.
[0109] In some possible implementations, the deep learning calibration model includes:
[0110] The ResNet-34 network is used to extract image features from three consecutive frames of images, resulting in three image feature vectors.
[0111] The ConvLSTM network is used to perform temporal modeling on three image feature vectors to obtain temporal features.
[0112] Fully connected networks are used to perform parametric regression processing on time series features to obtain and output extrinsic parameter corrections.
[0113] Among them, ResNet-34 is a deep convolutional neural network architecture based on residual connections. It can be implemented using stacked layers containing multiple residual blocks. Each residual block contains a convolutional layer, a batch normalization layer, and an activation function layer. Skip connections are used to directly add the input to the output to alleviate the gradient vanishing problem in deep network training, thereby effectively extracting multi-level spatial features from images. ConvLSTM is a recurrent neural network that combines convolutional operations and long short-term memory network characteristics. It can be implemented using a sequence of units containing convolutional gating units (such as input gates, forget gates, and output gates) and convolutional state update mechanisms. Each unit uses convolutional operations to process the input and state, thereby capturing spatial and temporal dependencies when processing sequential data. Fully connected networks refer to a neural network structure composed of multiple fully connected layers (or dense layers) stacked together. It can be implemented using a multilayer perceptron structure containing an input layer, one or more hidden layers, and an output layer. The neurons in each layer are connected to all neurons in the previous layer. Through weights and biases, linear transformations and non-linear activations are performed on the input data to achieve complex mapping and regression of input features.
[0114] The deep learning calibration model works as follows: After a neighboring camera adjusts its orientation, three consecutive frames of images are acquired. These images are then fed into a ResNet-34 network. The ResNet-34 network, acting as an image feature extractor, performs deep convolution on each frame, capturing rich spatial details and semantic information, and generating an image feature vector for each frame. Since there are three consecutive frames, three independent image feature vectors are obtained, each representing the spatial content of the image at a different time point. These three image feature vectors are then fed into a ConvLSTM network. The ConvLSTM network plays a role in temporal modeling here. It not only processes the spatial information of each image feature vector, but more importantly, it captures the dynamic correlation and changing trends between these three consecutive image feature vectors in the temporal dimension. Through its internal convolutional gating mechanism, the ConvLSTM network effectively learns and memorizes the temporal dependencies in the image sequence, thus integrating the three independent image feature vectors into a comprehensive temporal feature containing both spatial and temporal information. This temporal modeling is crucial for understanding subtle changes in camera pose, as the misalignment of extrinsic parameters is often a gradual or dynamic process. Finally, the temporal features output by the ConvLSTM network are input into a fully connected network. The fully connected network is responsible for parametric regression processing of these high-dimensional temporal features. Through multiple layers of nonlinear transformations, it maps the complex temporal features to low-dimensional, specific numerical values—the extrinsic parameter corrections. These corrections accurately reflect the deviation between the current extrinsic parameters of the camera and the actual extrinsic parameters. In this way, the deep learning calibration model can accurately calculate the required adjustments to the rotation angle and translation matrix based on the temporal changes of three consecutive frames.
[0115] This deep learning calibration model is tightly integrated with the steps in the aforementioned method, forming a complete calibration closed loop. After acquiring three consecutive frames of images, the model receives these images as input and outputs extrinsic parameter corrections. These corrections are then directly used to update the extrinsic parameters of neighboring cameras. This integration enables the system to perform real-time, dynamic, and refined extrinsic parameter calibration, overcoming the limitations of traditional fixed parameters in complex environments. By fully utilizing the spatial and temporal features of the images, the model can more accurately perceive subtle changes in camera pose, thereby providing more precise extrinsic parameter corrections. This ensures the accuracy and stability of multi-camera dynamic target tracking, especially in scenarios with frequent camera switching and environmental changes, significantly improving the system's environmental adaptability and robustness.
[0116] Preferably, step A507 may include:
[0117] B1. Acquire frame images captured by nearby cameras, perform vehicle identification from them, and obtain identification features of at least one vehicle; the identification features also include license plate number, vehicle model and vehicle color;
[0118] B2. Based on the identification features of the target vehicle obtained by the current camera and the license plate number, vehicle type and vehicle color of each vehicle obtained by the neighboring cameras, calculate the similarity between each vehicle identified by the neighboring cameras and the target vehicle.
[0119] B3. Determine the target vehicle from among the vehicles identified by nearby cameras based on similarity;
[0120] B4. Obtain the pixel size of the identified target vehicle's frame, denoted as the first pixel size, and obtain the corresponding actual size based on the vehicle model of the identified target vehicle;
[0121] B5. Calculate the ratio of the actual size to the size of the first pixel to obtain the actual conversion ratio. Combine the preset target conversion ratio and the current focal length of the adjacent camera to calculate the zoom factor that makes the actual conversion ratio equal to the target conversion ratio.
[0122] B6. Adjust the focal length of the adjacent camera according to the zoom level.
[0123] Step B1 can be referred to step A1 above, and will not be repeated here.
[0124] Similarity refers to a quantitative indicator that measures the degree of closeness between two or more entities in a specific or multiple dimensions. It can be calculated using methods such as feature distance-based measures (e.g., Euclidean distance, Manhattan distance), feature matching-based scores (e.g., hash matching, string edit distance), or classification probabilities based on machine learning models.
[0125] In step B3, a threshold-based method can be used to determine the target vehicle. For example, a preset similarity threshold (e.g., 95%) can be set, and vehicles with a similarity higher than this threshold can be considered candidate target vehicles. If only one vehicle has a similarity higher than the threshold, it is identified as the target vehicle. If multiple vehicles have a similarity higher than the threshold, the vehicle with the highest similarity can be selected as the target vehicle. Furthermore, if multiple vehicles have a similarity higher than the threshold, vehicle trajectory information can be used for auxiliary judgment; for example, the vehicle whose direction and speed more closely match the predicted trajectory of the target vehicle can be selected.
[0126] In step B4, the pixel dimensions of the bounding box can be directly obtained from the output of the object detection algorithm, such as the width and height of the bounding box. The actual dimensions of the vehicle can be obtained by querying a pre-established vehicle model database. This database can contain the length, width, and height information of various common vehicle models. Once the model of the target vehicle is determined, the system can automatically retrieve the corresponding actual dimension data from the database.
[0127] In step B5, the actual conversion ratio can be defined as the ratio of the vehicle's actual physical size (e.g., length or width) to the corresponding pixel size in the image. The preset target conversion ratio is the desired size ratio of the target vehicle in the new camera field of view; for example, it can be preset to make the target vehicle occupy a certain percentage of the image height. The zoom factor can be calculated based on optical imaging principles, such as the relationship between focal length, object distance, and image size. By comparing the actually observed size ratio with the desired target conversion ratio and combining it with the camera's current focal length, the required zoom factor is accurately calculated. This ensures that the focal length adjustment allows the target vehicle to achieve the preset ideal size and sharpness in the new camera field of view, optimizing the tracking effect.
[0128] This solution introduces a multi-feature similarity matching mechanism to accurately identify target vehicles within the fields of view of adjacent cameras, effectively solving the problem of difficulty in accurately determining the tracking target in multi-target scenarios or when there are subtle differences in target features. This ensures the accuracy of the pixel size of the marker box used to calculate zoom magnification and the actual size of the target vehicle, thus avoiding size information deviations caused by target recognition errors. Therefore, this solution enables precise adjustment of the focal length of adjacent cameras, allowing the target vehicle to quickly obtain clear and stable imaging after switching to a new camera's field of view. This significantly improves the smoothness of target vehicle switching and the robustness of tracking during cross-camera tracking, ensuring continuous and clear tracking of the target vehicle.
[0129] Preferably, step B2 may include:
[0130] B201. Calculate the license plate number consistency parameter according to the following formula:
[0131] ;
[0132] in, For license plate number consistency parameters, This is the encoding of the license plate number of the target vehicle currently captured by the camera. This is the encoding of the license plate number of a vehicle captured by a nearby camera. for and The coding distance between them for The encoding length, for The encoding length; where encoding distance is a measure of the difference between two encoded sequences, and is a dimensionless value; encoding length refers to the number of characters or bytes in the encoded sequence, and is a dimensionless value;
[0133] B202. Convert the color spaces of the target vehicle's color captured by the current camera and the vehicle's color captured by neighboring cameras to the HSV color space, and calculate the color distance using the following formula:
[0134] ;
[0135] in, For color distance, , , These are the H, S, and V channel values of the vehicle color of the target vehicle currently captured by the camera. , , These are the H, S, and V channel values of the vehicle color obtained from a nearby camera. This is a preset reference threshold;
[0136] B203. Using a hierarchical scoring rule based on a vehicle classification tree, determine the category similarity score Ts between the vehicle type of the target vehicle captured by the current camera and the vehicle type of the vehicles captured by neighboring cameras;
[0137] B204. Calculate the similarity between the vehicle identified by the nearby camera and the target vehicle using the following formula:
[0138] ;
[0139] in, For similarity, , , These are the preset weighting coefficients.
[0140] The license plate number consistency parameter quantifies the degree of similarity between two license plate numbers. Encoding refers to converting the character sequence of the license plate number into a computable form, which can be implemented using ASCII encoding, Unicode encoding, or a custom numeric encoding. Encoding distance is a measure of the difference between two encoded sequences, which can be calculated using algorithms such as Levenshtein distance, Hamming distance, or Jaccard distance. Encoding length refers to the number of characters or bytes in the encoded sequence, which can be obtained using string length functions or byte counting functions.
[0141] The HSV color space is a color model that represents colors as hue, saturation, and value, and can be implemented using standard color space conversion algorithms. Color distance is a numerical value that quantifies the difference between two colors in the HSV color space. A preset reference threshold is a pre-defined boundary value used to judge color similarity in color distance calculations; it can be determined based on empirical values, statistical analysis, or machine learning methods.
[0142] The vehicle model classification tree refers to a hierarchical classification structure that organizes vehicle models according to hierarchical relationships. It can be constructed using methods such as decision trees, hierarchical clustering, or ontology. Hierarchical scoring rules refer to a set of rules that assign scores based on the similarity of different levels and branches in the vehicle model classification tree. These rules can be designed using methods such as common ancestor nodes, path length, or semantic similarity. Category similarity scores are numerical values representing the degree of similarity between two vehicle model categories calculated according to the hierarchical scoring rules. These scores can be represented using methods such as normalized scoring, percentage matching, or fuzzy logic. For example, hierarchical scoring rules may include:
[0143] If the vehicle types are the same (e.g., both are sedans), then the category similarity score Ts is 1;
[0144] If the vehicle types belong to the same major category (e.g., SUVs and MPVs), then the category similarity score Ts is 0.7;
[0145] If the vehicle types do not belong to the same major category (e.g., cars and trucks), the category similarity score Ts is 0.1.
[0146] The preset weight coefficients refer to the values used to balance the importance of different features in the comprehensive similarity calculation. They can be determined by methods such as expert experience, regression analysis, or optimization algorithms.
[0147] Through the above technical solutions, this method can effectively address issues such as license plate recognition errors, color deviations caused by changes in lighting, and hierarchical relationships in vehicle classification that exist in practical applications. Specifically, the calculation of license plate number consistency parameters can handle subtle differences in license plate recognition, avoiding misjudgments due to minor errors; color distance calculation in the HSV color space can effectively reduce the impact of environmental factors such as lighting on color recognition, improving the accuracy and stability of color judgment; and the hierarchical scoring rules based on the vehicle classification tree can more finely and reasonably evaluate vehicle similarity, avoiding the limitations of simple matching. These multi-dimensional and refined similarity calculation methods, through weighted fusion, significantly improve the recognition accuracy and robustness of target vehicles when switching between cameras, thereby effectively solving the problem of misidentification or target loss caused by single or coarse similarity calculation methods, and thus improving the accuracy and stability of the entire tracking system.
[0148] In some implementations, step A6 includes:
[0149] A601. Based on the four-dimensional spatiotemporal coordinates of the trajectory points of each target vehicle's trajectory segment, the nearest neighbor matching algorithm is used to fuse all target vehicle trajectory segments to obtain a preliminary total trajectory.
[0150] A602. Using the Kalman filter method, the initial total motion trajectory is smoothed to obtain the final total motion trajectory.
[0151] The nearest neighbor matching algorithm is an algorithm used to find the data point in a dataset that is closest to a given data point. It can use Euclidean distance, Mahalanobis distance or other distance metrics to calculate the similarity between trajectory points and select the point with the smallest distance as the matching object.
[0152] Kalman filtering is an algorithm for estimating the state of dynamic systems, particularly suitable for processing noisy measurement data. It can use a prediction-update loop to combine the system's motion model and measurement data to estimate and correct the system state in real time, thereby achieving data smoothing and noise filtering.
[0153] This solution addresses the continuity and accuracy issues of trajectory fusion in multi-camera tracking through two main collaborative steps. First, in the initial fusion stage, based on the four-dimensional spatiotemporal coordinates of the trajectory points of each target vehicle's trajectory segment, a nearest neighbor matching algorithm is used to fuse all target vehicle trajectory segments, resulting in a preliminary overall trajectory. Here, the four-dimensional spatiotemporal coordinates used for the trajectory points not only include the target's spatial position information but also temporal information. This introduction of the temporal dimension enables accurate identification and association of trajectory points of the same target vehicle captured by different cameras at different time points during trajectory segment matching. The nearest neighbor matching algorithm, based on these time-informed coordinates, can handle overlaps or gaps between different trajectory segments, initially stitching together scattered trajectory points into a continuous trajectory. This spatiotemporal coordinate-based matching solves the problem of connecting target trajectories in cross-camera tracking, laying the foundation for subsequent smoothing processing.
[0154] Building upon this foundation, in the smoothing stage, the Kalman filter method is used to smooth the preliminary total motion trajectory, yielding the final total motion trajectory. The initially fused trajectory may still contain jitter and unevenness due to measurement noise, environmental interference, or matching errors. As an algorithm, the Kalman filter can dynamically adjust the estimated values of trajectory points based on the target's motion model and measurement data, thereby filtering out random errors in the trajectory and ensuring that the final total motion trajectory is smooth, continuous, and conforms to the actual motion patterns of the target. It is precisely because this approach first utilizes four-dimensional spatiotemporal coordinates with time information and a nearest neighbor matching algorithm for preliminary trajectory fusion, and then smooths the fused trajectory using Kalman filtering, that this solution can jointly address the continuity and accuracy issues of target vehicle trajectory fusion in multi-camera environments, thus providing a total motion trajectory with good quality and reliability.
[0155] In a specific embodiment, the step of fusing all target vehicle motion trajectory segments to obtain the total motion trajectory can be implemented as follows: First, in the initial total motion trajectory generation stage, the system receives motion trajectory segments of each target vehicle captured by different cameras. Each trajectory segment consists of a series of trajectory points, each containing four-dimensional spatiotemporal coordinates, i.e., a three-dimensional spatial position (X, Y, Z) and a corresponding timestamp (T). To fuse these trajectory segments, a global set of trajectory points can be constructed. Subsequently, a nearest neighbor matching algorithm is used; for example, each trajectory point in a trajectory segment can be traversed, and the trajectory point closest in four-dimensional spatiotemporal distance can be found in other trajectory segments. The distance can be calculated using Euclidean distance, combined with a weight in the time dimension. When a matching point is found, a preset matching threshold can be used to determine whether to connect or merge the points. For example, if the spatial and temporal distance between two trajectory points is less than a certain threshold, they are considered to belong to the continuous motion of the same target, and the corresponding trajectory segments are spliced together. This process can be iterated until all relevant trajectory segments have been attempted to be fused, resulting in a preliminary total motion trajectory that may still have slight jitter.
[0156] Following this, a Kalman filter can be applied to smooth the initial total motion trajectory. Specifically, a state model can be defined for the target vehicle's motion; for example, the state vector could contain the vehicle's position and velocity in the X, Y, and Z directions. The measured values are the four-dimensional spatiotemporal coordinates of the initial total motion trajectory. The Kalman filter iterates through two phases: prediction and update. In the prediction phase, the state at the next moment is predicted based on the vehicle's motion model; in the update phase, the predicted state is corrected using the actual measured trajectory point data. Through continuous iteration, the Kalman filter can estimate the true motion state of the target vehicle and filter out random noise and errors introduced during the measurement process, resulting in a final total motion trajectory with higher continuity and smoothness, more accurately reflecting the actual motion path of the target vehicle.
[0157] Please refer to Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device includes a processor 301 and a memory 302. The processor 301 and the memory 302 are interconnected and communicate with each other via a communication bus 303 and / or other connection mechanisms (not shown). The memory 302 stores a computer program executable by the processor 301. When the electronic device is running, the processor 301 executes the computer program to perform the dynamic target tracking method based on multiple cameras in any optional implementation of the above embodiments, to achieve the following functions: A1. Real-time acquisition of time-stamped frame images captured by the current camera, identification of the target vehicle from them, and acquisition of the target vehicle's identification features; the current camera is a camera whose field of view covers the target vehicle; the identification features include pixel position coordinates; A2. Based on the real-time intrinsic and extrinsic parameters of the current camera, the pixel position coordinates... A3. Convert to world coordinates to generate the target vehicle motion trajectory segment corresponding to the current camera; where, if the current orientation of the current camera is the initial orientation, the real-time intrinsic and extrinsic parameters are the intrinsic and extrinsic parameters pre-calibrated under the initial orientation; A4. When the target vehicle enters the field of view boundary area of the current camera, predict the world coordinate position of the target vehicle after a preset buffer time according to the target vehicle motion trajectory segment corresponding to the current camera, and record it as the predicted target position; A5. If the predicted target position is out of the field of view of the multi-camera tracking system, proceed to step A6, otherwise, proceed to step A5; A6. Determine a neighboring camera according to the predicted target position, adjust the orientation of the neighboring camera to align with the predicted target position, update the intrinsic and extrinsic parameters of the neighboring camera, and use the neighboring camera as the new current camera, return to step A1; A7. Fuse all target vehicle motion trajectory segments to obtain the total motion trajectory.
[0158] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the dynamic target tracking method based on multiple cameras in any optional implementation of the above embodiments to achieve the following functions: A1. Real-time acquisition of time-stamped frame images captured by the current camera, identifying the target vehicle from them, and obtaining the identification features of the target vehicle; the current camera is a camera whose field of view covers the target vehicle; the identification features include pixel position coordinates; A2. Based on the real-time intrinsic and extrinsic parameters of the current camera, converting the pixel position coordinates into world coordinates to generate the target vehicle motion trajectory segment corresponding to the current camera; wherein, if the current orientation of the current camera is the initial orientation, then real-time... The intrinsic and extrinsic parameters are pre-calibrated under the initial orientation; A3. When the target vehicle enters the field of view boundary area of the current camera, predict the world coordinate position of the target vehicle after a preset buffer time based on the target vehicle motion trajectory segment corresponding to the current camera, and record it as the predicted target position; A4. If the predicted target position is out of the field of view of the multi-camera tracking system, proceed to step A6; otherwise, proceed to step A5; A5. Determine a neighboring camera based on the predicted target position, adjust the orientation of the neighboring camera to align with the predicted target position, update the intrinsic and extrinsic parameters of the neighboring camera, and use the neighboring camera as the new current camera, then return to step A1; A6. Fuse all target vehicle motion trajectory segments to obtain the total motion trajectory.
[0159] The computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0160] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A dynamic target tracking method based on multiple cameras, wherein a vehicle is tracked using a multi-camera tracking system with multiple cameras having adjustable orientations, characterized in that, Including the following steps: A1. Real-time acquisition of timestamped frame images captured by the current camera, identification of the target vehicle from them, and acquisition of the identification features of the target vehicle; the current camera is a camera whose field of view covers the target vehicle; the identification features include pixel position coordinates; A2. Based on the real-time intrinsic and extrinsic parameters of the current camera, the pixel position coordinates are converted into world coordinates to generate the target vehicle motion trajectory segment corresponding to the current camera; wherein, if the current orientation of the current camera is the initial orientation, the real-time intrinsic and extrinsic parameters are the intrinsic and extrinsic parameters pre-calibrated under the initial orientation; A3. When the target vehicle enters the field of view boundary area of the current camera, predict the world coordinate position of the target vehicle after a preset buffer time based on the target vehicle's motion trajectory segment corresponding to the current camera, and record it as the predicted target position; A4. If the predicted target position is outside the field of view of the multi-camera tracking system, proceed to step A6; otherwise, proceed to step A5. A5. Determine a nearby camera based on the predicted target location, adjust the orientation of the nearby camera to align with the predicted target location, update the intrinsic and extrinsic parameters of the nearby camera, and use the nearby camera as the new current camera, then return to step A1; A6. Merge all the target vehicle motion trajectory segments to obtain the total motion trajectory; Step A5 includes: A501. Calculate the distance between the predicted target location and the field of view of all other cameras except the current camera, and record it as the deviation distance; A502. Select one of the other cameras whose deviation distance is less than a preset distance threshold as the neighboring camera; A503. Based on the predicted target position, the current rotation angle of the nearby camera, the current pitch angle of the nearby camera, and the real-time external parameters of the nearby camera, calculate the amount of rotation angle adjustment and pitch angle adjustment required to align the optical center of the nearby camera with the predicted target position; A504. Adjust the rotation angle and pitch angle of the adjacent camera according to the rotation angle adjustment amount and the pitch angle adjustment amount; A505. Acquire three consecutive frames of images captured by the adjacent camera after adjusting the rotation and pitch angles, input the three consecutive frames of images into a pre-trained deep learning calibration model, and obtain the external parameter correction amount output by the deep learning calibration model. A506. Update the extrinsic parameters of the nearby camera according to the extrinsic parameter correction amount; A507. Calculate the zoom factor based on the pixel size of the identification frame obtained by the neighboring camera in identifying the target vehicle and the actual size of the target vehicle, in order to adjust the focal length of the neighboring camera; A508. Update the intrinsic parameters of the neighboring cameras based on the adjusted focal length; A509. After completing the adjustment of the intrinsic and extrinsic parameters of the neighboring camera, use the neighboring camera as the new current camera and return to step A1.
2. The dynamic target tracking method based on multiple cameras according to claim 1, characterized in that, The identification features also include license plate number, vehicle model, and vehicle color; Step A1 includes: A101. Real-time acquisition of timestamped frame images captured by the current camera; A102. Use the YOLOv8 algorithm to perform target detection on the frame image to obtain the bounding box of the target vehicle, which is used to determine the pixel position coordinates of the target vehicle; A103. Using optical character recognition technology to identify the image within the identification frame to obtain the license plate number of the target vehicle; A104. Use a deep learning classifier to classify the image within the identification box to obtain the vehicle model and color of the target vehicle.
3. The dynamic target tracking method based on multiple cameras according to claim 1, characterized in that, The intrinsic parameters include an intrinsic parameter matrix, and the extrinsic parameters include a translation matrix and a rotation matrix; Step A2 includes: A201. Based on the real-time intrinsic parameters of the current camera, convert the pixel position coordinates into camera coordinates in the current camera coordinate system; A202. Based on the real-time extrinsic parameters of the current camera, convert the camera coordinates into three-dimensional world coordinates in the world coordinate system; A203. Add the timestamp of the current frame image to the three-dimensional world coordinates to form four-dimensional spatiotemporal coordinates, which will serve as the final world coordinates; A204. Based on the final world coordinates, generate the target vehicle motion trajectory segment with time information corresponding to the current camera.
4. The dynamic target tracking method based on multiple cameras according to claim 3, characterized in that, Step A3 includes: A301. Based on the current pixel position coordinates of the target vehicle and the pixel size of the frame image, determine whether the target vehicle has entered the field of view boundary area of the current camera; A302. If the target vehicle enters the field of view boundary area of the current camera, the current moving speed of the target vehicle is calculated based on the target vehicle's movement trajectory segment corresponding to the current camera. A303. Calculate the predicted target position based on the preset buffer time and the current moving speed.
5. The dynamic target tracking method based on multiple cameras according to claim 4, characterized in that, Step A301 includes: The feature distance from the edge of the target vehicle in the current camera's field of view is obtained using the following formula: L=min(u',1-u',v',1-v'); Where L is the feature distance from the edge, u' is the normalized value of the horizontal coordinate of the current pixel position coordinate of the target vehicle, and v' is the normalized value of the vertical coordinate of the current pixel position coordinate of the target vehicle. If the feature distance from the edge is less than or equal to a preset distance threshold, the target vehicle is determined to have entered the field of view boundary area of the current camera; otherwise, the target vehicle is determined not to have entered the field of view boundary area of the current camera.
6. The dynamic target tracking method based on multiple cameras according to claim 1, characterized in that, Step A507 includes: B1. Acquire frame images captured by the nearby camera, perform vehicle identification from them, and obtain identification features of at least one vehicle; the identification features also include license plate number, vehicle model, and vehicle color; B2. Based on the identification features of the target vehicle acquired by the current camera and the license plate number, vehicle type, and vehicle color of each vehicle acquired by the neighboring cameras, calculate the similarity between each vehicle identified by the neighboring cameras and the target vehicle. B3. Determine the target vehicle from among the vehicles identified by the nearby cameras based on the similarity score; B4. Obtain the pixel size of the identified target vehicle's frame, denoted as the first pixel size, and obtain the corresponding actual size based on the vehicle model of the identified target vehicle; B5. Calculate the ratio of the actual size to the first pixel size to obtain the actual conversion ratio. Combine the preset target conversion ratio and the current focal length of the adjacent camera to calculate the zoom factor that makes the actual conversion ratio equal to the target conversion ratio. B6. Adjust the focal length of the adjacent camera according to the zoom level.
7. The dynamic target tracking method based on multiple cameras according to claim 6, characterized in that, Step B2 includes: B201. Calculate the license plate number consistency parameter according to the following formula: ; in, The license plate number consistency parameter, This refers to the encoding of the license plate number of the target vehicle currently captured by the camera. The encoding of the vehicle's license plate number acquired by the nearby camera. for and The coding distance between them for The encoding length, for The encoding length; B202. Convert the color spaces of the target vehicle's color captured by the current camera and the vehicle's color captured by the neighboring camera into the HSV color space, and calculate the color distance using the following formula: ; in, The color distance, , , These are the H, S, and V channel values of the vehicle color of the target vehicle acquired by the current camera. , , These are the H, S, and V channel values of the vehicle color acquired by the nearby camera, respectively. This is a preset reference threshold; B203. Using a hierarchical scoring rule based on a vehicle model classification tree, determine the category similarity score Ts between the vehicle model of the target vehicle captured by the current camera and the vehicle model captured by the neighboring cameras; B204. Calculate the similarity between the vehicle identified by the nearby camera and the target vehicle according to the following formula: ; in, For the aforementioned similarity, , , These are the preset weighting coefficients.
8. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program executable by the processor, and when the processor executes the computer program, it performs the steps of the dynamic target tracking method based on multiple cameras as described in any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it performs the steps of the dynamic target tracking method based on multiple cameras as described in any one of claims 1-7.
Citation Information
Patent Citations
Vehicle real-time tracking method and computer readable storage medium
CN112633282A