Dynamic target tracking method based on multiple cameras and related equipment
By introducing timestamp and real-time parameter update mechanism in the multi-camera system and combining it with camera orientation adjustment, the problems of missing dynamic target prediction and poor environmental adaptability are solved, and accurate and robust tracking of dynamic targets is achieved.
Patent Information
- Application Number
- CN202511177669.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing multi-camera-based dynamic target tracking systems have deficiencies in dynamic target prediction and environmental adaptability, leading to target loss and trajectory disruption.
By introducing frame images with timestamps and real-time updated camera parameters, combined with dynamic target position prediction mechanism and camera orientation adjustment, accurate tracking of dynamic targets can be achieved.
The accuracy and robustness of dynamic target tracking are improved, ensuring continuous tracking in complex environments and avoiding target loss and trajectory breakage.
Smart Images

Figure CN120673089A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent transportation technology, and in particular to a dynamic target tracking method based on multiple cameras and related equipment. Background Art
[0002] With the widespread application of intelligent transportation systems, multi-camera dynamic target tracking technology has become a core requirement for scenarios such as vehicle monitoring and behavior analysis. Existing cross-camera target tracking systems typically rely on a three-dimensional spatial coordinate system for target positioning, but in practical applications, they face the following technical bottlenecks: Lack of dynamic target prediction: Traditional technologies primarily associate target positions with static coordinates and lack effective modeling of the temporal dimension. When a target vehicle is moving at high speeds, existing systems struggle to accurately predict its future trajectory, leading to target loss or tracking interruptions during cross-camera tracking. For example, on highways, the inability to predict the vehicle's next position can delay camera switching, disrupting the tracking chain.
[0003] Poor environmental adaptability: Existing technologies often rely on pre-calibrated and fixed camera extrinsic parameters. However, in complex real-world applications, sudden changes in lighting conditions, weather disturbances such as rain and fog, or temporary occlusions can cause inaccurate mapping relationships for these fixed parameters. For example, when camera images are overexposed due to strong light, traditional calibration parameters may become ineffective, increasing the coordinate mapping error of the target between adjacent cameras, leading to trajectory breakage and compromising system robustness and tracking accuracy.
[0004] In view of the above problems, the existing technology is in urgent need of improvement. Summary of the Invention
[0005] The purpose of this application is to provide a dynamic target tracking method and related equipment based on multiple cameras, which can effectively predict the position of dynamic targets and adaptively adjust camera parameters to improve environmental adaptability, thereby achieving accurate and robust tracking of dynamic targets.
[0006] In a first aspect, the present application provides a multi-camera-based dynamic target tracking method, which tracks a vehicle based on a multi-camera tracking system having multiple cameras with adjustable directions, comprising the steps of: A1. Real-time acquisition of a time-stamped frame image captured by the current camera, identifying a target vehicle from the frame, and obtaining identification features of the target vehicle; the current camera is a camera whose field of view covers the target vehicle; the identification features include pixel position coordinates; A2. Based on the real-time intrinsic and extrinsic parameters of the current camera, convert the pixel position coordinates into world coordinates to generate a target vehicle trajectory segment corresponding to the current camera. If the current camera orientation is the initial orientation, the real-time intrinsic and extrinsic parameters are pre-calibrated for the initial orientation. A3. When the target vehicle enters the boundary area of the current camera's field of view, the target vehicle's motion trajectory segment corresponding to the current camera is predicted to have a world coordinate position after a preset buffer time, recorded as the predicted target position; A4. If the predicted target position is out of the field of view of the multi-camera tracking system, execute step A6; otherwise, execute step A5; A5. Determine a neighboring camera based on the predicted target location, adjust the orientation of the neighboring camera to align with the predicted target location, update the intrinsic and extrinsic parameters of the neighboring camera, and use the neighboring camera as the new current camera, returning to step A1; A6. Fuse all the target vehicle motion trajectory segments to obtain a total motion trajectory.
[0007] In a second aspect, the present application provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program executable by the processor, and when the processor executes the computer program, it runs the steps in the dynamic target tracking method based on multiple cameras as described above.
[0008] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, runs the steps of the multi-camera-based dynamic target tracking method as described above.
[0009] Beneficial effect: The present application provides a dynamic target tracking method and related equipment based on multiple cameras. By introducing a target position prediction mechanism based on the time dimension and combining it with a strategy of dynamically adjusting the camera orientation and updating internal and external parameters in real time, it effectively solves the problems of lack of dynamic target prediction and poor environmental adaptability in the existing technology, thereby improving the accuracy and robustness of dynamic target tracking. It has the advantages of being able to effectively predict the position of dynamic targets and adaptively adjust camera parameters to improve environmental adaptability, thereby achieving accurate and robust tracking of dynamic targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This application provides a flowchart of a dynamic target tracking method based on multiple cameras.
[0011] Figure 2 This is a schematic diagram of the structure of an electronic device provided in this application.
[0012] Description of reference numerals: 301, processor; 302, memory; 303, communication bus. DETAILED DESCRIPTION
[0013] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.
[0014] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0015] Please refer to Figure 1 In some embodiments of the present application, a dynamic target tracking method based on multiple cameras is provided, wherein a vehicle is tracked based on a multiple-camera tracking system having multiple cameras with adjustable directions, comprising the steps of: A1. Acquire time-stamped frame images from the current camera in real time, identify the target vehicle, and obtain the target vehicle's identification features. The current camera's field of view covers the target vehicle; the identification features include pixel position coordinates. A2. Based on the current camera's real-time intrinsic and extrinsic parameters, convert the pixel position coordinates into world coordinates to generate the target vehicle's trajectory segment corresponding to the current camera. If the current camera's orientation is the initial orientation, the real-time intrinsic and extrinsic parameters are pre-calibrated for the initial orientation. A3. When the target vehicle enters the boundary area of the current camera's field of view, predict the target vehicle's world coordinate position after a preset buffer time based on the target vehicle's trajectory segment corresponding to the current camera. This is recorded as the predicted target position. A4. If the predicted target position is out of the field of view of the multi-camera tracking system, execute step A6; otherwise, execute step A5; A5. Determine a neighboring camera based on the predicted target position, adjust the orientation of the neighboring camera to align with the predicted target position, update the intrinsic and extrinsic parameters of the neighboring camera, and use the neighboring camera as the new current camera, returning to step A1; A6. Fuse all target vehicle trajectory segments to obtain the total trajectory.
[0016] Among them, the multi-camera tracking system refers to a monitoring network composed of multiple cameras, which can be implemented by distributed deployment or centralized management, such as deployment at urban traffic intersections or along highways. Its main purpose is to achieve continuous coverage and tracking of targets in a large area.
[0017] Among them, a camera with adjustable direction refers to a camera whose shooting direction can be adjusted according to instructions. It can be achieved by electric pan-tilt drive or robotic arm control, for example, by controlling its rotation and pitch angle through a stepper motor or servo motor. It is mainly used to achieve flexible capture of dynamic targets and field of view switching.
[0018] Among them, frame images with timestamps mean that each frame image is accompanied by precise time information, which can be achieved by using system clock synchronization or GPS timing modules. For example, the acquisition time is embedded in the image data packet. Its main purpose is to model the time dimension of target motion and provide a basis for subsequent trajectory prediction.
[0019] Among them, real-time intrinsic parameters and real-time extrinsic parameters refer to the internal parameters and external parameters of the camera in its current working state, and these parameters can be dynamically acquired or updated according to actual conditions. They can be implemented using online calibration algorithms or vision-based self-calibration technologies, such as by identifying known feature points in the scene or using deep learning models for parameter estimation. The main purpose is to achieve the accuracy of coordinate transformation to adapt to environmental changes or camera posture adjustments. Each camera has a default orientation, namely the initial orientation. The intrinsic and extrinsic parameters under this initial orientation are pre-calibrated (for example, by calibrating with a checkerboard). When the system starts, the orientation of each camera is the initial orientation. When the target vehicle enters the field of view of the first camera, the orientation of the first camera remains at the initial orientation. Therefore, its real-time intrinsic and real-time extrinsic parameters are equal to the pre-calibrated intrinsic and extrinsic parameters.
[0020] Among them, world coordinates refer to the position representation of the target in a unified three-dimensional space coordinate system, which can be implemented using a Cartesian coordinate system or a geodetic coordinate system, such as the X, Y, and Z coordinates with a fixed point as the origin. Its main purpose is to achieve a unified description of the target position and trajectory fusion under the field of view of different cameras.
[0021] Among them, the target vehicle motion trajectory segment refers to the world coordinate sequence of the target vehicle at different time points within the field of view of a single camera. It can be generated using linear interpolation or curve fitting methods, such as connecting a series of spatiotemporal coordinate points. Its main purpose is to record and analyze the local motion laws of the target and provide data support for subsequent motion prediction.
[0022] Among them, the field of view boundary area refers to the specific range at the edge of the camera's field of view, which can be defined by an image pixel distance threshold or a geometric projection area, such as a fixed pixel width area at the edge of the image. Its main purpose is to provide an early warning that the target is about to leave the current field of view and trigger subsequent prediction and switching mechanisms.
[0023] Among them, the predicted target position refers to the world coordinates of the target at a certain point in the future inferred based on the historical motion trajectory of the target. It can be implemented using Kalman filtering or deep learning prediction models. For example, the future position is calculated based on the current speed and direction of the target. Its main purpose is to achieve early prediction of the target movement and provide a forward-looking basis for camera switching.
[0024] Among them, the preset buffer time refers to the time interval from the target entering the boundary area of the field of view to the predicted target position. It can be implemented using a fixed time value or dynamically adjusted according to the target speed. Its main purpose is to reserve sufficient time for camera adjustment and switching to ensure the continuity of tracking.
[0025] Among them, the adjacent camera refers to a camera that is near the predicted target position and can relay tracking of the target. It can be determined based on geographic distance or field of view overlap, such as the camera that is closest to the predicted position and has covered field of view. It is mainly used to achieve smooth handover of tracking tasks.
[0026] The core innovation of this application lies in the introduction of frame images with timestamps and real-time updated camera parameters, combined with a dynamic target position prediction mechanism based on motion trajectory segments, and on this basis, the realization of prediction-driven camera intelligent switching and parameter adaptive updating, thereby solving the problems of lack of dynamic target prediction and poor environmental adaptability in the existing technology, and achieving the effect of continuous and reliable tracking of dynamic targets.
[0027] Through the above scheme, the present application effectively solves the problems of lack of dynamic target prediction and poor environmental adaptability in the existing dynamic target tracking methods based on multiple cameras. Specifically, by introducing frame images with timestamps and generating motion trajectory segments containing time information, the present application can model the motion of the target vehicle in the time dimension, thereby achieving accurate prediction of the future position of the target, avoiding tracking interruption and target loss due to the inability to predict the target movement. At the same time, by adopting real-time updated camera internal and external parameters, and combining prediction-driven camera orientation adjustment and parameter adaptive update mechanism, the present application significantly enhances the adaptability of the system in complex and changing environments, ensures the accuracy of coordinate transformation and the continuity of tracking, and effectively avoids the problem of trajectory breakage caused by inaccuracy of fixed parameters. Overall, the present application realizes continuous and reliable tracking of dynamic targets, and improves the reliability and practicality of the tracking system.
[0028] Among them, the identification features may also include license plate number, vehicle model and vehicle color.
[0029] In some embodiments, step A1 comprises: A101. Obtain the time-stamped frame images captured by the current camera in real time. A102. Use the YOLOv8 algorithm to perform target detection on the frame image and obtain a target vehicle identification box to determine the pixel position coordinates of the target vehicle. A103. Use optical character recognition technology to identify the image within the identification frame and obtain the target vehicle's license plate number; A104. Use a deep learning classifier to classify the image within the identification box to obtain the model and color of the target vehicle.
[0030] The YOLOv8 algorithm is a real-time object detection algorithm based on deep learning. It can be implemented using a single-stage detection framework. It achieves efficient and accurate target localization by directly predicting the bounding box (i.e., landmark box) and category of the target in the image. The pixel coordinates of the center point of the landmark box can be extracted as the pixel position coordinates of the target vehicle.
[0031] Among them, optical character recognition technology is a technology that converts text characters in images into machine-editable text. It can be implemented using an end-to-end model based on convolutional neural networks and recurrent neural networks, or by combining traditional image processing with machine learning, such as through image preprocessing, character segmentation, feature extraction, and classifier recognition.
[0032] Among them, the deep learning classifier is an image classification model based on deep neural networks, which can be implemented using a convolutional neural network architecture, such as ResNet, VGG, or EfficientNet, to classify objects in the image by learning the high-level semantic features of the image.
[0033] By incorporating license plate number, vehicle model, and color as identification features, and employing the YOLOv8 algorithm for target detection, optical character recognition for license plate number, and a deep learning classifier for vehicle model and color, this solution enhances the accuracy and stability of target vehicle identification in a multi-camera dynamic target tracking system. This overcomes the potential for tracking interruptions and misidentifications that can occur when relying solely on pixel position coordinates in complex scenarios, such as multi-camera switching, the presence of multiple similar vehicles, or temporary occlusion of the target vehicle. This solution provides more reliable identity confirmation information, enabling the system to stably and continuously track the target vehicle.
[0034] Specifically, the intrinsic parameters include the intrinsic matrix, and the extrinsic parameters include the translation matrix and the rotation matrix.
[0035] The intrinsic parameter matrix is a mathematical model that describes the internal optical properties and imaging geometry of the camera. It is used to map the two-dimensional pixel coordinates on the image plane to the three-dimensional spatial coordinates in the camera coordinate system. The intrinsic parameter matrix typically contains parameters such as focal length, principal point coordinates, and distortion coefficients. The extrinsic parameters, namely the translation matrix and rotation matrix, describe the position and posture of the camera in the world coordinate system. They transform the three-dimensional coordinates in the camera coordinate system into a unified world coordinate system. The translation matrix represents the displacement of the camera's optical center relative to the origin of the world coordinate system, while the rotation matrix represents the rotation of the camera coordinate system relative to the world coordinate system. By clarifying the specific form of these parameters, a clear and standard mathematical model can be provided for subsequent precise coordinate transformations, ensuring the accuracy and operability of the transformation.
[0036] In some embodiments, step A2 comprises: A201. Convert the pixel position coordinates to camera coordinates in the camera coordinate system of the current camera based on the real-time internal parameters of the current camera. A202. Convert the camera coordinates to 3D world coordinates based on the real-time external parameters of the current camera. A203. Add the timestamp of the current frame image to the three-dimensional world coordinates to form four-dimensional space-time coordinates as the final world coordinates; A204. Generate a target vehicle motion trajectory segment with time information corresponding to the current camera based on the final world coordinates.
[0037] The coordinate conversion methods of step A201 and step A202 are prior art and will not be described in detail here.
[0038] The timestamp of the current frame image is added to the 3D world coordinates to form 4D space-time coordinates, which serve as the final world coordinates. This process expands the static 3D spatial coordinates into 4D space-time coordinates that include time information. By binding time information to spatial position, each target location carries the time information of its occurrence, enabling more accurate capture of the target's motion state and providing critical time dimension data for subsequent dynamic target motion analysis and prediction.
[0039] Among them, the four-dimensional space-time coordinate points can be connected in time sequence to form a trajectory segment containing time information, that is, the target vehicle motion trajectory segment is obtained.
[0040] By clarifying the specific composition of the camera's intrinsic and extrinsic parameters and innovatively integrating the timestamp of the current frame image into the three-dimensional world coordinates to form four-dimensional space-time coordinates, this method can overcome the limitations of the traditional three-dimensional space coordinate system in dynamic target modeling. This four-dimensional space-time coordinate with time information allows each target position point to be accurately bound to the time of its occurrence, thereby accurately reflecting the dynamic motion state of the target vehicle and effectively capturing its motion patterns even in high-speed moving scenarios. As a result, subsequent motion trajectory predictions can be based on more comprehensive and accurate space-time data, significantly improving the accuracy of the prediction. This directly guarantees the continuity and accuracy of cross-camera tracking, effectively avoids the problem of target loss due to inaccurate predictions, and improves the robustness of the entire tracking system.
[0041] In some embodiments, step A3 comprises: A301. According to the current pixel position coordinates of the target vehicle and the pixel size of the frame image, determine whether the target vehicle enters the boundary area of the current camera's field of view; A302. If the target vehicle enters the boundary area of the current camera's field of view, the current moving speed of the target vehicle is calculated based on the target vehicle's trajectory segment corresponding to the current camera; A303. Calculate the predicted target position based on the preset buffer time and the current moving speed.
[0042] The field of view boundary area can be defined using a preset pixel distance threshold or a percentage of the image size, for example, a certain pixel range inward from the image edge or a certain percentage range of the image width / height. The pixel size of a frame image refers to the width and height of the current frame image, in pixels.
[0043] The current moving speed refers to the instantaneous speed of the target vehicle at the current moment or the average speed in the recent period. It can be represented by a velocity vector or velocity component, for example, , where are the x-, y-, and z-axis velocity components, respectively, and T represents the transposed sign. Linear regression analysis or the finite difference method can be used to calculate the target vehicle's velocity components along the x-, y-, and z-axis world coordinates, thereby obtaining its current velocity vector.
[0044] The preset buffer time is actually a period of time reserved for camera switching and adjustment operations. It can be set to a fixed time value, a time value dynamically adjusted according to system load, or a time value adaptively adjusted according to the target vehicle speed.
[0045] Among them, the predicted target position can be calculated by the following formula: ; in, 、 、 are the x, y, and z axis coordinate values of the predicted target position, 、 、 are the x, y, and z axis coordinate values of the target vehicle’s current position, The preset buffer time.
[0046] This solution effectively solves the problem of insufficient prediction accuracy in the existing technology by refining the prediction process of the target position. By judging whether the target vehicle has entered the boundary area of the field of view based on the current pixel position coordinates and the pixel size of the frame image, the system can trigger the prediction mechanism in real time and accurately, avoiding prediction lag. Furthermore, the current moving speed of the target vehicle is calculated using the target vehicle's motion trajectory segment, so that the prediction is no longer a simple static estimate, but fully considers the actual dynamic motion state of the target vehicle, significantly improving the accuracy of the prediction. Finally, the predicted target position is determined by combining the preset buffer time and the accurately calculated current moving speed, providing a reliable and forward-looking basis for subsequent camera switching and adjustment. This enables the system to switch cameras more promptly and smoothly, thereby ensuring continuous and stable target tracking even when the target vehicle is moving at high speed or in a complex motion state, effectively avoiding target loss.
[0047] Preferably, step A301 may include: The feature distance from the edge of the target vehicle in the current camera's field of view is obtained according to the following formula: L=min(u',1-u',v',1-v'); Where L is the feature distance (indicates the closest distance between the target vehicle and the boundary of the current camera's field of view), u' is the normalized value of the horizontal coordinate in the current pixel position coordinate of the target vehicle, and v' is the normalized value of the vertical coordinate in the current pixel position coordinate of the target vehicle; If the feature distance from the edge is less than or equal to the preset distance from the edge threshold, it is determined that the target vehicle has entered the boundary area of the field of view of the current camera; otherwise, it is determined that the target vehicle has not entered the boundary area of the field of view of the current camera.
[0048] Normalization maps the original pixel coordinates to floating-point values between 0 and 1, where 0 represents the left or top edge of the image and 1 represents the right or bottom edge. This normalization process eliminates the effects of varying camera resolutions, ensuring universality in subsequent judgments.
[0049] The preset distance-from-edge threshold is a critical value used to determine whether the target vehicle enters the boundary area of the field of view. It can be set based on the needs of actual application scenarios through experiments or experience to balance the timeliness and stability of tracking. For example, it can be 0.1, but is not limited to this.
[0050] This solution introduces the concept of feature off-edge distance and uses the normalized values of the target vehicle's pixel position coordinates for calculation, providing an accurate, standardized, and robust method for determining whether the target vehicle has entered the boundary area of the current camera's field of view. This method overcomes the limitations of relying solely on raw pixel coordinates and frame image pixel size for judgment, and can accurately identify when the target vehicle is truly at the edge of the camera's effective tracking range. Therefore, this solution can optimize the triggering timing of target prediction and camera switching, ensuring that the subsequent target position prediction steps can be triggered in a consistent and reliable manner under different camera configurations and scenarios, thereby ensuring the continuity and smoothness of the multi-camera tracking system when the target vehicle crosses the camera's field of view.
[0051] Specifically, step A4 may include: If the predicted target position is outside the optical center scanning range of all cameras except the current camera, it is determined that the predicted target position is out of the field of view of the multi-camera tracking system.
[0052] The optical center scanning range refers to the range of positions that the optical center can reach within the camera's orientation adjustment range. If the predicted target position is outside the optical center scanning range of all cameras other than the current camera, it indicates that the optical centers of the other cameras cannot align with the predicted target position, and thus no camera is available for relaying. At this point, step A6 is executed to fuse all target vehicle trajectory segments to obtain the total trajectory.
[0053] In some embodiments, step A5 comprises: A501. Calculate the distance between the predicted target position and the field of view of all cameras other than the current camera, and record it as the deviation distance. A502. Select another camera whose deviation distance is less than the preset distance threshold as a neighboring camera; A503. Based on the predicted target position, the current rotation angle of the neighboring camera, the current pitch angle of the neighboring camera, and the real-time external parameters of the neighboring camera, calculate the rotation angle adjustment amount and the pitch angle adjustment amount required to align the optical center of the neighboring camera with the predicted target position; A504. Adjust the rotation and pitch angles of the adjacent cameras according to the rotation and pitch adjustment amounts. A505. Obtain three consecutive frames of images captured by the adjacent camera after adjusting the rotation angle and pitch angle, input the three consecutive frames of images into the pre-trained deep learning calibration model, and obtain the external parameter correction value output by the deep learning calibration model; A506. Update the extrinsic parameters of the neighboring cameras according to the extrinsic parameter correction; A507. Calculate the zoom factor based on the pixel size of the identification frame obtained by the neighboring camera to identify the target vehicle and the actual size of the target vehicle to adjust the focal length of the neighboring camera; A508. Update the intrinsic parameters of the adjacent cameras according to the adjusted focal length; A509. After completing the adjustment of the internal and external parameters of the adjacent camera, use the adjacent camera as the new current camera and return to step A1.
[0054] In step A501, the field of view of each camera refers to the three-dimensional spatial area covered by each camera in its current posture (orientation). Specifically, the camera's intrinsic and extrinsic parameters, as well as its field of view angle parameters, can be used to define and calculate its field of view boundary. For example, a frustum model can be constructed for each camera to determine the location of the frustum boundary. The deviation distance can be calculated using various geometric distance metrics, such as the Euclidean distance from the predicted target position to the center point of each camera's field of view, or the closest distance from the predicted target position to the field of view boundary.
[0055] In step A502, the preset distance threshold is a configurable parameter used to limit the selection range of neighboring cameras. The specific setting can be based on the actual application scenario, camera deployment density, and tracking accuracy requirements. Various strategies can be used to select neighboring cameras. For example, a camera with the smallest deviation distance can be selected, a camera with a deviation distance within the threshold range and offering the best image quality and the lowest occlusion risk can be selected, or a camera with the smallest angular adjustment required to align its optical center with the predicted target position (e.g., the sum of the rotation and pitch angle adjustments) can be selected as the neighboring camera.
[0056] In step A503, the current rotation and pitch angles of the neighboring camera can be acquired in real time via the camera's own sensors (e.g., encoders). The rotation and pitch angle adjustments can be calculated using an inverse kinematics algorithm. For example, the world coordinates of the predicted target position are converted to coordinates in the neighboring camera's camera coordinate system. Then, based on the vector relationship between the neighboring camera's optical center and the target position, combined with the neighboring camera's current posture (rotation and pitch angles), the rotation and translation transformations required to align the optical axis with the target are calculated. These transformations are then decomposed into the rotation and pitch angle adjustments.
[0057] In step A504, the rotation and pitch angles of the adjacent camera can be adjusted by controlling the stepper motor or servo motor within the adjacent camera. For example, control commands can be sent to a PTZ (Pan-Tilt-Zoom) head to precisely rotate it according to the calculated adjustment amount. This adjustment process can utilize closed-loop control, ensuring accuracy by providing feedback on the camera's actual rotation angle.
[0058] In step A505, three consecutive frames of images are acquired to provide temporal information, helping the deep learning calibration model better understand scene changes and motion. The pre-trained deep learning calibration model can be a convolutional neural network (CNN) combined with a recurrent neural network (RNN) or a Transformer. The training data can include images under different lighting, viewing angles, and occlusion conditions, along with their corresponding true extrinsic parameter deviations. Extrinsic parameter corrections can be small adjustments to the rotation and translation matrices, such as rotation corrections expressed in Euler angles or Lie algebraic form, and three-dimensional translation corrections. For example, the extrinsic parameter corrections output by the deep learning calibration model can be represented in array form: ΔP = [Δrx, Δry, Δrz, Δtx, Δty, Δtz], where ΔP represents the extrinsic parameter corrections output by the deep learning calibration model, Δrx, Δry, and Δrz represent the rotation corrections along the three axes, and Δtx, Δty, and Δtz represent the translation corrections along the three axes.
[0059] In step A506, the extrinsic parameters of the neighboring camera can be updated by adding the correction value obtained in step A505 to the current real-time extrinsic parameters of the camera. For example, if the extrinsic parameter correction value is an increment of the rotation matrix and the translation matrix, it can be updated by matrix multiplication and vector addition. Specifically, it can be updated by the following formula: ; ; ; ; in, is the updated rotation matrix, is the current rotation matrix of the neighboring camera, is the updated translation matrix, is the current translation matrix of the neighboring camera, is a vector composed of three axial rotation corrections, is a vector composed of three axial translation corrections, Indicates based on The resulting rotation matrix.
[0060] In step A507, the pixel size of the identification box obtained by the neighboring cameras identifying the target vehicle can be obtained using an object detection algorithm (such as YOLOv8). The actual size of the target vehicle can be obtained from a preset database based on its vehicle model information. For example, different vehicle models such as sedans, SUVs, and trucks have different average lengths, widths, and heights. The zoom factor can be calculated based on the principles of perspective projection. For example, by comparing the ratio of the target vehicle's pixel size in the image to its actual size and combining it with the current focal length, the required focal length adjustment can be calculated to achieve the desired size of the target in the image.
[0061] In step A508, updating the intrinsic parameters of the adjacent cameras refers to updating the focal length parameters in the intrinsic parameter matrix. The intrinsic parameter matrix typically includes the focal length parameters, principal point coordinate parameters, and distortion coefficients. When the focal length is adjusted, the corresponding focal length parameters in the intrinsic parameter matrix are updated.
[0062] In step A509, using the neighboring camera as the new current camera means that the system switches the current tracking task's leading camera to the neighboring camera, using its latest internal and external parameters as real-time parameters. Returning to step A1 means that the system will continue acquiring images, identifying targets, converting coordinates, and generating trajectory segments in real time, starting with the new current camera, thus achieving seamless cross-camera tracking.
[0063] This solution effectively solves the problem of inaccurate updates of intrinsic and extrinsic parameters in multi-camera dynamic target tracking by introducing a series of refined steps. By calculating the deviation distance between the predicted target position and the field of view of other cameras and selecting adjacent cameras, it ensures that the most suitable camera to take over tracking is selected when switching, avoiding unnecessary switching or inappropriate selection. By accurately calculating and adjusting the rotation and pitch angles of adjacent cameras based on the predicted target position, it ensures that the camera can accurately align with the future position of the target vehicle, improving the efficiency and accuracy of target capture. Of particular importance, by obtaining continuous images collected after the orientation is adjusted and inputting them into a pre-trained deep learning calibration model to obtain extrinsic parameter corrections, and then updating the extrinsic parameters accordingly, this solution can adaptively and robustly correct extrinsic parameter deviations caused by changes in camera orientation or environmental factors, significantly improving the accuracy of the extrinsic parameters, thereby ensuring the accuracy of the conversion of pixel position coordinates to world coordinates. Furthermore, by calculating the zoom factor and adjusting the focal length based on the pixel size and actual size of the target vehicle's identification frame, and then updating the internal parameters, this solution dynamically optimizes image quality, ensuring that the target vehicle is presented at the optimal size in the new camera field of view, improving the clarity and accuracy of target recognition. These improvements work together to enable the system to achieve seamless, high-precision tracking of target vehicles across different camera fields of view, significantly reducing the risk of target loss and improving the continuity and reliability of overall tracking.
[0064] In some possible implementations, the deep learning calibration model includes: ResNet-34 network is used to extract image features of three consecutive frames and obtain three image feature vectors; ConvLSTM network is used to perform time series modeling on the three image feature vectors to obtain time series features; The fully connected network is used to perform parameter regression processing on the time series features to obtain and output the external parameter correction.
[0065] The ResNet-34 network is a deep convolutional neural network architecture based on residual connections. It can be implemented as a stack of multiple residual blocks. Each residual block contains a convolutional layer, a batch normalization layer, and an activation function layer. Skip connections are used to directly add the input to the output, alleviating the vanishing gradient problem in deep network training and effectively extracting multi-level spatial features from images. The ConvLSTM network is a recurrent neural network that combines the characteristics of convolution and long short-term memory networks. It can be implemented as a sequence of units containing convolutional gated units (such as input gates, forget gates, and output gates) and a convolutional state update mechanism. Each unit uses convolution operations to process input and state, thereby capturing dependencies in both spatial and temporal dimensions when processing sequential data. A fully connected network is a neural network structure composed of multiple stacked fully connected layers (or dense layers). It can be implemented as a multilayer perceptron structure consisting of an input layer, one or more hidden layers, and an output layer. Neurons in each layer are connected to all neurons in the previous layer. Weights and biases are used to perform linear transformations and nonlinear activations on the input data, thereby achieving complex mapping and regression of input features.
[0066] The deep learning calibration model works as follows: When adjacent cameras adjust their orientation, three consecutive frames are captured. These images are fed into a ResNet-34 network. The ResNet-34 network, acting as an image feature extractor, performs deep convolution on each frame, capturing rich spatial details and semantic information, and generates an image feature vector for each frame. Since these three consecutive frames are captured, three independent image feature vectors are generated, representing the spatial content of the image at different points in time. These three image feature vectors are then fed into a ConvLSTM network. The ConvLSTM network plays a role in temporal modeling. It not only processes the spatial information of each image feature vector but, more importantly, captures the dynamic correlations and changing trends between these three consecutive image feature vectors in the temporal dimension. Through its internal convolutional gating mechanism, the ConvLSTM network effectively learns and memorizes the temporal dependencies in the image sequence, integrating the three independent image feature vectors into a comprehensive temporal feature containing both spatial and temporal information. This temporal modeling is crucial for understanding subtle changes in camera pose, as misalignment of extrinsic parameters is often a gradual or dynamic process. Finally, the temporal features output by the ConvLSTM network are fed into a fully connected network. This network performs parameter regression on these high-dimensional temporal features. Through multiple layers of nonlinear transformations, it maps the complex temporal features into low-dimensional, specific numerical values, known as extrinsic parameter corrections. These corrections accurately reflect the deviation between the camera's current extrinsic parameters and the true extrinsic parameters. In this way, the deep learning calibration model can accurately calculate the required rotation angle and translation matrix adjustments based on the temporal changes in three consecutive image frames.
[0067] The deep learning calibration model is closely integrated with the steps in the aforementioned method to form a complete calibration closed loop. After acquiring three consecutive frames of images, the model receives these images as input and outputs an extrinsic parameter correction. This correction is then directly used to update the extrinsic parameters of the adjacent camera. This combination enables the system to perform fine-grained calibration of extrinsic parameters in real time and dynamically, overcoming the limitations of traditional fixed parameters in complex environments. By making full use of the spatial and temporal features of the image, the model can more accurately perceive subtle changes in the camera posture, thereby providing more accurate extrinsic parameter corrections, thereby ensuring the accuracy and stability of multi-camera dynamic target tracking, especially in scenarios with frequent camera switching and environmental changes, significantly improving the environmental adaptability and robustness of the system.
[0068] Preferably, step A507 may include: B1. Obtaining a frame image captured by a neighboring camera, performing vehicle identification, and obtaining at least one vehicle identification feature; identification features also include license plate number, vehicle model, and vehicle color; B2. Calculate the similarity between each vehicle identified by the neighboring camera and the target vehicle based on the target vehicle's identification features (license plate number, vehicle model, and vehicle color) acquired by the current camera and the neighboring camera. B3. Determine the target vehicle from among the vehicles identified by the neighboring cameras based on similarity; B4 obtains the pixel size of the identification frame of the target vehicle determined, recorded as the first pixel size, and obtains the corresponding actual size according to the model of the target vehicle determined; B5. Calculate the ratio of the actual size to the first pixel size to obtain the actual conversion ratio. Combined with the preset target conversion ratio and the current focal length of the adjacent camera, calculate the zoom factor that can make the actual conversion ratio equal to the target conversion ratio. B6. Adjust the focus of the adjacent camera according to the zoom ratio.
[0069] Among them, step B1 can refer to step A1 above and will not be repeated here.
[0070] Among them, similarity refers to a quantitative indicator that measures the degree of proximity between two or more entities in a specific dimension or multiple dimensions. It can be calculated using metrics based on feature distance (such as Euclidean distance, Manhattan distance), scoring based on feature matching (such as hash matching, string edit distance), or classification probability based on machine learning models.
[0071] In step B3, a threshold determination method can be used to determine the target vehicle. For example, a preset similarity threshold (e.g., 95%) can be set, and vehicles with similarities exceeding this threshold are selected as candidate target vehicles. If only one vehicle has a similarity above the threshold, it is determined as the target vehicle. If multiple vehicles have similarities above the threshold, the vehicle with the highest similarity can be selected as the target vehicle. Furthermore, if multiple vehicles have similarities above the threshold, the vehicle's trajectory information can be used to assist in the determination, for example, selecting the vehicle whose direction and speed more closely match the predicted trajectory of the target vehicle.
[0072] In step B4, the pixel dimensions of the identification box can be directly obtained from the output of the object detection algorithm, such as the width and height of the identification box. The actual dimensions of the vehicle can be obtained by querying a pre-established vehicle model database. This database may contain length, width, and height information for various common vehicle models. Once the target vehicle model is determined, the system can automatically retrieve the corresponding actual dimensions from the database.
[0073] In step B5, the actual conversion ratio can be defined as the ratio of the vehicle's actual physical dimensions (e.g., length or width) to the corresponding pixel size in the image. The preset target conversion ratio is the desired size ratio of the target vehicle in the new camera field of view. For example, it can be preset so that the target vehicle occupies a certain percentage of the image height. The zoom factor calculation can be based on optical imaging principles, such as the relationship between focal length, object distance, and image size. By comparing the actual observed size ratio with the desired target conversion ratio and combining it with the current camera focal length, the required zoom factor can be accurately calculated. This ensures that the focus adjustment enables the target vehicle to achieve the preset ideal size and clarity in the new camera field of view, optimizing tracking effectiveness.
[0074] By introducing a multi-feature similarity matching mechanism, this solution can accurately identify target vehicles in the field of view of adjacent cameras, effectively solving the problem of difficulty in accurately determining the tracking target in multi-target scenarios or when there are subtle differences in target features. This ensures the accuracy of the pixel size of the identification frame used to calculate the zoom factor and the actual size of the target vehicle, thereby avoiding size information deviations caused by target recognition errors. Therefore, this solution can achieve precise adjustment of the focal length of adjacent cameras, allowing the target vehicle to quickly obtain clear and stable images after switching to the new camera's field of view. This significantly improves the smoothness of target vehicle switching and the robustness of tracking during cross-camera tracking, ensuring continuous and clear tracking of the target vehicle.
[0075] Preferably, step B2 may include: B201. Calculate the license plate number consistency parameter according to the following formula: ; in, is the license plate number consistency parameter, The encoding of the license plate number of the target vehicle obtained by the current camera, The license plate number of the vehicle captured by the nearby camera. for and The coding distance between for The encoding length, for The encoding length of ; where the encoding distance is a measure of the difference between two encoding sequences and is a dimensionless value; the encoding length refers to the number of characters or bytes in the encoding sequence and is a dimensionless value; B202. Convert the color space of the target vehicle captured by the current camera and the vehicle color captured by the neighboring camera into the HSV color space, and calculate the color distance using the following formula: ; in, is the color distance, 、 、 The H, S, and V channel values of the target vehicle's color are obtained by the current camera. 、 、 are the H, S, and V channel values of the vehicle color obtained by the neighboring cameras, is the preset reference threshold; B203. Determine the category similarity score Ts between the target vehicle model captured by the current camera and the vehicle model captured by the adjacent camera using a hierarchical scoring rule designed based on the vehicle model classification tree; B204. Calculate the similarity between the vehicle identified by the neighboring camera and the target vehicle using the following formula: ; in, is the similarity, 、 、 is the preset weight coefficient.
[0076] The license plate number consistency parameter is used to quantify the numerical value of the similarity between two license plate numbers. Encoding refers to converting the character sequence of the license plate number into a computable form, which can be achieved using ASCII encoding, Unicode encoding, or a custom digital encoding. The encoding distance is a measure of the difference between two encoding sequences, which can be calculated using algorithms such as Levenshtein distance, Hamming distance, or Jaccard distance. The encoding length refers to the number of characters or bytes in the encoding sequence, which can be obtained using a string length function or a byte count function.
[0077] The HSV color space is a color model that represents colors as hue, saturation, and value, and can be implemented using standard color space conversion algorithms. Color distance is a numerical value that quantifies the difference between two colors in the HSV color space. The preset reference threshold is a pre-set limit used to determine color similarity in color distance calculations. This threshold can be determined based on empirical values, statistical analysis, or machine learning methods.
[0078] Among them, the vehicle model classification tree refers to a classification structure that organizes vehicle models in a hierarchical relationship, which can be constructed using decision trees, hierarchical clustering or ontology. The hierarchical scoring rules refer to a set of rules that assign scores based on the similarity of different levels and branches in the vehicle model classification tree. It can be designed based on methods such as common ancestor nodes, path length or semantic similarity. The category similarity score refers to the numerical value of the similarity between two vehicle model categories calculated according to the hierarchical scoring rules, which can be expressed using methods such as normalized scoring, percentage matching or fuzzy logic. For example, the hierarchical scoring rules may include: If the vehicle types are of the same category (e.g., both are sedans), the category similarity score Ts is 1; If the vehicle types belong to the same category (e.g., SUV and MPV), the category similarity score Ts is 0.7; If the vehicle type categories do not belong to the same general category (for example, cars and trucks), the category similarity score Ts is 0.1.
[0079] The preset weight coefficient refers to a value used to balance the importance of different features in the comprehensive similarity calculation, which can be determined by expert experience, regression analysis, or optimization algorithm.
[0080] Through the above technical solutions, this method can effectively deal with problems existing in actual applications such as license plate recognition errors, color deviations caused by changes in lighting, and hierarchical relationships in vehicle classification. Specifically, the calculation of license plate number consistency parameters can handle subtle differences in license plate recognition and avoid misjudgments due to minor errors; color distance calculations in the HSV color space can effectively reduce the impact of environmental factors such as lighting on color recognition, improving the accuracy and stability of color judgment; hierarchical scoring rules based on vehicle classification trees can more accurately and reasonably evaluate vehicle model similarity, avoiding the limitations of simple matching. These multi-dimensional and refined similarity calculation methods, through weighted fusion, significantly improve the recognition accuracy and robustness of target vehicles when switching across cameras, thereby effectively solving the problems of misidentification or target loss caused by a single or rough similarity calculation method, thereby improving the accuracy and stability of the entire tracking system.
[0081] In some embodiments, step A6 includes: A601. Based on the four-dimensional space-time coordinates of the trajectory points of each target vehicle's trajectory segment, all target vehicle trajectory segments are fused using a nearest neighbor matching algorithm to obtain a preliminary total trajectory. A602. Use the Kalman filter method to smooth the preliminary total motion trajectory to obtain the final total motion trajectory.
[0082] Among them, the nearest neighbor matching algorithm refers to an algorithm used to find the data point closest to a given data point in a data set. It can use Euclidean distance, Mahalanobis distance or other distance metrics to calculate the similarity between trajectory points and select the point with the smallest distance as the matching object.
[0083] Among them, the Kalman filtering method refers to an algorithm used to estimate the state of a dynamic system, which is particularly suitable for processing measurement data containing noise. It can use a prediction-update cycle method, combined with the system's motion model and measurement data, to estimate and correct the system state in real time, thereby achieving data smoothing and noise filtering.
[0084] This solution utilizes two key steps, working in tandem, to address the issues of continuity and accuracy in trajectory fusion in multi-camera tracking. First, in the preliminary fusion stage, all target vehicle trajectory segments are fused using a nearest neighbor matching algorithm based on the four-dimensional spatiotemporal coordinates of the trajectory points within each segment to produce a preliminary overall trajectory. The four-dimensional spatiotemporal coordinates used for the trajectory points not only include the target's spatial position but also its temporal information. This inclusion of the temporal dimension enables accurate identification and association of trajectory points from the same target vehicle captured by different cameras at different time points during segment matching. Based on these temporally informed coordinates, the nearest neighbor matching algorithm can handle overlaps or gaps between different trajectory segments, initially stitching together scattered trajectory points into a single continuous trajectory. This spatiotemporal coordinate-based matching solves the problem of target trajectory continuity in cross-camera tracking, laying the foundation for subsequent smoothing.
[0085] On this basis, during the smoothing phase, the Kalman filter method is used to smooth the preliminary total motion trajectory to obtain the final total motion trajectory. The preliminary fused trajectory may still contain jitter and unevenness due to measurement noise, environmental interference, or matching errors. The Kalman filter method, as an algorithm, can dynamically adjust the estimated values of the trajectory points based on the target's motion model and measurement data, thereby filtering out random errors in the trajectory, making the final total motion trajectory smooth, continuous, and consistent with the target's actual motion patterns. It is precisely because the four-dimensional spatiotemporal coordinates with time information and the nearest neighbor matching algorithm are first used for preliminary trajectory fusion, and the fused trajectory is then smoothed through Kalman filtering that this solution can jointly solve the continuity and accuracy issues of target vehicle trajectory fusion in a multi-camera environment, thereby providing a total motion trajectory with good quality and reliability.
[0086] In a specific embodiment, the step of fusing all target vehicle trajectory segments to obtain a total trajectory can be implemented as follows. First, during the preliminary total trajectory generation phase, the system receives trajectory segments for each target vehicle captured by different cameras. Each trajectory segment consists of a series of trajectory points, each of which contains four-dimensional space-time coordinates, namely, a three-dimensional spatial position (X, Y, Z) and a corresponding timestamp (T). To fuse these trajectory segments, a global set of trajectory points can be constructed. Subsequently, a nearest neighbor matching algorithm can be used. For example, each trajectory point in a trajectory segment can be traversed and the closest trajectory point in four-dimensional space-time can be found in other trajectory segments. Distance can be calculated using Euclidean distance, incorporating a weighting in the time dimension. When matching points are found, a predetermined matching threshold can be used to determine whether to connect or merge them. For example, if the spatial and temporal distance between two trajectory points is less than a specific threshold, they are considered to belong to the continuous motion of the same target, and the corresponding trajectory segments are then spliced. This process can be repeated until all relevant trajectory segments have been attempted to be fused, resulting in a preliminary total trajectory, which may still contain slight jitter.
[0087] After this, the Kalman filter method can be applied to smooth the preliminary total motion trajectory. Specifically, a state model can be defined for the target vehicle's motion. For example, the state vector can contain the vehicle's position and velocity in the X, Y, and Z directions. The measurement value is the four-dimensional space-time coordinate point in the preliminary total motion trajectory. The Kalman filter iterates through two stages: prediction and update. In the prediction stage, the state at the next moment is predicted based on the vehicle's motion model; in the update stage, the predicted state is corrected based on the actual measured trajectory point data. Through continuous iteration, the Kalman filter can estimate the target vehicle's true motion state and filter out random noise and errors introduced during the measurement process, so that the final total motion trajectory exhibits higher continuity and smoothness, more accurately reflecting the target vehicle's actual motion path.
[0088] Please refer to Figure 2, is a structural diagram of an electronic device provided in an embodiment of the present application. The present application provides an electronic device, including: a processor 301 and a memory 302. The processor 301 and the memory 302 are interconnected and communicate with each other through a communication bus 303 and / or other forms of connection mechanisms (not marked). The memory 302 stores a computer program executable by the processor 301. When the electronic device is running, the processor 301 executes the computer program to execute the dynamic target tracking method based on multiple cameras in any optional implementation of the above embodiments to achieve the following functions: A1. Acquire a frame image with a timestamp captured by the current camera in real time, identify the target vehicle therefrom, and obtain the identification features of the target vehicle; the current camera is a camera whose field of view covers the target vehicle; the identification features include pixel position coordinates; A2. Based on the real-time internal parameters and real-time external parameters of the current camera, the pixel position coordinates Converted into world coordinates to generate the target vehicle motion trajectory segment corresponding to the current camera; wherein, if the current orientation of the current camera is the initial orientation, the real-time intrinsic parameters and real-time extrinsic parameters are the intrinsic parameters and extrinsic parameters pre-calibrated under the initial orientation; A3. When the target vehicle enters the boundary area of the field of view of the current camera, the world coordinate position of the target vehicle after a preset buffer time is predicted according to the target vehicle motion trajectory segment corresponding to the current camera, and recorded as the predicted target position; A4. If the predicted target position is out of the field of view of the multi-camera tracking system, execute step A6, otherwise, execute step A5; A5. Determine a neighboring camera according to the predicted target position, adjust the orientation of the neighboring camera to align with the predicted target position, update the intrinsic parameters and extrinsic parameters of the neighboring camera, and use the neighboring camera as the new current camera, and return to step A1; A6. Fuse all target vehicle motion trajectory segments to obtain the total motion trajectory.
[0089] The embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the multi-camera-based dynamic target tracking method in any optional implementation of the above embodiment is executed to achieve the following functions: A1. Real-time acquisition of a frame image with a timestamp captured by the current camera, identification of the target vehicle therefrom, and acquisition of identification features of the target vehicle; the current camera is a camera whose field of view covers the target vehicle; the identification features include pixel position coordinates; A2. Based on the real-time internal parameters and real-time external parameters of the current camera, the pixel position coordinates are converted into world coordinates to generate a target vehicle motion trajectory segment corresponding to the current camera; wherein, if the current orientation of the current camera is the initial orientation, the real-time The intrinsic parameters and real-time extrinsic parameters are pre-calibrated intrinsic parameters and extrinsic parameters under the initial orientation; A3. When the target vehicle enters the boundary area of the field of view of the current camera, the world coordinate position of the target vehicle after a preset buffer time is predicted based on the target vehicle motion trajectory segment corresponding to the current camera, and recorded as the predicted target position; A4. If the predicted target position is out of the field of view of the multi-camera tracking system, execute step A6; otherwise, execute step A5; A5. Determine a neighboring camera based on the predicted target position, adjust the orientation of the neighboring camera to align with the predicted target position, update the intrinsic parameters and extrinsic parameters of the neighboring camera, and use the neighboring camera as the new current camera, and return to step A1; A6. Fuse all target vehicle motion trajectory segments to obtain the total motion trajectory.
[0090] Among them, the computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0091] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A dynamic target tracking method based on multiple cameras, which tracks a vehicle based on a multi-camera tracking system with multiple cameras with adjustable directions, characterized in that: Including steps: A1. Real-time acquisition of a time-stamped frame image captured by the current camera, identifying a target vehicle from the frame, and obtaining identification features of the target vehicle; the current camera is a camera whose field of view covers the target vehicle; the identification features include pixel position coordinates; A2. Based on the real-time intrinsic and extrinsic parameters of the current camera, convert the pixel position coordinates into world coordinates to generate a target vehicle trajectory segment corresponding to the current camera. If the current camera orientation is the initial orientation, the real-time intrinsic and extrinsic parameters are pre-calibrated for the initial orientation. A3. When the target vehicle enters the boundary area of the current camera's field of view, the target vehicle's motion trajectory segment corresponding to the current camera is predicted to have a world coordinate position after a preset buffer time, recorded as the predicted target position; A4. If the predicted target position is out of the field of view of the multi-camera tracking system, execute step A6; otherwise, execute step A5; A5. Determine a neighboring camera based on the predicted target location, adjust the orientation of the neighboring camera to align with the predicted target location, update the intrinsic and extrinsic parameters of the neighboring camera, and use the neighboring camera as the new current camera, returning to step A1; A6. Fuse all the target vehicle motion trajectory segments to obtain a total motion trajectory.
2. The dynamic target tracking method based on multiple cameras according to claim 1, characterized in that: The identification features also include license plate number, vehicle model and vehicle color; Step A1 includes: A101. Obtain the time-stamped frame images captured by the current camera in real time. A102 uses the YOLOv8 algorithm to perform target detection on the frame image to obtain the identification box of the target vehicle to determine the pixel position coordinates of the target vehicle; A103 uses optical character recognition technology to identify the image within the identification frame to obtain the target vehicle's license plate number; A104. Use a deep learning classifier to classify the image in the identification frame to obtain the model and color of the target vehicle.
3. The dynamic target tracking method based on multiple cameras according to claim 1, characterized in that: The internal parameters include an internal parameter matrix, and the external parameters include a translation matrix and a rotation matrix; Step A2 includes: A201. According to the real-time internal parameters of the current camera, the pixel position coordinates are converted into camera coordinates in the camera coordinate system of the current camera; A202. According to the real-time external parameters of the current camera, the camera coordinates are converted into three-dimensional world coordinates in the world coordinate system; A203. Add the timestamp of the current frame image to the three-dimensional world coordinates to form four-dimensional space-time coordinates as the final world coordinates; A204. Generate a target vehicle motion trajectory segment with time information corresponding to the current camera based on the final world coordinates.
4. The method for dynamic target tracking based on multiple cameras according to claim 3, wherein: Step A3 includes: A301. According to the pixel position coordinates of the target vehicle and the pixel size of the frame image, it is determined whether the target vehicle enters the boundary area of the current camera's field of view; A302. If the target vehicle enters the boundary area of the current camera's field of view, the current moving speed of the target vehicle is calculated based on the target vehicle's motion trajectory segment corresponding to the current camera; A303. Calculate the predicted target position based on the preset buffer time and the current moving speed.
5. The dynamic target tracking method based on multiple cameras according to claim 4, characterized in that: Step A301 includes: The feature distance from the edge of the target vehicle in the field of view of the current camera is obtained according to the following formula: L=min(u',1-u',v',1-v'); Wherein, L is the feature distance from the edge, u' is the normalized value of the horizontal coordinate in the current pixel position coordinate of the target vehicle, and v' is the normalized value of the vertical coordinate in the current pixel position coordinate of the target vehicle; If the characteristic distance from the edge is less than or equal to a preset distance from the edge threshold, it is determined that the target vehicle has entered the boundary area of the field of view of the current camera; otherwise, it is determined that the target vehicle has not entered the boundary area of the field of view of the current camera.
6. The dynamic target tracking method based on multiple cameras according to claim 1, characterized in that: Step A5 includes: A501. Calculate the distance between the predicted target position and the field of view of all cameras other than the current camera, recorded as the deviation distance; A502. Select one of the other cameras whose deviation distance is less than a preset distance threshold as the adjacent camera; A503. Calculate the rotation angle adjustment and pitch angle adjustment required to align the optical center of the neighboring camera with the predicted target position based on the predicted target position, the current rotation angle of the neighboring camera, the current pitch angle of the neighboring camera, and the real-time external parameters of the neighboring camera; A504. Adjust the rotation angle and pitch angle of the adjacent camera according to the rotation angle adjustment amount and the pitch angle adjustment amount; A505. Obtain three consecutive frames of images captured by the adjacent camera after adjusting the rotation angle and pitch angle, input the three consecutive frames of images into a pre-trained deep learning calibration model, and obtain the external parameter correction value output by the deep learning calibration model; A506. Update the external parameters of the adjacent camera according to the external parameter correction amount; A507. Calculate the zoom factor based on the pixel size of the identification frame obtained by the neighboring camera to identify the target vehicle and the actual size of the target vehicle to adjust the focal length of the neighboring camera; A508. Update the internal parameters of the adjacent camera according to the adjusted focal length; A509. After completing the adjustment of the internal parameters and external parameters of the adjacent camera, use the adjacent camera as the new current camera and return to step A1.
7. The method for dynamic target tracking based on multiple cameras according to claim 6, characterized in that: Step A507 includes: B1 obtains the frame image captured by the adjacent camera, performs vehicle identification, and obtains at least one vehicle identification feature; the identification feature also includes the license plate number, vehicle model and vehicle color; B2. Calculate the similarity between each vehicle identified by the neighboring camera and the target vehicle based on the license plate number, vehicle model, and vehicle color of the target vehicle identified by the neighboring camera; B3 determines the target vehicle based on the similarity from each vehicle identified by the neighboring camera; B4 obtains the pixel size of the identification frame of the target vehicle determined, recorded as the first pixel size, and obtains the corresponding actual size according to the model of the target vehicle determined; B5. Calculate the ratio of the actual size to the first pixel size to obtain an actual conversion ratio, and calculate the zoom factor that can make the actual conversion ratio equal to the target conversion ratio by combining the preset target conversion ratio and the current focal length of the adjacent camera; B6. Adjust the focal length of the adjacent camera according to the zoom factor.
8. The dynamic target tracking method based on multiple cameras according to claim 7, characterized in that: Step B2 includes: B201. Calculate the license plate number consistency parameter according to the following formula: ; in, is the license plate number consistency parameter, is the code of the license plate number of the target vehicle acquired by the current camera, is the code of the vehicle license plate number obtained by the proximity camera, for and The coding distance between for The encoding length, for The encoding length; B202. Convert the color space of the target vehicle obtained by the current camera and the vehicle color of the vehicle obtained by the adjacent camera into the HSV color space, and calculate the color distance using the following formula: ; in, is the color distance, 、 、 are the H, S, and V channel values of the vehicle color of the target vehicle acquired by the current camera, 、 、 are the H, S, and V channel values of the vehicle color acquired by the neighboring camera, is the preset reference threshold; B203. Using a hierarchical scoring rule based on a vehicle classification tree design, determine the category similarity score Ts between the target vehicle model captured by the current camera and the vehicle model captured by the adjacent camera; B204. Calculate the similarity between the vehicle identified by the neighboring camera and the target vehicle according to the following formula: ; in, is the similarity, 、 、 is the preset weight coefficient.
9. An electronic device, characterized in that: It includes a processor and a memory, the memory stores a computer program executable by the processor, and when the processor executes the computer program, it runs the steps in the dynamic target tracking method based on multiple cameras as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multi-camera-based dynamic target tracking method according to any one of claims 1 to 8 are executed.
Citation Information
Patent Citations
Vehicle real-time tracking method and computer readable storage medium
CN112633282A