Target identification tracking method and system based on multi-source fusion imaging
By adopting a multi-source fusion imaging method in the target recognition and tracking technology, multi-modal data processing and feature fusion is used to use heterogeneous sensor arrays and cross-attention mechanisms to perform multi-modal data processing and feature fusion, and building a cascading target recognition model for target recognition and tracking, the technical difficulties of target recognition and tracking in complex environments are solved, and efficient target tracking and trajectory repair are achieved.
Patent Information
- Application Number
- CN202510644767.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-20
AI Technical Summary
Existing target identification and tracking technologies face problems in complex environments such as insufficient spatial and temporal alignment accuracy of multi-source data, low cross-modal feature correlation efficiency, and lack of trajectory repair mechanism after target loss, making it difficult to achieve continuous and stable tracking of multiple targets in complex scenarios in key areas such as intelligent monitoring.
The object recognition and tracking method based on multi-source fusion imaging is adopted, and multi-modal environmental monitoring data is obtained through heterogeneous sensor arrays, spatial registration and image enhancement are performed, cross-attention mechanism is introduced for multi-modal feature fusion, and a cascading object recognition model is constructed for target recognition and tracking. At the same time, tracking tracking track disconnection repair is used to ensure target identification accuracy and tracking reliability.
It improves the accuracy of target recognition and tracking reliability in complex environments, enhances the ability to continuously and stably track targets in monitoring scenarios, and solves the problems of insufficient spatial and temporal alignment accuracy of multi-source data and low cross-modal feature association efficiency in traditional methods.
Smart Images

Figure CN120182323A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target recognition and tracking, and particularly to a target recognition and tracking method and system based on multi-source fusion imaging. Background Art
[0002] In the technical field of target recognition and tracking, traditional methods usually rely on a single sensor to collect data, facing significant technical bottlenecks in complex environments. Visible light imaging is vulnerable to changes in illumination, bad weather, and target occlusion. Although infrared thermal imaging has the ability to perceive at night, its spatial resolution is limited. Millimeter-wave radar has inherent defects in target contour recognition. Existing methods based on multi-modal data fusion have improved environmental adaptability to a certain extent, but there are still problems in practical applications, such as insufficient spatio-temporal alignment accuracy of multi-source data, low cross-modal feature association efficiency, and the lack of a trajectory repair mechanism after target loss. In addition, the phenomenon of trajectory interruption caused by environmental interference or sensor disconnection is common during target tracking. Traditional algorithms usually rely on linear prediction or trajectory interpolation for repair, and it is difficult to ensure the accuracy and continuity of trajectory association in unstructured dynamic scenarios. These technical defects seriously restrict the demand for continuous and stable tracking of multiple targets in complex scenarios in key fields such as intelligent monitoring, and there is an urgent need to achieve a technical breakthrough through multi-modal deep collaborative perception and an intelligent repair mechanism. Summary of the Invention
[0003] The present invention overcomes the defects of the prior art and provides a target recognition and tracking method and system based on multi-source fusion imaging.
[0004] To achieve the above object, the first aspect of the present invention provides a target recognition and tracking method based on multi-source fusion imaging, including: Performing environmental monitoring on a target scene to obtain multi-modal environmental monitoring data and preprocessing the data. Based on the preprocessed multi-modal environmental monitoring data, image enhancement is performed to obtain image-enhanced multi-modal environmental monitoring data; Performing multi-modal feature extraction on the image-enhanced multi-modal environmental monitoring data, and introducing a cross-attention mechanism to fuse the extracted multi-modal features to construct cross-modal fusion features; Constructing a cascaded target recognition model, inputting the cross-modal fusion features to perform target recognition on the current monitoring scene, and obtaining scene target recognition information; Tracking the scene recognition targets in the current monitoring scene based on the cross-modal fusion features and the scene target recognition information to generate scene target tracking information and storing it in a preset trajectory memory pool; When a tracking target disconnection occurs during target tracking, obtaining a mismatched tracking trajectory and combining it with the trajectory memory pool to repair the tracking trajectory disconnection.
[0005] In this solution, the environmental monitoring of the target scenario to obtain multimodal environmental monitoring data and perform preprocessing specifically includes: Install a heterogeneous sensor array in the target scenario, and use the installed heterogeneous sensor array to perform environmental monitoring on the target scenario to obtain multimodal environmental monitoring data, where the heterogeneous sensor array includes a high-resolution imaging camera, an infrared thermal imaging camera, and a millimeter-wave radar; Based on the obtained multimodal environmental monitoring data, perform spatial registration on the data obtained from different monitoring sources. Introduce the ORB feature extraction algorithm, extract the scale pyramid features from the initial visible light image, and eliminate the illumination difference through gray value normalization; After completing the elimination of the illumination difference, extract the image region features of the initial visible light image after the illumination difference elimination, and perform threshold segmentation processing on the corresponding region of the initial infrared image through the extracted image region features; Extract feature points from the initial visible light image and the initial infrared image respectively, perform preliminary matching on the extracted feature points through the Hamming distance to generate a set of matching points. After completing the preliminary matching, use the RANSAC algorithm to iteratively screen the inlier set; Randomly select 4 pairs of feature points each time to calculate the affine transformation matrix, and calculate the reprojection error of all matching segments with respect to this affine transformation matrix. Select the points with a reprojection error less than the preset reprojection error and include them in the inlier set. After repeated iteration, retain the final affine transformation matrix and the corresponding inlier set to obtain the spatial registration result of the initial visible light image and the initial infrared image; Extract the radar point cloud data from the multimodal environmental monitoring data, convert the polar coordinates corresponding to the radar point cloud data to Cartesian coordinate system three-dimensional coordinates, and perform rigid body transformation projection onto the imaging plane of the visible light camera in combination with the external parameter calibration matrix of the sensor, and finally generate a two-dimensional point cloud heat map spatially aligned with the optical image; Combine the spatial registration result of the initial visible light image and the initial infrared image and the two-dimensional point cloud heat map to obtain the preprocessed multimodal environmental monitoring data.
[0006] In this solution, the image enhancement based on the preprocessed multimodal environmental monitoring data to obtain image-enhanced multimodal environmental monitoring data specifically includes: Obtain the preprocessed multimodal environmental monitoring data, input it into the HSI model to convert the original RGB image to the HSI color space, and obtain the hue component, saturation component, and brightness component through the HSI color space; Introduce the Retinex algorithm to decompose the brightness component into a reflection component and an illumination component, perform digital domain processing on the reflection component to enhance the dark area of the image, perform non-linear adaptive adjustment on the separated saturation component, and enhance the color concentration through the dynamic range expansion function; The Canny edge detection operator is used to extract the structural features of the whole image and form an edge map. The edge map is pixel-level weighted fused with the enhanced brightness component. After processing each component, the optimized H, S, and I channels are re-fused, and the RGB image is reconstructed through inverse color space conversion to obtain the initial enhanced image; The Markov algorithm is introduced to reprocess the initial enhanced image. The initial enhanced image is imported into the convolutional autoencoder network for encoding and mapped to the latent space to generate the initial latent variable as the starting state of the Markov chain; The state distribution is defined through the generated initial latent variable and the state transition probability is calculated. Candidate latent variables are generated according to the state transition probability, and the Metropolis criterion is used to judge whether to accept the target candidate latent variable; Iterative optimization is performed until the stopping criterion is met, and then the final latent variable is output. The final latent variable is input into the pre-trained U-Net decoder network, and the high-resolution image is gradually reconstructed through the deconvolution layer and skip connection to obtain the image-enhanced multimodal environmental monitoring data.
[0007] In this solution, for the multimodal feature extraction of the image-enhanced multimodal environmental monitoring data, a cross-attention mechanism is introduced to fuse the extracted multimodal features to construct cross-modal fusion features, which specifically includes: The image-enhanced multimodal environmental monitoring data is obtained and input into a multi-channel feature extraction network for multimodal feature extraction. The multi-channel feature extraction network includes a first feature extraction channel, a second feature extraction channel, and a third feature extraction channel; The visible light image is imported into the first feature extraction channel to extract multi-scale context features based on the hybrid dilated convolutional layer. The context information under different receptive fields is captured according to the parallel convolution branches, and then compressed into a multi-dimensional global vector through global average pooling to obtain the first feature extraction information; The infrared image is imported into the second feature extraction channel for spatial feature extraction. The extracted spatial features are input into the SE attention module to perform global average pooling to generate a channel description vector. The channel weight matrix is obtained by using the channel description vector and multiplied with the original feature map to extract the local thermal feature vector to obtain the second feature extraction information; The radar point cloud data is imported into the third feature extraction channel to extract three-dimensional spatial features. k key points are selected as the initial point set through farthest point sampling, and the geometric contour and motion trajectory of the target object are captured by combining the local aggregation mechanism to generate the third feature extraction information; The cross-attention mechanism is introduced to perform feature cross-modal association based on the first feature extraction information, the second feature extraction information, and the third feature extraction information. Taking the visible light image features as the query, and the infrared image features and the radar point cloud features as the key values, the spatial correlation weights are calculated through multi-head attention to generate cross-modal fusion features.
[0008] In this solution, to construct the cascaded target recognition model, the cross-modal fusion features are input to perform target recognition on the current monitoring scene to obtain the scene target recognition information, which specifically includes: Construct a cascaded target recognition model, which is constructed with Bi-LSTM and a multi-scale detection head as the framework. Obtain the cross-modal fusion features and input them into the cascaded target recognition model to analyze the current scene; Perform temporal context modeling on the input cross-modal fusion features using the Bi-LSTM network. Based on the stacked two-layer LSTM units, capture the spatio-temporal context features of the objects in the target scene respectively by forward and backward passes, and output spatio-temporal enhanced features; Input the spatio-temporal enhanced features into the multi-scale detection head for target parsing. Execute multi-level feature abstraction through the lightweight GhostNet module to obtain multiple groups of low-rank feature maps, and use the cross-stage local connection module to extract local detail features; Perform depthwise separable convolution compression on the obtained multiple groups of low-rank feature maps to obtain the channel dimension, perform channel concatenation with the local detail features to obtain cascaded fusion features, introduce spatial pyramid pooling and input the cascaded fusion features, and extract cross-receptive field features through parallel multi-scale max pooling and splice them with the original feature map to obtain a three-level feature pyramid; Perform scene target recognition on the target scene based on the obtained three-level feature pyramid to generate several candidate recognition frames, and perform non-maximum suppression and mask threshold filtering on the candidate recognition frames to obtain the scene target recognition information.
[0009] In this solution, to track the scene recognition targets in the current monitoring scene based on the cross-modal fusion features and the scene target recognition information to generate scene target tracking information and store it in the preset trajectory memory pool, it specifically includes: Obtain the cross-modal fusion features and the scene target recognition information, extract the motion state features of the scene recognition targets based on the cross-modal fusion features, and import the extracted running state features into the extended Kalman filter to construct a state vector; Predict the position distribution of the scene recognition targets in the next frame of image based on the constructed state vector to obtain target prediction frames, calculate the Mahalanobis distance between the target prediction frames and the target detection frames in combination with the scene target recognition information, and generate a motion cost matrix; Obtain the appearance state features of the scene recognition target through the cross-modal fusion features, calculate the cosine similarity with the target appearance embedding in the adjacent frame image, generate an appearance cost matrix and perform weighted fusion with the motion cost matrix to generate a comprehensive cost matrix as the input of the Hungarian algorithm for solving the optimal matching pairs, and generate scene target tracking information according to the obtained optimal matching pairs and store it in the preset trajectory memory pool.
[0010] In this solution, when a disconnection situation of the tracking target occurs during target tracking, obtain the mismatched tracking trajectory, and combine the trajectory memory pool to repair the disconnection of the tracking trajectory. Specifically, it includes: Obtain the mismatched tracking trajectory, extract the last frame image where the trajectory mismatch occurs based on the mismatched tracking trajectory, define it as the disconnection repair start frame, and extract the frame timing features and multi-modal features of the disconnection repair start frame to obtain the disconnection repair frame feature information; Obtain the trajectory memory pool, extract the scene target tracking trajectory after the frame timing features of the disconnection repair frame based on the disconnection frame feature information, define it as the initial repair tracking trajectory, and extract the motion state attributes of the disconnection repair frame through the multi-modal features of the disconnection repair start frame; Extract the motion state attributes of each initial repair tracking trajectory, calculate the deviation from the motion state attributes of the disconnection repair frame, and select the initial repair tracking trajectories with deviations within the preset range as candidate repair tracking trajectories; Extract the multi-modal features of each candidate repair trajectory, calculate the similarity with the disconnection repair frame feature information to obtain several candidate repair tracking trajectories with a similarity greater than the preset similarity threshold, and select the candidate repair tracking trajectory with the largest similarity value as the final repair tracking trajectory to connect with the mismatched tracking trajectory.
[0011] In this solution, when a disconnection situation of the tracking target occurs during target tracking, obtain the mismatched tracking trajectory, and combine the trajectory memory pool to repair the disconnection of the tracking trajectory. It also includes: If the similarity values between the multi-modal features of all candidate repair tracking trajectories and the disconnection repair frame feature information are less than the preset similarity threshold, obtain the spatial coordinate data of the scene tracking target in the disconnection repair frame as the origin to establish a polar coordinate system, and generate the predicted trajectories for the next k frames based on the motion state attributes of the disconnection repair frame, and calibrate the predicted trajectory coordinates in the established polar coordinate system; Extract the spatial coordinate data of the scene tracking target in the start frame of each candidate repair tracking trajectory according to the multi-modal features of the scene target associated with the start frame of each candidate repair tracking trajectory, and calibrate it in the established polar coordinate system, which is defined as the point to be determined; Segment the predicted trajectory in polar coordinates according to the temporal characteristics, and set the confidence levels of different predicted trajectory segments based on the order of the temporal characteristics. Calculate the Manhattan distances between each point to be determined and the predicted trajectory calibrated to the polar coordinate system at different times in the constructed polar coordinate system. Select the predicted trajectory segment with the smallest Manhattan distance as the target predicted trajectory segment corresponding to the point to be determined according to the calculated Manhattan distances, and perform weighted fusion in combination with the confidence level corresponding to the target predicted trajectory segment to obtain the association degree between each point to be determined and the predicted trajectory. Judge the association degree between each point to be determined and the preset trajectory and the preset association degree threshold. If there is no point to be determined greater than the preset association threshold, mark the target mismatched tracking trajectory as an irreparable trajectory. If there is a point to be determined greater than the preset association threshold, select the point to be determined with the largest association degree as the target determination point, define the candidate repair tracking trajectory corresponding to the target determination point as the target repair tracking trajectory, and connect it with the target mismatched tracking trajectory to complete the disconnection repair of the tracking trajectory.
[0012] The second aspect of the present invention provides a target recognition and tracking system based on multi-source fusion imaging. The system includes: a memory and a processor. The memory contains a program for the target recognition and tracking method based on multi-source fusion imaging. When the program for the target recognition and tracking method based on multi-source fusion imaging is executed by the processor, the following steps are implemented: Conduct environmental monitoring on the target scene to obtain multi-modal environmental monitoring data and perform preprocessing. Based on the preprocessed multi-modal environmental monitoring data, perform image enhancement to obtain image-enhanced multi-modal environmental monitoring data. Extract multi-modal features from the image-enhanced multi-modal environmental monitoring data, and introduce a cross-attention mechanism to fuse the extracted multi-modal features to construct cross-modal fusion features. Construct a cascaded target recognition model, input the cross-modal fusion features to perform target recognition on the current monitoring scene, and obtain scene target recognition information. Based on the cross-modal fusion features and the scene target recognition information, track the scene recognition targets in the current monitoring scene to generate scene target tracking information and store it in a preset trajectory memory pool. When a tracking target disconnection occurs during target tracking, obtain the mismatched tracking trajectory and combine it with the trajectory memory pool to repair the disconnection of the tracking trajectory.
[0013] The present invention discloses a target recognition and tracking method and system based on multi-source fusion imaging, including: acquiring multi-modal environmental monitoring data and performing preprocessing, performing image enhancement based on the preprocessed multi-modal environmental monitoring data to obtain image-enhanced multi-modal environmental monitoring data; extracting multi-modal features from the image-enhanced multi-modal environmental monitoring data, introducing a cross-attention mechanism to fuse the extracted multi-modal features to construct cross-modal fusion features; constructing a cascaded target recognition model, inputting the cross-modal fusion features to recognize targets in the current monitoring scene; tracking the scene recognition targets in the current monitoring scene to generate scene target tracking information and storing it in a preset trajectory memory pool; when a tracking target disconnection occurs during target tracking, acquiring a mismatched tracking trajectory and combining it with the trajectory memory pool to repair the tracking trajectory disconnection. Thereby improving the recognition accuracy and tracking reliability of targets in the monitoring scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments or exemplary examples of the present invention, the following will briefly introduce the drawings required for use in the embodiments or exemplary descriptions. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the drawings shown.
[0015] Figure 1 It is a flowchart of a target recognition and tracking method based on multi-source fusion imaging provided by an embodiment of the present invention; Figure 2 It is a flowchart of a tracking disconnection repair method for target recognition and tracking provided by an embodiment of the present invention; Figure 3 It is a block diagram of a target recognition and tracking system based on multi-source fusion imaging provided by an embodiment of the present invention; The realization, functional features and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] In order to be able to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0017] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0018] Figure 1Flowchart of a target recognition and tracking method based on multi-source fusion imaging provided by an embodiment of the present invention; As Figure 1 shown, the present invention provides a flowchart of a target recognition and tracking method based on multi-source fusion imaging, including: S102, perform environmental monitoring on the target scene to obtain multi-modal environmental monitoring data and perform preprocessing, and perform image enhancement based on the preprocessed multi-modal environmental monitoring data to obtain image-enhanced multi-modal environmental monitoring data; S104, perform multi-modal feature extraction on the image-enhanced multi-modal environmental monitoring data, and introduce a cross-attention mechanism to fuse the extracted multi-modal features to construct cross-modal fusion features; S106, construct a cascaded target recognition model, input the cross-modal fusion features to perform target recognition on the current monitoring scene, and obtain scene target recognition information; S108, track the scene recognition targets in the current monitoring scene based on the cross-modal fusion features and the scene target recognition information to generate scene target tracking information and store it in a preset trajectory memory pool; S110, when a tracking target disconnection occurs during target tracking, obtain the mismatched tracking trajectory, and combine the trajectory memory pool to repair the tracking trajectory disconnection.
[0019] Further, in a preferred embodiment of the present invention, the performing environmental monitoring on the target scene to obtain multi-modal environmental monitoring data and perform preprocessing specifically includes: Install a heterogeneous sensor array in the target scene, and perform environmental monitoring on the target scene through the installed heterogeneous sensor array to obtain multi-modal environmental monitoring data, where the heterogeneous sensor array includes a high-resolution imaging camera, an infrared thermal imaging camera, and a millimeter-wave radar; Perform spatial registration on the data obtained from different monitoring sources based on the obtained multi-modal environmental monitoring data, introduce the ORB feature extraction algorithm, extract scale pyramid features from the initial visible light image, and eliminate the illumination difference through gray value normalization; After completing the illumination difference elimination, extract the image region features of the initial visible light image after illumination difference elimination, and perform threshold segmentation processing on the corresponding region of the initial infrared image through the extracted image region features; Extract feature points from the initial visible light image and the initial infrared image respectively, perform preliminary matching on the extracted feature points through the Hamming distance to generate a set of matching points, and after completing the preliminary matching, use the RANSAC algorithm to iteratively screen the inlier set; Randomly select 4 pairs of feature points each time to calculate the affine transformation matrix, and calculate the reprojection error of all matching segment pairs under this affine transformation matrix. Select the points with reprojection error less than the preset value and include them in the inlier set. After repeated iterations, retain the final affine transformation matrix and the corresponding inlier set to obtain the spatial registration result of the initial visible light image and the initial infrared image; Extract radar point cloud data from multi-modal environmental monitoring data, convert the polar coordinates corresponding to the radar point cloud data into three-dimensional Cartesian coordinates, and project them onto the imaging plane of the visible light camera through rigid body transformation in combination with the external parameter calibration matrix of the sensor, and finally generate a two-dimensional point cloud thermal map spatially aligned with the optical image; Combine the spatial registration result of the initial visible light image and the initial infrared image and the two-dimensional point cloud thermal map to obtain the preprocessed multi-modal environmental monitoring data.
[0020] It should be noted that a heterogeneous sensor array composed of a high-resolution imaging camera, an infrared thermal imaging camera, and a millimeter-wave radar is deployed in the target scene. Through the collaborative work of multiple sensors, multi-modal environmental monitoring data such as high-resolution image information, infrared thermal radiation distribution, and millimeter-wave radar point cloud of the scene are collected in real time. Based on the acquired multi-source data, first perform cross-modal spatial registration on the visible light and infrared images: for the visible light image, use the ORB feature extraction algorithm to construct a scale pyramid, detect directional FAST corner points at different resolution levels and calculate rotation-invariant BRIEF binary descriptors, and at the same time eliminate the influence of illumination intensity differences on feature stability through gray value normalization; for the infrared image, perform adaptive threshold segmentation based on the region features (such as edge gradient distribution) of the visible light image with illumination differences eliminated, and extract the thermal radiation contour corresponding to the visible light target area. Subsequently, extract ORB feature points from the two-modal images respectively, use the Hamming distance to measure the similarity of the binary descriptors to generate a preliminary matching point set, and then iteratively optimize the registration parameters through the RANSAC algorithm: randomly select 4 pairs of feature points each time to calculate the affine transformation matrix, calculate the reprojection error of all matching point pairs based on this matrix (i.e., the Euclidean distance between the projected coordinates and the measured coordinates), screen the points with error less than the preset threshold and include them in the inlier set, and after multiple rounds of iteration, select the affine transformation matrix with the largest number of inliers and the smallest average error as the final registration parameter to achieve high-precision spatial alignment of the visible light and infrared images.
[0021] Furthermore, for the millimeter-wave radar point cloud data, its original polar coordinates (distance, azimuth angle, elevation angle) are converted into three-dimensional coordinates in the Cartesian coordinate system, and a rigid body transformation is performed in combination with the external sensor calibration matrix (the rotation matrix and translation vector pre-acquired through a multi-view calibration board), and the three-dimensional point cloud is projected onto the imaging plane coordinate system of the visible light camera. During the projection process, a perspective projection model is used to calculate the two-dimensional coordinates of each radar point on the image plane, and a gray value mapping is generated according to the point cloud reflection intensity to form a two-dimensional point cloud heat map that is spatially aligned with the visible light image. Finally, the visible light-infrared registration image and the radar point cloud heat map are integrated to construct a spatio-temporally consistent multi-modal environmental monitoring data set, providing a fusion data basis with cross-modal spatial alignment characteristics for subsequent target detection and tracking.
[0022] Further, in a preferred embodiment of the present invention, image enhancement is performed on the preprocessed multi-modal environmental monitoring data to obtain image-enhanced multi-modal environmental monitoring data, which specifically includes: Obtain the preprocessed multi-modal environmental monitoring data, input it into the HSI model to convert the original RGB image to the HSI color space, and obtain the hue component, saturation component, and brightness component through the HSI color space; Introduce the Retinex algorithm to decompose the brightness component into a reflection component and an illumination component, perform digital domain processing on the reflection component to enhance the dark area of the image, perform non-linear adaptive adjustment on the separated saturation component, and enhance the color concentration through a dynamic range expansion function; Use the Canny edge detection operator to extract the full-image structure features and form an edge map, perform pixel-level weighted fusion of the edge map and the enhanced brightness component. After processing each component, re-fuse the optimized H, S, and I channels, and reconstruct the RGB image through inverse color space conversion to obtain an initial enhanced image; Introduce the Markov algorithm to reprocess the initial enhanced image, import the initial enhanced image into the convolutional autoencoder network for encoding and mapping it to the latent space, and generate an initial latent variable as the starting state of the Markov chain; Define the state distribution through the generated initial latent variable and calculate the state transition probability. Generate candidate latent variables according to the state transition probability, and use the Metropolis criterion to determine whether to accept the target candidate latent variable; Perform iterative optimization until the stop criterion is met, then output the final latent variable. Input the final latent variable into the pre-trained U-Net decoder network, and gradually reconstruct a high-resolution image through deconvolution layers and skip connections to obtain image-enhanced multi-modal environmental monitoring data.
[0023] It should be noted that first, the visible light RGB image is input into the HSI (Hue, Saturation, Intensity) color space conversion model, and the hue component, saturation component representing color attributes, and the intensity component representing illumination intensity are separated through a non-linear formula. For the intensity component, the Retinex algorithm is introduced for illumination correction: based on multi-scale Gaussian filtering, the intensity layer is decomposed into a reflection component (representing the inherent brightness of the target) and an illumination component (representing environmental illumination interference). A logarithmic transformation is applied to the reflection component to expand the dynamic range of the dark area, and at the same time, adaptive gamma correction is used to further suppress uneven illumination; for the saturation component, a piecewise non-linear stretching function is adopted, with exponential enhancement in the low saturation area to improve color discrimination and linear mapping in the high saturation area to avoid over-saturation. Subsequently, the Canny edge detection operator is used to extract the full-image structure features and form an edge map: through Gaussian filtering to smooth noise, non-maximum suppression, and double-threshold hysteresis processing, a binary edge map with complete edge connection is generated. The edge map is pixel-level weighted fused with the enhanced intensity component, and the fusion weight is dynamically adjusted according to the edge density to enhance texture details while retaining illumination consistency. After independent optimization of the H, S, and I components, the color space is reconstructed through the reverse HSI to RGB conversion formula to generate an initial enhanced image.
[0024] Furthermore, a Markov algorithm is introduced to improve the image quality. The initial enhanced image is input into a pre-trained convolutional autoencoder. The encoder network extracts multi-scale features and compresses them into the latent space to generate an initial latent variable as the initial state of the Markov chain. The state transition process is defined as a random walk in the latent space. The acceptance probability of the candidate latent variable is calculated, and whether to accept the candidate state is determined according to the Metropolis criterion (based on the energy function difference and probability threshold). After multiple iterative optimizations, the latent variable with the minimum energy function is selected and input into the U-Net decoder network. Through the deconvolution layer, it is gradually upsampled and combined with skip connections to fuse the low-level features in the encoding stage to reconstruct a high-resolution image. This process enhances the detail sharpness while suppressing noise through the probabilistic exploration and deterministic decoding in the latent space, and finally outputs multi-modal environmental monitoring data with balanced colors and clear edges, providing high-quality cross-modal input for subsequent target recognition and tracking.
[0025] It is worth mentioning that the energy function is expressed as:
[0026] where is to map the latent variable to the image space to generate an enhanced image, is the original image, is the regularization coefficient, is the image gradient.
[0027] Further, in a preferred embodiment of the present invention, for the multi-modal feature extraction of the image-enhanced multi-modal environmental monitoring data, a cross-attention mechanism is introduced to fuse the extracted multi-modal features to construct cross-modal fusion features, which specifically includes: Obtain the image-enhanced multi-modal environmental monitoring data, and input the image-enhanced multi-modal environmental monitoring data into a multi-channel feature extraction network for multi-modal feature extraction. The multi-channel feature extraction network includes a first feature extraction channel, a second feature extraction channel, and a third feature extraction channel; Import the visible light image into the first feature extraction channel to extract multi-scale context features based on the hybrid dilated convolutional layer, capture context information under different receptive fields according to the parallel convolution branches, and then compress it into a multi-dimensional global vector through global average pooling to obtain the first feature extraction information; Import the infrared image into the second feature extraction channel for spatial feature extraction, input the extracted spatial features into the SE attention module to perform global average pooling to generate a channel description vector, use the channel description vector to obtain a channel weight matrix and multiply it with the original feature map to extract a local thermal feature vector, and obtain the second feature extraction information; Import the radar point cloud data into the third feature extraction channel to extract three-dimensional spatial features, select k key points as the initial point set through farthest point sampling, and capture the geometric contour and motion trajectory of the target object in combination with the local aggregation mechanism to generate the third feature extraction information; Introduce a cross-attention mechanism to perform feature cross-modal association based on the first feature extraction information, the second feature extraction information, and the third feature extraction information. Use the visible light image feature as the query, and the infrared image feature and the radar point cloud feature as the key values, and calculate the spatial correlation weight through multi-head attention to generate cross-modal fusion features.
[0028] It should be noted that after obtaining the multi-modal environmental monitoring data after image enhancement, it is input into a multi-channel feature extraction network composed of a first feature extraction channel (visible light), a second feature extraction channel (infrared), and a third feature extraction channel (radar point cloud) to extract the deep semantic features of each modality respectively. After the visible light image is input into the first feature extraction channel, convolution kernels with different dilation rates (such as 1, 3, and 5) are fused through a hybrid dilated convolution layer to capture multi-scale context information while maintaining the resolution. Among them, the parallel convolution branches use 1×1, 3×3, and 5×5 convolution kernels to synchronously extract local detail and global structure features. Subsequently, the feature map is compressed into a multi-dimensional vector representing global semantics through global average pooling, generating the first feature extraction information containing target contour, texture, and illumination information. For the infrared image, first, the spatial convolution network (SCN) is used to extract the thermal radiation distribution features in the spatial dimension, and then it is input into the SE attention module: global average pooling is performed on the feature map to generate a channel description vector, the channel weight matrix is calculated through a fully connected layer and the Sigmoid activation function, and it is multiplied with the original feature map channel by channel to strengthen the thermal response of the target area and suppress background noise, finally extracting the second feature extraction information focusing on local high-temperature or low-temperature targets. For the millimeter-wave radar point cloud data, farthest point sampling is used to select k key points from the original point cloud as the initial point set, and the local neighborhood features of the point cloud are constructed based on the Set Abstraction layer in the local aggregation PointNet++: the coordinate normalization is performed on the neighborhood point set (points within a radius r) of each key point, and the coordinates, reflection intensity, and velocity information of the neighborhood points are aggregated through a multi-layer perceptron (MLP) to generate the third feature extraction information describing the three-dimensional geometric contour and motion trajectory of the target object. Finally, a cross-attention mechanism is introduced to achieve multi-modal feature fusion: using the visible light feature vector as the query, the infrared thermal feature vector and the radar point cloud feature vector as the key and value, the spatial correlation weights between the visible light feature and the features of the other modalities are calculated through multi-head self-attention, and the infrared and radar features are weighted and aggregated after Softmax normalization to generate a cross-modal joint feature that fuses the target appearance (visible light), thermal radiation (infrared), and three-dimensional motion (radar). It realizes the adaptive enhancement of the contribution of key modalities in complex scenarios, such as increasing the weights of infrared and radar features in haze weather, ensuring the robustness and environmental adaptability of multi-modal fusion.
[0029] Furthermore, in a preferred embodiment of the present invention, when constructing the cascaded target recognition model, the cross-modal fusion feature is input to perform target recognition on the current monitoring scenario to obtain the scene target recognition information, which specifically includes: Construct a cascaded target recognition model, which is constructed with a Bi-LSTM and a multi-scale detection head as the framework, obtain the cross-modal fusion feature and input it into the cascaded target recognition model to analyze the current scene; Perform temporal context modeling using a Bi-LSTM network based on the input cross-modal fusion features. Capture the spatio-temporal context features of objects in the target scene by forward and backward passes based on two stacked LSTM units, and output spatio-temporal enhanced features. Input the spatio-temporal enhanced features into a multi-scale detection head for object parsing. Perform multi-level feature abstraction through a lightweight GhostNet module to obtain multiple groups of low-rank feature maps, and use a cross-stage local connection module to extract local detail features. Perform depthwise separable convolution compression on the obtained multiple groups of low-rank feature maps to obtain the channel dimension, concatenate them with the local detail features at the channel level to obtain concatenated fusion features, introduce spatial pyramid pooling and input the concatenated fusion features, extract cross-receptive field features through parallel multi-scale max pooling and splice them with the original feature map to obtain a three-level feature pyramid. Perform scene object recognition on the target scene based on the obtained three-level feature pyramid to generate a number of candidate recognition frames, and perform non-maximum suppression and mask threshold filtering on the candidate recognition frames to obtain scene object recognition information.
[0030] It should be noted that a cascaded object recognition model is constructed based on the design of a model architecture with a bidirectional long short-term memory network (Bi-LSTM) and a multi-scale detection head framework, and the cross-modal fusion features are input into the model for scene analysis. The Bi-LSTM network performs forward and backward propagation respectively through two stacked LSTM units to perform temporal modeling on the input feature sequence: the forward LSTM captures the motion trajectory features of the target object from the initial state to the current frame, and the backward LSTM learns the context dependence reversely. The cell state is dynamically updated through a gating mechanism (input gate, forget gate, output gate) to fuse historical and future information in the spatio-temporal dimension, generating a spatio-temporal enhanced feature vector. This feature vector encodes the joint representation of the dynamic behavior patterns (such as motion speed, direction change) and static attributes (such as shape, size) of objects in the scene.
[0031] Subsequently, the spatio-temporal enhanced features are input into a multi-scale detection head for object parsing, and multiple groups of low-rank feature maps are generated through a lightweight GhostNet module. GhostNet adopts a linear transformation and channel splitting strategy to decompose the original feature map into backbone features and "ghost" features. The backbone features are extracted through conventional convolutions, and the ghost features are generated through inexpensive linear operations (such as per-channel convolutions). After splicing the two, a multi-level feature abstraction is formed, which retains the multi-scale characteristics of the object while reducing the computational cost. Then, the feature map is processed in stages through a Cross-Stage Partial Connection (CSP) module: the feature map is divided into two parts. One part extracts local detailed features (such as edges and corners) through depthwise convolutions, and the other part is directly passed and fused with the result of the previous part to suppress gradient redundancy and enhance the ability to express details. For the multiple groups of low-rank feature maps output by GhostNet, depthwise separable convolutions are used to compress the channel dimension. By separating the feature learning of the spatial and channel dimensions through per-channel convolutions and pointwise convolutions, the number of parameters is reduced and the computational efficiency is improved. The compressed features are concatenated with the local detailed features extracted by the CSP module to form cascaded fusion features. Subsequently, spatial pyramid pooling is introduced, and the fusion features are input into parallel multi-scale max pooling layers (such as 4×4, 8×8, 16×16 grids) to extract global context information under different receptive fields. The multi-scale pooling results are concatenated with the original feature map along the channel dimension to construct a three-level feature pyramid containing fine-grained details, medium-scale structures, and macroscopic semantics. Finally, based on the three-level feature pyramid, an anchor box mechanism is used to generate candidate recognition boxes covering different scales and aspect ratios, and the class probability, bounding box offset, and mask confidence of each anchor box are predicted through convolutional layers. Non-maximum suppression is applied to candidate boxes with a high overlap rate, and they are sorted based on the intersection over union threshold and class scores to eliminate redundant detection results. At the same time, low-confidence predictions are removed by combining mask threshold filtering (confidence > 0.7). Finally, the position, class, and motion state information of the scene object are output. This enhances the robustness of subsequent dynamic object tracking. The multi-scale detection head balances computational efficiency and recognition accuracy, improving the real-time object perception ability and accuracy of multi-modal data in complex environments.
[0032] Further, in a preferred embodiment of the present invention, tracking the scene recognition target in the current monitoring scene based on the cross-modal fusion feature and the scene object recognition information to generate scene object tracking information and storing it in a preset trajectory memory pool specifically includes: Obtain the cross-modal fusion feature and the scene object recognition information, extract the motion state feature of the scene recognition target based on the cross-modal fusion feature, and import the extracted running state feature into an extended Kalman filter to construct a state vector; Predict the position distribution of the scene recognition target in the next frame of image based on the constructed state vector to obtain the target prediction box, calculate the Mahalanobis distance between the target prediction box and the target detection box in combination with the scene target recognition information, and generate a motion cost matrix; Obtain the appearance state features of the scene recognition target through the cross-modal fusion features, calculate the cosine similarity with the target appearance embedding in the adjacent frame image, generate an appearance cost matrix, and perform weighted fusion with the motion cost matrix to generate a comprehensive cost matrix as the input of the Hungarian algorithm for solving the optimal matching pair, and generate scene target tracking information according to the obtained optimal matching pair and store it in the preset trajectory memory pool.
[0033] It should be noted that first, the motion state features of the target (including dynamic parameters such as position, velocity, and acceleration) are extracted based on the cross-modal fusion features and used as the state vector to input the Extended Kalman Filter (EKF). The EKF predicts the position distribution of the target in the next frame through a non-linear state transition model, calculates the mean and covariance of the predicted state in combination with the process noise covariance matrix (characterizing the uncertainty of the motion model), and generates the probability distribution of the target prediction box. At the same time, the Mahalanobis distance between the prediction box and the detection box is calculated using the scene target recognition information of the current frame, and the Euclidean distance is normalized through the covariance matrix to quantify the matching degree between the two in the motion state space, and a motion cost matrix is constructed to characterize the motion consistency of the target in time series. To enhance the matching robustness, the appearance state features of the target (such as color histogram, texture descriptor, or depth embedding vector) are further extracted from the cross-modal fusion features, and the cosine similarity is calculated with the appearance embedding of the target in the adjacent frame to generate an appearance cost matrix reflecting the appearance consistency. The motion cost matrix (weight coefficient α) and the appearance cost matrix (weight coefficient β) are linearly weighted and fused to generate a comprehensive cost matrix (total cost = α × motion cost + β × appearance cost), where α and β are adaptively adjusted according to the scene dynamic characteristics (such as increasing α in high-speed motion scenes and increasing β in scenes with frequent occlusions). This comprehensive cost matrix is used as the input of the Hungarian algorithm. By constructing a bipartite graph and solving the minimum weight matching, the optimal detection box and trajectory association pair are obtained, and false matches and repeated tracking are excluded. Finally, the successfully matched trajectories are updated to the trajectory memory pool, recording the target ID, position history, and motion state. At the same time, new trajectories are initialized for the unmatched detection boxes, continuously outputting the continuous tracking information of the scene target to solve problems such as occlusion, deformation, and cross-modal feature drift.
[0034] Further, in a preferred embodiment of the present invention, when a disconnection occurs in the tracking target during target tracking, the mismatched tracking trajectory is obtained, and the tracking trajectory disconnection is repaired in combination with the trajectory memory pool, specifically including: Obtain the mismatch tracking trajectory, extract the last frame image where trajectory mismatch occurs based on the mismatch tracking trajectory, define it as the disconnection repair start frame, and extract the frame timing features and multi-modal features of the disconnection repair start frame to obtain disconnection repair frame feature information; Obtain the trajectory memory pool, extract the scene target tracking trajectory after the frame timing features of the disconnection repair frame based on the disconnection frame feature information, define it as the initial repair tracking trajectory, and extract the motion state attributes of the disconnection repair frame through the multi-modal features of the disconnection repair start frame; Extract the motion state attributes of each initial repair tracking trajectory, calculate the deviation from the motion state attributes of the disconnection repair frame, and select the initial repair tracking trajectories with deviations within the preset range as candidate repair tracking trajectories; Extract the multi-modal features of each candidate repair trajectory, calculate the similarity with the disconnection repair frame feature information to obtain several candidate repair tracking trajectories with a similarity greater than the preset similarity threshold, and select the candidate repair tracking trajectory with the largest similarity value as the final repair tracking trajectory to connect with the mismatch tracking trajectory.
[0035] It should be noted that after detecting the mismatch tracking trajectory, first locate the last frame before the disconnection of the trajectory (i.e., the disconnection repair start frame) based on the trajectory memory pool, extract the timing features (such as timestamp, motion speed history) and multi-modal features (including visible light appearance embedding, infrared thermal radiation distribution, radar point cloud geometric features) of this frame to form the disconnection repair frame feature information. Subsequently, retrieve all active tracking trajectories in the trajectory memory pool after the timestamp of the disconnection repair start frame as the initial repair tracking trajectories, and extract the motion state attributes (such as position, velocity vector, acceleration covariance matrix) of these trajectories at the disconnection moment. Extrapolate the predicted motion state of each initial repair tracking trajectory from the disconnection moment to the current moment through the Extended Kalman Filter (EKF), and calculate the deviation from the actual motion state of the disconnection repair frame: use the Mahalanobis distance to measure the statistical difference between the predicted state and the measured state, and screen out the trajectories with a Mahalanobis distance less than the preset threshold as candidate repair tracking trajectories. Further extract the multi-modal features of the candidate trajectories (such as visible light appearance embedding vector, infrared thermal radiation histogram, radar point cloud geometric descriptor), and calculate the cross-modal similarity with the disconnection repair frame features. Screen out the candidate trajectories with a similarity greater than the threshold, and select the one with the highest similarity as the final repair trajectory. Connect the mismatch trajectory and the repair trajectory through a trajectory interpolation algorithm (such as cubic spline interpolation), update the target ID, motion state, and multi-modal features in the trajectory memory pool, restore the tracking interruption caused by occlusion or detection failure, and ensure the spatio-temporal continuity of the target motion trajectory.
[0036] Furthermore, the specific calculation formula for using the Mahalanobis distance to measure the statistical difference between the predicted state and the measured state is as follows:
[0037] Among them, is the predicted state mean, is the disconnection frame observation mean, is the inverse matrix of the covariance matrix, is the residual vector between the predicted state and the observed state.
[0038] Figure 2 is the flowchart of a tracking disconnection repair method for target recognition and tracking provided by an embodiment of the present invention; As Figure 2 shown, the present invention provides a flowchart of a tracking disconnection repair method for target recognition and tracking, including: S202, if the similarity values between the multi-modal features of all candidate repair tracking trajectories and the disconnection repair frame feature information are all less than the preset similarity threshold, then obtain the spatial coordinate data of the scene tracking target in the disconnection repair frame as the origin to establish a polar coordinate system, and generate the predicted trajectories of the future k frames based on the motion state attributes of the disconnection repair frame, and perform predicted trajectory coordinate calibration in the established polar coordinate system; S204, extract the spatial coordinate data of the scene tracking target in the starting frame of each candidate repair tracking trajectory according to the multi-modal features of the scene target associated with the starting frame of each candidate repair tracking trajectory, and perform calibration in the established polar coordinate system, which is defined as the point to be determined; S206, segment the predicted trajectories in the polar coordinates according to the temporal features and set the confidence levels of different predicted trajectory segments based on the sequence of the temporal features, and calculate the Manhattan distances between each point to be determined and the predicted trajectories calibrated in the polar coordinate system at different time sequences; S208, select the predicted trajectory segment with the smallest Manhattan distance as the target predicted trajectory segment corresponding to the point to be determined according to the calculated Manhattan distances, and perform weighted fusion in combination with the confidence level corresponding to the target predicted trajectory segment to obtain the association degree between each point to be determined and the predicted trajectory; S210, judge the association degree between each point to be determined and the preset trajectory and the preset association threshold. If there is no point to be determined greater than the preset association threshold, then mark the target mismatched tracking trajectory as an irreparable trajectory; S212, if there is a point to be determined greater than the preset association threshold, then select the point to be determined with the largest association degree as the target determination point, define the candidate repair tracking trajectory corresponding to the target determination point as the target repair tracking trajectory, and connect it with the target mismatched tracking trajectory to complete the trajectory disconnection repair.
[0039] It should be noted that during the trajectory repair process, if the multi-modal feature similarity of all candidate repair tracking trajectories is lower than the preset threshold, an alternative repair strategy based on motion state and spatial relationship is activated. First, a polar coordinate system is constructed with the spatial coordinates of the target in the disconnection repair starting frame (coordinates in the Cartesian coordinate system) as the origin. The motion direction of the target (direction of the velocity vector) is defined as the polar axis, and the radial distance r from the origin and the angle θ with the polar axis are used as coordinate parameters. Based on the motion state attributes of the disconnection repair frame (such as velocity v and acceleration a), the predicted trajectory of the next k frames is extrapolated through an Extended Kalman Filter (EKF) or a constant velocity model, and the coordinate sequence of the predicted points is calibrated at time intervals in the polar coordinate system Meanwhile, the spatial coordinates of the target in the starting frame of each candidate repair tracking trajectory are extracted and converted to the points to be determined in the same polar coordinate system. The predicted trajectory is divided into several segments in chronological order (each segment is 1 frame), and a decreasing confidence weight is assigned to each segment (for example, the weight of the first segment is 0.8, the second segment is 0.7, and the weights of subsequent segments decrease gradually) to reflect the cumulative effect of prediction uncertainty. For each point to be determined, the Manhattan distance between it and each segment of the predicted trajectory is calculated: the polar coordinates are converted to a local Cartesian grid (the grid resolution matches the target size), and the sum of the Manhattan distances from it to all grid points within the predicted segment is calculated, and the minimum value is taken as the measure of spatial proximity between the segment and the point to be determined. The segment of the predicted trajectory corresponding to the minimum Manhattan distance of each point to be determined is selected, and the correlation degree is calculated in combination with the confidence weight of this segment. If all correlation degrees are lower than the preset threshold, the mismatched trajectory is determined to be irreparable and marked as the termination state; if there are points to be determined with a correlation degree greater than the preset correlation threshold, the point to be determined with the highest correlation degree is selected, and the corresponding candidate repair trajectory and the mismatched trajectory are connected by cubic spline interpolation, and the motion history and multi-modal features in the trajectory memory pool are updated to complete the disconnection repair. The trajectory continuity is ensured by motion consistency first, improving the accuracy and continuity of target tracking in complex scenarios where the target appearance changes significantly after occlusion
[0040] Furthermore, the specific formula for the correlation degree is as follows:
[0041] where is the confidence weight of the th segment of the predicted trajectory, is the minimum Manhattan distance between the point i to be determined and the jth segment, is the preset maximum correlation distance
[0042] Figure 3A target recognition and tracking system 3 based on multi-source fusion imaging provided by an embodiment of the present invention, the system includes: a memory 31, a processor 32, the memory 31 contains a target recognition and tracking method program based on multi-source fusion imaging, and when the target recognition and tracking method program based on multi-source fusion imaging is executed by the processor 32, the following steps are implemented: Perform environmental monitoring on the target scene to obtain multi-modal environmental monitoring data and perform preprocessing, and perform image enhancement based on the preprocessed multi-modal environmental monitoring data to obtain image-enhanced multi-modal environmental monitoring data; Perform multi-modal feature extraction on the image-enhanced multi-modal environmental monitoring data, introduce a cross-attention mechanism to fuse the extracted multi-modal features to construct cross-modal fusion features; Construct a cascaded target recognition model, input the cross-modal fusion features to perform target recognition on the current monitoring scene, and obtain scene target recognition information; Track the scene recognition targets in the current monitoring scene based on the cross-modal fusion features and the scene target recognition information to generate scene target tracking information and store it in a preset trajectory memory pool; When a tracking target disconnection occurs during target tracking, obtain a mismatched tracking trajectory, and combine the trajectory memory pool to repair the tracking trajectory disconnection.
[0043] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0044] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0045] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0046] Those of ordinary skill in the art will understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, and other various media that can store program codes.
[0047] Alternatively, if the above integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks, and other various media that can store program codes.
[0048] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A target recognition and tracking method based on multi-source fusion imaging, characterized in that: include: Performing environmental monitoring on the target scene to obtain multimodal environmental monitoring data and preprocessing it, performing image enhancement based on the preprocessed multimodal environmental monitoring data to obtain image enhanced multimodal environmental monitoring data; Performing multimodal feature extraction on the image enhanced multimodal environment monitoring data, and introducing a cross-attention mechanism to fuse the extracted multimodal features to construct cross-modal fusion features; Constructing a cascade target recognition model, inputting the cross-modal fusion features to perform target recognition on the current monitoring scene, and obtaining scene target recognition information; Based on the cross-modal fusion features and the scene target recognition information, the scene recognition target in the current monitoring scene is tracked to generate the scene target tracking information and store it in the preset trajectory memory pool; When a target disconnection occurs during target tracking, a mismatched tracking trajectory is obtained and the tracking trajectory disconnection repair is performed in combination with the trajectory memory pool.
2. The target recognition and tracking method based on multi-source fusion imaging according to claim 1 is characterized in that: The step of performing environmental monitoring on the target scene to obtain multimodal environmental monitoring data and preprocessing the data specifically includes: Installing a heterogeneous sensor array in a target scene, and performing environmental monitoring on the target scene through the installed heterogeneous sensor array to obtain multimodal environmental monitoring data, wherein the heterogeneous sensor array includes a high-resolution imaging camera, an infrared thermal imaging camera, and a millimeter-wave radar; Based on the acquired multimodal environmental monitoring data, the data obtained from different monitoring sources are spatially registered, and the ORB feature extraction algorithm is introduced to extract scale pyramid features from the initial visible light image, and the illumination difference is eliminated by gray value normalization; After the illumination difference elimination is completed, the image region features of the initial visible light image for illumination difference elimination are extracted, and the initial infrared image is subjected to threshold segmentation processing in the corresponding region through the extracted image region features; Feature points are extracted from the initial visible light image and the initial infrared image respectively, and the extracted feature points are preliminarily matched to generate a matching point set through the Hamming distance. After the preliminary matching is completed, the RANSAC algorithm is used to iteratively screen the internal point set. Each time, 4 pairs of feature points are randomly selected to calculate the affine transformation matrix, and the reprojection errors of all matching segments in the affine transformation matrix are calculated. Points with a reprojection error smaller than the preset error are selected and counted into the internal point set. After repeated iterations, the final affine transformation matrix and the corresponding internal point set are retained to obtain the spatial registration result of the initial visible light image and the initial infrared image. Extract radar point cloud data through multimodal environmental monitoring data, convert the polar coordinates corresponding to the radar point cloud data into three-dimensional coordinates in a Cartesian coordinate system, perform rigid body transformation and projection to the imaging plane of a visible light camera in combination with the sensor's external parameter calibration matrix, and finally generate a two-dimensional point cloud heat map aligned with the optical image space; The preprocessed multimodal environmental monitoring data is obtained by combining the spatial registration results of the initial visible light image and the initial infrared image and the two-dimensional point cloud thermal map.
3. The target recognition and tracking method based on multi-source fusion imaging according to claim 1, characterized in that: The image enhancement is performed based on the preprocessed multimodal environment monitoring data to obtain image enhanced multimodal environment monitoring data, specifically including: Obtain preprocessed multimodal environmental monitoring data, input it into the HSI model to convert the original RGB image into the HSI color space, and obtain the hue component, saturation component, and brightness component through the HSI color space; The Retinex algorithm is introduced to decompose the brightness component into the reflection component and the illumination component, and the reflection component is processed in the digital domain to enhance the dark area of the image. The separated saturation component is adjusted nonlinearly and adaptively, and the color concentration is enhanced through the dynamic range extension function. The Canny edge detection operator is used to extract the structural features of the entire image and form an edge map. The edge map is weightedly fused with the enhanced brightness component at the pixel level. After processing each component, the optimized H, S, and I channels are re-fused, and the RGB image is reconstructed through inverse color space conversion to obtain the initial enhanced image. A Markov algorithm is introduced to reprocess the initial enhanced image, the initial enhanced image is imported into a convolutional autoencoder network for encoding and mapping to a latent space, and an initial latent variable is generated as a starting state of a Markov chain; Define the state distribution and calculate the state transition probability through the generated initial latent variables, generate candidate latent variables according to the state transition probability, and use the Metropolis criterion to determine whether to accept the target candidate latent variables; Iterative optimization is performed until the stopping criteria are met and the final latent variables are output. The final latent variables are input into the pre-trained U-Net decoder network, and the high-resolution image is gradually reconstructed through deconvolution layers and jump connections to obtain image-enhanced multimodal environmental monitoring data.
4. The target recognition and tracking method based on multi-source fusion imaging according to claim 1 is characterized in that: The multimodal feature extraction is performed on the image enhanced multimodal environment monitoring data, and a cross attention mechanism is introduced to fuse the extracted multimodal features to construct cross-modal fusion features, which specifically includes: Acquire image-enhanced multimodal environmental monitoring data, and input the image-enhanced multimodal environmental monitoring data into a multi-channel feature extraction network for multimodal feature extraction, wherein the multi-channel feature extraction network includes a first feature extraction channel, a second feature extraction channel, and a third feature extraction channel; Importing the visible light image into the first feature extraction channel to extract multi-scale context features based on the hybrid hole convolution layer, capturing context information under different receptive fields according to the parallel convolution branch, and then compressing it into a multi-dimensional global vector through global average pooling to obtain the first feature extraction information; The infrared image is imported into the second feature extraction channel for spatial feature extraction, and the extracted spatial features are input into the SE attention module to perform global average pooling to generate a channel description vector. The channel description vector is used to obtain the channel weight matrix and multiply it with the original feature map to extract the local thermal feature vector to obtain the second feature extraction information; The radar point cloud data is imported into the third feature extraction channel to extract the three-dimensional spatial features, k key points are selected as the initial point set through the farthest point sampling, and the geometric contour and motion trajectory of the target object are captured by combining the local aggregation mechanism to generate the third feature extraction information; The cross-attention mechanism is introduced to perform cross-modal feature association based on the first feature extraction information, the second feature extraction information and the third feature extraction information. The visible light image features are used as queries, and the infrared image features and radar point cloud features are used as keys. The spatial correlation weights are calculated through multi-head attention to generate cross-modal fusion features.
5. The target recognition and tracking method based on multi-source fusion imaging according to claim 1, characterized in that: The cascade target recognition model is constructed, and the cross-modal fusion feature is input to perform target recognition on the current monitoring scene to obtain scene target recognition information, specifically including: Construct a cascade target recognition model, which is constructed with Bi-LSTM and multi-scale detection head as the framework, obtain cross-modal fusion features and input them into the cascade target recognition model to analyze the current scene; Using the Bi-LSTM network to perform temporal context modeling based on the input cross-modal fusion features, using forward and backward transfer based on the stacked two-layer LSTM unit to respectively capture the spatiotemporal context features of the objects in the target scene, and output spatiotemporal enhanced features; The spatiotemporal enhanced features are input into a multi-scale detection head for target parsing, a multi-level feature abstraction is performed through a lightweight GhostNet module to obtain multiple sets of low-rank feature maps, and a cross-stage local connection module is used to extract local detail features; Performing deep separable convolution compression on the obtained multiple groups of low-rank feature maps to obtain channel dimensions, performing channel cascade with the local detail features to obtain cascade fusion features, introducing spatial pyramid pooling and inputting the cascade fusion features, extracting cross-receptive field features through parallel multi-scale maximum pooling and splicing them with the original feature maps to obtain a three-level feature pyramid; Based on the obtained three-level feature pyramid, the target scene is recognized to generate several candidate recognition frames, and the candidate recognition frames are subjected to non-maximum suppression and mask threshold filtering to obtain scene target recognition information.
6. The target recognition and tracking method based on multi-source fusion imaging according to claim 1 is characterized in that: The step of tracking the scene recognition target in the current monitoring scene based on the cross-modal fusion feature and the scene target recognition information to generate the scene target tracking information and storing it in a preset trajectory memory pool specifically includes: Acquire cross-modal fusion features and scene target recognition information, extract motion state features of the scene recognition target based on the cross-modal fusion features, and import the extracted motion state features into an extended Kalman filter to construct a state vector; Based on the constructed state vector, the position distribution of the scene recognition target in the next frame image is predicted to obtain a target prediction frame, and the Mahalanobis distance between the target prediction frame and the target detection frame is calculated in combination with the scene target recognition information to generate a motion cost matrix; The appearance state features of the scene recognition target are obtained through the cross-modal fusion features, and the cosine similarity is calculated with the target appearance embedding in the adjacent frame image, and an appearance cost matrix is generated and weighted fused with the motion cost matrix to generate a comprehensive cost matrix as the input of the Hungarian algorithm to solve the optimal matching pair. According to the optimal matching pair obtained by the solution, the scene target tracking information is generated and stored in a preset trajectory memory pool.
7. The target recognition and tracking method based on multi-source fusion imaging according to claim 1 is characterized in that: When a target disconnection occurs during target tracking, a mismatched tracking trajectory is obtained, and the tracking trajectory disconnection repair is performed in combination with the trajectory memory pool, specifically including: Acquire a mismatch tracking trajectory, extract the last frame image where the trajectory mismatch occurs based on the mismatch tracking trajectory, define it as a disconnection repair start frame, extract the frame timing features and multimodal features of the disconnection repair start frame, and obtain disconnection repair frame feature information; Acquire a trajectory memory pool, extract a scene target tracking trajectory after the frame timing feature of the disconnected repair frame based on the disconnected frame feature information, define it as an initial repair tracking trajectory, and extract the motion state attribute of the disconnected repair frame through the multimodal feature of the disconnected repair start frame; Extract the motion state attributes of each initial repair tracking trajectory, calculate the deviation with the motion state attributes of the disconnected repair frame, and select the initial repair tracking trajectory with the deviation within a preset range as the candidate repair tracking trajectory; The multimodal features of each candidate repair trajectory are extracted, and similarity calculation is performed with the disconnected repair frame feature information to obtain several candidate repair tracking trajectories greater than a preset similarity threshold, and the candidate repair tracking trajectory with the largest similarity value is selected as the final repair tracking trajectory to be connected with the mismatch tracking trajectory.
8. The target recognition and tracking method based on multi-source fusion imaging according to claim 1, characterized in that: When a target disconnection occurs during target tracking, a mismatched tracking trajectory is obtained, and the tracking trajectory disconnection repair is performed in combination with the trajectory memory pool, further comprising: If the similarity values of the multimodal features of all candidate repair tracking trajectories and the feature information of the disconnected repair frame are less than the preset similarity threshold, the spatial coordinate data of the scene tracking target in the disconnected repair frame is obtained as the origin to establish a polar coordinate system, and the predicted trajectory of the next k frames is generated based on the motion state attributes of the disconnected repair frame, and the predicted trajectory coordinates are calibrated in the established polar coordinate system; The spatial coordinate data of the scene tracking target in the starting frame of each candidate repair tracking trajectory is extracted according to the multimodal features of the scene target associated with the starting frame of each candidate repair tracking trajectory, and the spatial coordinate data is calibrated in the established polar coordinate system to be defined as the point to be determined; The predicted trajectory in the polar coordinates is segmented according to the time series features and the confidence of different predicted trajectory segments is set based on the order of the time series features. The Manhattan distance between each point to be determined and the predicted trajectory calibrated to the polar coordinate system at different time series is calculated in the constructed polar coordinate system. According to the calculated Manhattan distance, the predicted trajectory segment with the smallest Manhattan distance is selected as the target predicted trajectory segment corresponding to the point to be determined, and weighted fusion is performed in combination with the confidence corresponding to the target predicted trajectory segment to obtain the correlation between each point to be determined and the predicted trajectory; The correlation between each point to be determined and the preset trajectory is judged with a preset correlation threshold. If there is no point to be determined with a correlation greater than the preset correlation threshold, the target mismatch tracking trajectory is marked as an unrepairable trajectory; If there is a point to be determined that is greater than a preset correlation threshold, the point to be determined with the largest correlation is selected as the target determination point, and the candidate repair tracking trajectory corresponding to the target determination point is defined as the target repair tracking trajectory, which is connected to the target mismatch tracking trajectory to complete the trajectory disconnection repair.
9. A target recognition and tracking system based on multi-source fusion imaging, characterized in that: The system includes: a memory and a processor, wherein the memory contains a target recognition and tracking method program based on multi-source fusion imaging, and when the target recognition and tracking method program based on multi-source fusion imaging is executed by the processor, the following steps are implemented: Performing environmental monitoring on the target scene to obtain multimodal environmental monitoring data and preprocessing it, performing image enhancement based on the preprocessed multimodal environmental monitoring data to obtain image enhanced multimodal environmental monitoring data; Performing multimodal feature extraction on the image enhanced multimodal environment monitoring data, and introducing a cross-attention mechanism to fuse the extracted multimodal features to construct cross-modal fusion features; Constructing a cascade target recognition model, inputting the cross-modal fusion features to perform target recognition on the current monitoring scene, and obtaining scene target recognition information; Based on the cross-modal fusion features and the scene target recognition information, the scene recognition target in the current monitoring scene is tracked to generate the scene target tracking information and store it in the preset trajectory memory pool; When a target disconnection occurs during target tracking, a mismatched tracking trajectory is obtained and the tracking trajectory disconnection repair is performed in combination with the trajectory memory pool.
Citation Information
Patent Citations
Three-dimensional single target tracking method based on multi-modal information fusion
CN115880333A
Image processing method and device for golf course image, and equipment
WO2018166084A1
Method for high-precision multi-target tracking against complex background
WO2022217840A1
Cited By
Intelligent road inspection multi-terminal collaborative management method and system based on edge calculation
CN120499349A
Automatic target identification and tracking method and system for intelligent pod
CN120522691A
Intelligent pod automatic target recognition and tracking method and system
CN120522691B
Dynamic spatial data set construction system and method for bridge structural member
CN120544009A
A system and method for constructing dynamic spatial data sets of bridge structural components
CN120544009B