Multi-target pedestrian re-identification system based on multi-mode and vector database

Through a multi-object pedestrian re-identification system combining multimodal and vector database, the accuracy and robustness of target pedestrian recognition in complex indoor environments are solved, and high-precision cross-camera matching and trajectory association are achieved.

CN120496174AActive Publication Date: 2025-08-15YUNTU DATA TECH (ZHENGZHOU) CO LTD

Patent Information

Application Number
CN202510562660.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high accuracy, high robustness and high real-time target pedestrian recognition and cross-camera matching in complex indoor environments. Especially when lighting conditions are complex, viewing angle differences are significant and occlusion frequently, the existing methods lack the diversity and complementary support for feature expressions, resulting in a decrease in recognition accuracy and robustness.

Method used

A multi-objective pedestrian re-identification system based on multi-modal and vector database is adopted. Through the combination of monocular tracking module, multi-modal extraction module, trajectory generation module, multi-objective matching module and global search module, the fuzzy dynamic decision-making mechanism, quality reconstruction judgment mechanism, timing moving trajectory generation mechanism and space-time constraint mechanism are used to improve the accuracy and applicability of cross-modal pedestrian re-identification.

Benefits of technology

It significantly improves the accuracy and applicability of cross-modal pedestrian re-identification, enhances the modeling ability of target dynamic characteristics in complex scenarios, and ensures the logical consistency and global optimization of target character trajectory associations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496174A_ABST
    Figure CN120496174A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target pedestrian re-identification system based on multiple modes and a vector database, relates to the technical field of network communication and positioning, and solves the problem of cross-target and cross-mode trajectory association in a complex multi-camera scene. The multi-target pedestrian re-recognition system comprises a monocular tracking module, a multimode extraction module, a trajectory generation module, a multi-objective matching module and a global retrieval module, through organic combination of multi-modal features and a multi-modal multi-path recall strategy, the accuracy and applicability of cross-modal pedestrian re-identification are significantly improved. Through track-level feature generation and storage design, the modeling capability of dynamic features of a target in a complex scene is enhanced; through collaborative design of a space-time constraint mechanism and multi-modal features, logic consistency and global optimality of target person trajectory association are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network communication and positioning technology, and in particular to a multi-target pedestrian re-identification system based on a multimodal and vector database. Background Art

[0002] In intelligent surveillance systems, the demand for multi-target multi-camera (MTMC) pedestrian re-identification technology in indoor scenarios is increasing, such as smart building management, shopping mall behavior analysis, and office security.

[0003] In indoor scenes, due to complex lighting conditions, significant differences in viewpoints, and frequent occlusions, relying solely on single-modal features (such as visual appearance or motion trajectory) is difficult to comprehensively and stably characterize target pedestrians, resulting in reduced recognition accuracy and robustness. Existing methods lack support for diverse and complementary feature representation, limiting their adaptability in complex environments.

[0004] In indoor scenes with multiple cameras with non-overlapping fields of view, the appearance of a target pedestrian often varies significantly from camera to camera, and some features may even be missing or inconsistent. Traditional person re-identification techniques rely heavily on appearance similarity for matching, but struggle with complex situations involving similar-looking pedestrians or incomplete features, limiting cross-camera correlation performance.

[0005] Although the number of target pedestrians in indoor scenes is limited, due to the dense deployment of cameras, the large amount of collected data and the diversity of target features, existing methods have bottlenecks in multimodal feature fusion and retrieval efficiency, and cannot meet the actual needs of high real-time and scalability.

[0006] Patent No. CN202411711350.9 discloses a method and system for realizing personalized recommendation based on multi-feature fusion face recognition, aiming to provide real-time, efficient and accurate recommendation services. The method includes the following steps: real-time facial images of the target audience are collected through camera equipment, and the images are pre-processed to extract the face area; a face recognition algorithm is run in the edge computing device to extract multimodal features including age, gender, appearance and style features; a comprehensive feature vector is generated by the Transformer model through the feature fusion module to realize style classification and appearance classification; personalized recommendation content is generated by matching user feature vectors with content databases based on cosine similarity, and the classification model weights are dynamically optimized based on user feedback. The above invention accurately captures user dynamic behavior and appearance characteristics through time series analysis, dynamic style analysis and deep appearance analysis modules; and comprehensively generates user feature descriptions by combining visual, voice and text multimodal information.

[0007] Patent No. CN202311203484.5 discloses an object recognition method based on multimodal features. It includes: obtaining a gait graph sequence set BT of the video to be detected; according to BT, obtaining an optimal frame image list set BTY; performing feature extraction on BT and BTY to obtain a feature vector set TZ; according to BT and BTY, obtaining an image quality score list set ZL; obtaining a feature vector list MZ of the target object; according to MZ and TZ, obtaining a matching degree list set P; according to ZL and P, obtaining a comprehensive matching degree set PY; if the maximum comprehensive matching degree in PY is greater than a preset matching degree threshold, then the object to be detected corresponding to the maximum comprehensive matching degree in PY is marked as a key object. The above application comprehensively considers the three factors of facial features, body features and gait features of the object to be detected, and determines the weight according to the corresponding quality scores to obtain a comprehensive matching degree. Compared with identity recognition based solely on gait features, it is more accurate.

[0008] However, the above patents and existing systems face many challenges in complex indoor environments, making it difficult to achieve high-precision, high-robustness, and high-real-time recognition and cross-camera matching of target pedestrians. Summary of the Invention

[0009] The purpose of the present invention is to provide a multi-target pedestrian re-identification system based on multimodality and vector database. It can significantly improve the accuracy and applicability of cross-modal pedestrian re-identification by organically combining multimodal features, vector database and multimodal multi-path recall strategy; enhance the modeling ability of target dynamic characteristics in complex scenarios through trajectory-level feature generation and storage design; and ensure the logical consistency and global optimality of target person trajectory association through the collaborative design of spatiotemporal constraint mechanism and multimodal features.

[0010] The present invention utilizes the following technical solutions:

[0011] A multi-target pedestrian re-identification system based on multimodal and vector database includes a monocular tracking module, a multimodal extraction module, a trajectory generation module, a multi-target matching module and a global retrieval module; wherein,

[0012] The monocular tracking module is used to locate and track the target person in each frame of the original image based on the fuzzy dynamic decision-making mechanism combined with the image enhancement judgment model, and generate multimodal information;

[0013] The multimodal extraction module is used to process multimodal information according to the quality reconstruction judgment mechanism, extract and fuse multimodal features, and store them in the vector database;

[0014] The trajectory generation module is used to extract trajectory-level feature vectors of multimodal features based on the temporal motion trajectory generation mechanism;

[0015] The multi-camera matching module is used to match the trajectory-level feature vectors of different cameras according to the spatiotemporal constraint mechanism to generate the target person's trajectory;

[0016] The global retrieval module is used to search and optimize in the vector database based on the multi-modal and multi-path recall strategy and the new target trajectory features.

[0017] Preferably, the multimodal information includes target trajectory data, boundary detection frame and segmentation mask; the monocular tracking module adopts a decomposition extraction algorithm to convert the videos captured by all cameras into several frames of original images; the blurriness of each frame of the original image is detected and quantified according to the fuzzy dynamic decision mechanism, and judged with a preset blur threshold to obtain a clear image and a blurred image, and then the quality of the blurred image is enhanced to obtain an enhanced image; a target detector is used to detect each target person in the enhanced image and the clear image to obtain a boundary detection frame and a confidence score; an instance segmentation algorithm is used to segment the boundary detection frame to generate a segmentation mask; a target tracking algorithm is used to combine the boundary detection frame with the segmentation mask, associate and track the trajectory of the target person in the continuous frame original image, obtain target trajectory data, and assign a unique identifier.

[0018] Preferably, the operation process of the fuzzy dynamic decision-making mechanism is:

[0019] The gradient analysis algorithm is combined with the frequency domain analysis algorithm to quantify the blur degree of each frame of the original image and obtain the image blur;

[0020] Compare the image blur with the preset blur threshold: if the image blur is less than the blur threshold, the current image is determined to be a clear image, and target detection, target segmentation and target tracking operations are directly performed;

[0021] If the image blur is greater than or equal to the blur threshold, the current image is judged to be a blurred image, and the feasibility prediction is input into the quantization judgment layer of the image enhancement judgment model to obtain the feasibility quantization value, which is then compared with the preset quantization threshold:

[0022] If the feasibility quantization value is less than the quantization threshold, the image completion layer of the input image enhancement judgment model is combined with the unique identifier and timestamp, and the generative adversarial network is used to enhance the current image to obtain a clear image for target detection, target segmentation and target tracking operations, while the current image is discarded.

[0023] If the feasibility quantization value is greater than or equal to the quantization threshold, it is input into the feature extraction layer of the image enhancement judgment model, and the fuzzy image is extracted with the principal component analysis algorithm to obtain the global fuzzy matrix;

[0024] The enhanced feedback layer of the image enhancement decision model is combined with the PID algorithm to perform resolution adaptation on the global fuzzy matrix to generate a super-resolution feedback matrix;

[0025] The iterative training layer of the image enhancement decision model combines the ESPCN network and the FSRCNN network to perform several iterative training on the super-resolution feedback matrix to obtain the full-image super-resolution weight matrix;

[0026] The prediction decision layer of the image enhancement judgment model adopts the divergence-cross entropy loss function to update and optimize the full-image super-resolution weight matrix and output the enhanced image at the same time.

[0027] Preferably, the multimodal extraction module preprocesses the multimodal information according to the quality reconstruction determination mechanism:

[0028] A blur detector is used to detect the clarity of the boundary detection box and segmentation mask, and then compared with the preset clarity threshold:

[0029] If the clarity of the boundary detection box or segmentation mask is less than the clarity threshold, the EDSR network combined with the ESRGAN network is used to perform local super-resolution reconstruction on the current boundary detection box or segmentation mask;

[0030] If the clarity of the boundary detection box or segmentation mask is greater than or equal to the clarity threshold, the edge detection algorithm is used to check the integrity of the boundary detection box and segmentation mask:

[0031] If the boundary detection box is incomplete, the current boundary detection box is estimated according to the timestamp and background space based on the adjacent frame image or target trajectory data, combined with the target motion direction and trajectory prediction, to obtain the actual boundary box;

[0032] If the segmentation mask is incomplete, the mask completion algorithm is used to restore and complete the current segmentation mask based on the key points of the human body and the actual bounding box to obtain a complete mask, thereby completing the multimodal information preprocessing.

[0033] Preferably, the multimodal extraction module uses a visual transformer combined with an actual bounding box to extract global visual features of the target person to obtain an overall appearance vector; the overall appearance vector includes clothing, posture, height, body shape, skin color, hairstyle and texture;

[0034] At the same time, a human pose estimator is used in combination with a complete mask to extract local visual features of the target person and obtain component-level feature vectors; component-level feature vectors include the head, torso, and limbs;

[0035] The face recognition algorithm is used in combination with the actual bounding box to detect the face of the target person, and the facial feature vector and target attribute information are obtained; the target attribute information includes gender, age group and expression;

[0036] The image-language model is used to describe the actual bounding box in language to obtain language feature information; and the language embedding model is used to combine the language feature information with the target attribute information to obtain a semantic vector.

[0037] A multimodal fusion algorithm is used to jointly optimize the semantic vector with the overall appearance vector, component-level feature vector, and facial feature vector to obtain a multimodal feature vector.

[0038] The spatial position of the target person entering and leaving the camera's field of view is extracted to obtain the center coordinates of the target detection frame. Combined with the monocular ranging algorithm, it is converted into 3D world coordinates, and the appearance and disappearance timestamps are recorded at the same time.

[0039] At the same time, the trajectory points of the target person in the field of view of the camera device are captured to obtain a trajectory point sequence, and then the motion characteristics of the target person are calculated; the motion characteristics include speed, direction and acceleration;

[0040] The multimodal feature vectors are fused and appended with the trajectory point sequence respectively to be stored in the corresponding vector database.

[0041] Preferably, the trajectory generation module uses a sliding window moving average algorithm to perform local denoising on the multimodal feature vector of each frame according to the time-series motion trajectory generation mechanism to obtain a frame-level smooth feature vector; the frame-level smooth feature vector is temporally clustered according to the appearance timestamp and disappearance timestamp to generate several modal clustering clusters; each modal clustering cluster is averaged and fused to generate a cluster-level feature vector; the cluster-level feature vector is weightedly fused according to the number of cluster frames to generate a trajectory-level feature vector; the cluster-level feature vector and the trajectory-level feature vector are stored in a vector database according to the trajectory point sequence.

[0042] Preferably, the multi-camera matching module determines and counts the adjacent cameras of each camera according to the deployment position and field of view of different cameras to obtain an adjacency table of each camera; calculates the field of view intersection between cameras according to the three-dimensional world coordinates of each camera; and performs trajectory matching and fusion of the target person based on the trajectory-level feature vector of each camera in the vector database combined with the spatiotemporal constraint mechanism:

[0043] If the timestamps of N trajectory-level feature vectors are the same and they are located in the field of view of the same camera device, then each trajectory-level feature vector is determined to belong to a different target person and cannot be merged into the same target person trajectory;

[0044] If N trajectory-level feature vectors have the same timestamp but are located in the field of view of different cameras and have no field of view intersection, then each trajectory-level feature vector is determined to belong to a different target person and cannot be merged into the same target person trajectory;

[0045] If N trajectory-level feature vectors have the same timestamp but are located in the field of view of different cameras and have overlapping fields of view, the Euclidean distance between the trajectory-level feature vectors is calculated and compared with the preset trajectory distance threshold:

[0046] If the Euclidean distance is less than the trajectory distance threshold, the current trajectory-level feature vectors are determined to belong to the same target person and are merged into the same target person trajectory;

[0047] If the Euclidean distance is greater than or equal to the trajectory distance threshold, it is determined that the current trajectory-level feature vectors belong to different target persons and cannot be merged into the same target person trajectory;

[0048] If the timestamps of N trajectory-level feature vectors are different and different camera devices are all in the adjacency table, the modal comprehensive matching degree of the multimodal feature vectors in different field of view spaces is calculated and compared with the preset matching threshold:

[0049] If the modal comprehensive matching degree is greater than or equal to the matching threshold, the current trajectory-level feature vector is determined to belong to the same target person, and is merged into the same target person trajectory. The unmatched trajectory-level feature vector and the last segment of the target person trajectory are marked to obtain the candidate trajectory to be matched;

[0050] If the modal comprehensive matching degree is less than the matching threshold, it is determined that the current trajectory-level feature vectors belong to different target persons and cannot be merged into the same target person trajectory.

[0051] Preferably, the global retrieval module performs multimodal feature extraction on the new target trajectory to obtain a new frame-level smoothed feature vector, and generates a new cluster-level feature vector and a new trajectory-level feature vector through a moving average algorithm and a time series aggregation algorithm. The process of searching the trajectory-level feature vector in the vector database according to the multimodal multi-path recall strategy is as follows:

[0052] The new trajectory-level feature vector of each modality is decomposed to obtain several trajectory element points. At the same time, the candidate trajectories to be matched are parsed, matched and sorted with several trajectory element points to obtain the trajectory matching scores of different modalities.

[0053] The trajectory matching scores of each modality are weighted and fused according to the boundary detection box to obtain the total confidence trajectory matching score;

[0054] The new trajectory-level feature vector is verified with spatiotemporal constraints based on the total confidence trajectory matching score: if the new trajectory-level feature vector successfully matches the unmatched trajectory-level feature vector, the two trajectory-level feature vectors are merged into the same target person trajectory;

[0055] If the new trajectory-level feature vector successfully matches the last segment of the target person's trajectory, the new trajectory-level feature vector is incorporated into the matched target person's trajectory; if neither is successfully matched, the next layer of cluster-level feature vector retrieval is performed;

[0056] The process of hierarchical retrieval of cluster-level feature vectors and frame-level smooth feature vectors of the vector database by the multi-mode multi-way recall strategy is as follows:

[0057] Locally aggregate the unmatched trajectory-level feature vector and the last segment of the target person's trajectory to obtain the cluster-level features to be matched;

[0058] The new cluster-level feature vector of each modality is decomposed to obtain several cluster-level element points. At the same time, the cluster-level features to be matched are analyzed, matched and sorted with several cluster-level element points to obtain cluster-level matching scores of different modalities.

[0059] The cluster-level matching scores of each modality are weighted and fused according to the boundary detection box to obtain the total confidence cluster-level matching score;

[0060] The new cluster-level feature vector is verified with spatiotemporal constraints based on the total confidence cluster-level matching score: if the new cluster-level feature vector successfully matches the unmatched cluster-level feature vector, the two cluster-level feature vectors are merged into the same target person trajectory;

[0061] If the new cluster-level feature vector successfully matches the last segment of the target person's cluster-level features, the new cluster-level feature vector is incorporated into the matched target person's cluster-level features;

[0062] If no match is successful, the computing resource status is estimated: if the computing resources are idle, the next layer of frame-level feature vector search is performed; if the computing resources are busy, the new cluster-level feature vector is saved and marked until the computing resources are idle and the frame-level feature vector search is performed;

[0063] The cluster-level features to be matched are retrieved frame by frame to obtain the frame-level features to be matched;

[0064] The new frame-level feature vector of each modality is decomposed to obtain several frame-level element points. At the same time, the frame-level features to be matched are analyzed, matched and sorted with several frame-level element points to obtain the frame-level matching scores of different modalities.

[0065] The frame-level matching scores of each modality are weighted and fused according to the boundary detection box to obtain the total confidence frame-level matching score;

[0066] The new frame-level feature vector is verified for spatiotemporal constraints based on the total confidence frame-level matching score: if the new frame-level feature vector successfully matches the unmatched frame-level feature vector, the two frame-level feature vectors are merged into the same target person trajectory;

[0067] If the new frame-level feature vector successfully matches the last segment of the target person's frame-level features, the new frame-level feature vector is incorporated into the matched target person's frame-level features;

[0068] If no matching is successful, the new frame-level feature vector is saved and marked as an unmatched trajectory.

[0069] Preferably, the operation process of the global search module further includes the following steps:

[0070] The time intervals between frame-level feature vectors, cluster-level feature vectors, and trajectory-level feature vectors are calculated based on the timestamps, and compared with the preset time thresholds:

[0071] If the time interval is greater than the time threshold, the time connection between the two trajectory feature vectors is judged to be unreasonable and the association is cancelled;

[0072] If the time interval is less than or equal to the time threshold, the two trajectory feature vectors are marked as trajectories to be completed;

[0073] According to the field of view space combined with the homography matrix, the spatial distance between frame-level feature vectors, cluster-level feature vectors, and trajectory-level feature vectors is calculated and compared with the preset distance threshold:

[0074] If the spatial distance is greater than the distance threshold, the spatial movement of the two trajectory feature vectors is judged to be unreasonable and the association is cancelled;

[0075] If the spatial distance is less than or equal to the distance threshold, the two trajectory feature vectors are marked as trajectories to be completed;

[0076] The completed trajectory is analyzed according to time and space to obtain time feature parameters and space feature parameters;

[0077] The temporal characteristic parameters include the trajectory start time, trajectory end time and time interval; the spatial characteristic parameters include the starting point coordinates, end point coordinates and spatial distance;

[0078] Based on the temporal and spatial feature parameters combined with the time threshold and distance threshold, the similarity of all target person trajectory segments in the vector database is calculated and compared with the preset similarity threshold:

[0079] If the similarity is greater than or equal to the similarity threshold, the pedestrian re-identification algorithm is used to match the start and end points of the completed segment;

[0080] If the similarity is less than the similarity threshold, the Bezier curve interpolation algorithm is used to generate smooth trajectory completion segments to fill the gaps between the target person's trajectory segments.

[0081] The present invention significantly improves the accuracy and applicability of cross-modal person re-identification by organically combining multimodal features with a multimodal multi-path recall strategy. It enhances the modeling capability of target dynamic characteristics in complex scenarios through trajectory-level feature generation and storage design. It ensures the logical consistency and global optimality of target person trajectory association through the collaborative design of spatiotemporal constraint mechanism and multimodal features. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0083] Figure 1 This is the schematic diagram of the multi-target person re-identification system;

[0084] Figure 2 This is a schematic diagram of the principle of the image enhancement judgment model;

[0085] Figure 3 Store flow charts for multimodal feature extraction;

[0086] Figure 4 This is the flow chart of the multi-mode and multi-channel recall strategy. DETAILED DESCRIPTION

[0087] The present invention is described in detail below with reference to the accompanying drawings and embodiments:

[0088] like Figures 1 to 4 As shown, the multi-target pedestrian re-identification system based on multimodal and vector database of the present invention includes a monocular tracking module, a multimodal extraction module, a trajectory generation module, a multi-target matching module and a global retrieval module; wherein,

[0089] The monocular tracking module is used to locate and track the target person in each frame of the original image based on the fuzzy dynamic decision-making mechanism combined with the image enhancement judgment model, and generate multimodal information;

[0090] The multimodal extraction module is used to process multimodal information according to the quality reconstruction judgment mechanism, extract and fuse multimodal features, and store them in the vector database;

[0091] The trajectory generation module is used to extract trajectory-level feature vectors of multimodal features based on the temporal motion trajectory generation mechanism;

[0092] The multi-camera matching module is used to match the trajectory-level feature vectors of different cameras according to the spatiotemporal constraint mechanism to generate the target person's trajectory;

[0093] The global retrieval module is used to search and optimize in the vector database based on the multi-modal and multi-path recall strategy and the new target trajectory features.

[0094] In the present invention, multimodal information includes target trajectory data, boundary detection frame and segmentation mask; the monocular tracking module adopts a decomposition and extraction algorithm to convert the videos captured by all cameras into several frames of original images; the fuzziness of each frame of the original image is detected and quantified according to the fuzzy dynamic decision mechanism, and judged with a preset fuzzy threshold to obtain a clear image and a blurred image, and then the quality of the blurred image is enhanced to obtain an enhanced image; a target detector is used to detect each target person in the enhanced image and the clear image to obtain a boundary detection frame and a confidence score; an instance segmentation algorithm is used to segment the boundary detection frame to generate a segmentation mask; a target tracking algorithm is used to combine the boundary detection frame with the segmentation mask, associate and track the trajectory of the target person in the continuous frame original image, obtain target trajectory data, and assign a unique identifier.

[0095] In the present invention, the operation process of the fuzzy dynamic decision-making mechanism is as follows:

[0096] The gradient analysis algorithm is combined with the frequency domain analysis algorithm to quantify the blur degree of each frame of the original image and obtain the image blur;

[0097] Compare the image blur with the preset blur threshold: if the image blur is less than the blur threshold, the current image is determined to be a clear image, and target detection, target segmentation and target tracking operations are directly performed;

[0098] If the image blur B(I) is greater than or equal to the blur threshold Tb, the current image I is determined to be a blurred image, and the quantization decision layer fq of the image enhancement decision model is input for feasibility prediction to obtain the feasibility quantization value Q, which is then compared with the preset quantization threshold Tq:

[0099] In this embodiment, Where F represents the set of blurred images; Q = fq(I;θq), Among them, θq represents the quantization parameter, θG represents the completion parameter, and T represents the transpose symbol;

[0100] If the feasibility quantization value is less than the quantization threshold, the image completion layer of the input image enhancement judgment model is combined with the unique identification ID and timestamp t, and the generative adversarial network G is used to enhance the current image to obtain a clear image Ien for target detection, target segmentation and target tracking operations, while the current image is discarded;

[0101] If the feasibility quantization value is greater than or equal to the quantization threshold, it is input into the feature extraction layer of the image enhancement judgment model, and the principal component analysis algorithm P is combined to extract the features of the blurred image to obtain the global fuzzy matrix Mg;

[0102] The enhanced feedback layer of the image enhancement decision model is combined with the PID algorithm to perform resolution adaptation on the global fuzzy matrix to generate a super-resolution feedback matrix Mf;

[0103] In this embodiment, Wherein, e(t) = ||Mg - Mtarget||2, Kp, Ki, Kd are PID coefficients, e(t) represents the matrix difference, Mtarget represents the preset resolution matrix, and τ represents the number of sampling times;

[0104] The iterative training layer of the image enhancement judgment model combines the ESPCN network and the FSRCNN network to perform several iterative training on the super-resolution feedback matrix to obtain the full-image super-resolution weight matrix

[0105] In this embodiment, ⊕ represents the network fusion operation, k is the number of iterations;

[0106] The prediction decision layer of the image enhancement judgment model adopts the divergence-cross entropy loss function L to update and optimize the full-image super-resolution weight matrix and output the enhanced image Ienh.

[0107] In this embodiment,

[0108]

[0109] Among them, α represents the allocation weight coefficient; D KL represents the divergence function, p represents the ideal distribution probability vector, q represents the predicted distribution probability vector, H represents the cross entropy function, and y represents the predicted label. represents the true label, W represents the set of full-image super-resolution weight matrices;

[0110] In the present invention, the multimodal extraction module preprocesses the multimodal information according to the quality reconstruction judgment mechanism:

[0111] A blur detector is used to detect the clarity of the boundary detection box and segmentation mask, and then compared with the preset clarity threshold:

[0112] If the clarity of the boundary detection box or segmentation mask is less than the clarity threshold, the EDSR network combined with the ESRGAN network is used to perform local super-resolution reconstruction on the current boundary detection box or segmentation mask;

[0113] If the clarity of the boundary detection box or segmentation mask is greater than or equal to the clarity threshold, the edge detection algorithm is used to check the integrity of the boundary detection box and segmentation mask:

[0114] If the boundary detection box is incomplete, the current boundary detection box is estimated according to the timestamp and background space based on the adjacent frame image or target trajectory data, combined with the target motion direction and trajectory prediction, to obtain the actual boundary box;

[0115] If the segmentation mask is incomplete, the mask completion algorithm is used to restore and complete the current segmentation mask based on the key points of the human body and the actual bounding box to obtain a complete mask, thereby completing the multimodal information preprocessing.

[0116] In the present invention, the multimodal extraction module uses a visual transformer combined with an actual bounding box to extract global visual features of the target person and obtain an overall appearance vector; the overall appearance vector includes clothing, posture, height, body shape, skin color, hairstyle and texture;

[0117] At the same time, a human pose estimator is used in combination with a complete mask to extract local visual features of the target person and obtain component-level feature vectors; component-level feature vectors include the head, torso, and limbs;

[0118] The face recognition algorithm is used in combination with the actual bounding box to detect the face of the target person, and the facial feature vector and target attribute information are obtained; the target attribute information includes gender, age group and expression;

[0119] The image-language model is used to describe the actual bounding box in language to obtain language feature information; and the language embedding model is used to combine the language feature information with the target attribute information to obtain a semantic vector.

[0120] A multimodal fusion algorithm is used to jointly optimize the semantic vector with the overall appearance vector, component-level feature vector, and facial feature vector to obtain a multimodal feature vector.

[0121] The spatial position of the target person entering and leaving the camera's field of view is extracted to obtain the center coordinates of the target detection frame. Combined with the monocular ranging algorithm, it is converted into 3D world coordinates, and the appearance and disappearance timestamps are recorded at the same time.

[0122] At the same time, the trajectory points of the target person in the field of view of the camera device are captured to obtain a trajectory point sequence, and then the motion characteristics of the target person are calculated; the motion characteristics include speed, direction and acceleration;

[0123] The multimodal feature vectors are fused and appended with the trajectory point sequence respectively to be stored in the corresponding vector database.

[0124] In the present invention, the trajectory generation module adopts a sliding window moving average algorithm to perform local denoising on the multimodal feature vector of each frame according to the temporal motion trajectory generation mechanism to obtain a frame-level smooth feature vector; the frame-level smooth feature vector is temporally clustered according to the appearance timestamp and disappearance timestamp to generate several modal clustering clusters; each modal clustering cluster is averaged and fused to generate a cluster-level feature vector; the cluster-level feature vector is weightedly fused according to the number of cluster frames to generate a trajectory-level feature vector; the cluster-level feature vector and the trajectory-level feature vector are stored in a vector database according to the trajectory point sequence.

[0125] In this invention, the multi-camera matching module determines and counts the adjacent cameras of each camera according to the deployment position and field of view of different cameras, and obtains an adjacency table of each camera. Based on the three-dimensional world coordinates of each camera, the field of view intersection between cameras is calculated. The trajectory-level feature vector of each camera in the vector database is combined with the spatiotemporal constraint mechanism to perform trajectory matching and fusion of the target person.

[0126] If the timestamps of N trajectory-level feature vectors are the same and they are located in the field of view of the same camera device, then each trajectory-level feature vector is determined to belong to a different target person and cannot be merged into the same target person trajectory;

[0127] If N trajectory-level feature vectors have the same timestamp but are located in the field of view of different cameras and have no field of view intersection, then each trajectory-level feature vector is determined to belong to a different target person and cannot be merged into the same target person trajectory;

[0128] If N trajectory-level feature vectors have the same timestamp but are located in the field of view of different cameras and have overlapping fields of view, the Euclidean distance between the trajectory-level feature vectors is calculated and compared with the preset trajectory distance threshold:

[0129] If the Euclidean distance is less than the trajectory distance threshold, the current trajectory-level feature vectors are determined to belong to the same target person and are merged into the same target person trajectory;

[0130] If the Euclidean distance is greater than or equal to the trajectory distance threshold, it is determined that the current trajectory-level feature vectors belong to different target persons and cannot be merged into the same target person trajectory;

[0131] If the timestamps of N trajectory-level feature vectors are different and different camera devices are all in the adjacency table, the modal comprehensive matching degree of the multimodal feature vectors in different field of view spaces is calculated and compared with the preset matching threshold:

[0132] If the modal comprehensive matching degree is greater than or equal to the matching threshold, the current trajectory-level feature vector is determined to belong to the same target person, and is merged into the same target person trajectory. The unmatched trajectory-level feature vector and the last segment of the target person trajectory are marked to obtain the candidate trajectory to be matched;

[0133] If the modal comprehensive matching degree is less than the matching threshold, it is determined that the current trajectory-level feature vectors belong to different target persons and cannot be merged into the same target person trajectory.

[0134] In the present invention, the global retrieval module performs multimodal feature extraction on the new target trajectory to obtain a new frame-level smoothed feature vector, and generates a new cluster-level feature vector and a new trajectory-level feature vector through a moving average algorithm and a time series aggregation algorithm. The process of searching the trajectory-level feature vector in the vector database according to the multimodal multi-path recall strategy is as follows:

[0135] The new trajectory-level feature vector of each modality is decomposed to obtain several trajectory element points. At the same time, the candidate trajectories to be matched are parsed, matched and sorted with several trajectory element points to obtain the trajectory matching scores of different modalities.

[0136] The trajectory matching scores of each modality are weighted and fused according to the boundary detection box to obtain the total confidence trajectory matching score;

[0137] The new trajectory-level feature vector is verified with spatiotemporal constraints based on the total confidence trajectory matching score: if the new trajectory-level feature vector successfully matches the unmatched trajectory-level feature vector, the two trajectory-level feature vectors are merged into the same target person trajectory;

[0138] If the new trajectory-level feature vector successfully matches the last segment of the target person's trajectory, the new trajectory-level feature vector is incorporated into the matched target person's trajectory; if neither is successfully matched, the next layer of cluster-level feature vector retrieval is performed;

[0139] The process of hierarchical retrieval of cluster-level feature vectors and frame-level smooth feature vectors of the vector database by the multi-mode multi-way recall strategy is as follows:

[0140] Locally aggregate the unmatched trajectory-level feature vector and the last segment of the target person's trajectory to obtain the cluster-level features to be matched;

[0141] The new cluster-level feature vector of each modality is decomposed to obtain several cluster-level element points. At the same time, the cluster-level features to be matched are analyzed, matched and sorted with several cluster-level element points to obtain cluster-level matching scores of different modalities.

[0142] The cluster-level matching scores of each modality are weighted and fused according to the boundary detection box to obtain the total confidence cluster-level matching score;

[0143] The new cluster-level feature vector is verified with spatiotemporal constraints based on the total confidence cluster-level matching score: if the new cluster-level feature vector successfully matches the unmatched cluster-level feature vector, the two cluster-level feature vectors are merged into the same target person trajectory;

[0144] If the new cluster-level feature vector successfully matches the last segment of the target person's cluster-level features, the new cluster-level feature vector is incorporated into the matched target person's cluster-level features;

[0145] If no match is successful, the computing resource status is estimated: if the computing resources are idle, the next layer of frame-level feature vector search is performed; if the computing resources are busy, the new cluster-level feature vector is saved and marked until the computing resources are idle and the frame-level feature vector search is performed;

[0146] The cluster-level features to be matched are retrieved frame by frame to obtain the frame-level features to be matched;

[0147] The new frame-level feature vector of each modality is decomposed to obtain several frame-level element points. At the same time, the frame-level features to be matched are analyzed, matched and sorted with several frame-level element points to obtain the frame-level matching scores of different modalities.

[0148] The frame-level matching scores of each modality are weighted and fused according to the boundary detection box to obtain the total confidence frame-level matching score;

[0149] The new frame-level feature vector is verified for spatiotemporal constraints based on the total confidence frame-level matching score: if the new frame-level feature vector successfully matches the unmatched frame-level feature vector, the two frame-level feature vectors are merged into the same target person trajectory;

[0150] If the new frame-level feature vector successfully matches the last segment of the target person's frame-level features, the new frame-level feature vector is incorporated into the matched target person's frame-level features;

[0151] If no matching is successful, the new frame-level feature vector is saved and marked as an unmatched trajectory.

[0152] In the present invention, the operation process of the global search module further includes the following steps:

[0153] The time intervals between frame-level feature vectors, cluster-level feature vectors, and trajectory-level feature vectors are calculated based on the timestamps, and compared with the preset time thresholds:

[0154] If the time interval is greater than the time threshold, the time connection between the two trajectory feature vectors is judged to be unreasonable and the association is cancelled;

[0155] If the time interval is less than or equal to the time threshold, the two trajectory feature vectors are marked as trajectories to be completed;

[0156] According to the field of view space combined with the homography matrix, the spatial distance between frame-level feature vectors, cluster-level feature vectors, and trajectory-level feature vectors is calculated and compared with the preset distance threshold:

[0157] If the spatial distance is greater than the distance threshold, the spatial movement of the two trajectory feature vectors is judged to be unreasonable and the association is cancelled;

[0158] If the spatial distance is less than or equal to the distance threshold, the two trajectory feature vectors are marked as trajectories to be completed;

[0159] The completed trajectory is analyzed according to time and space to obtain time feature parameters and space feature parameters;

[0160] The temporal characteristic parameters include the trajectory start time, trajectory end time and time interval; the spatial characteristic parameters include the starting point coordinates, end point coordinates and spatial distance;

[0161] Based on the temporal and spatial feature parameters combined with the time threshold and distance threshold, the similarity of all target person trajectory segments in the vector database is calculated and compared with the preset similarity threshold:

[0162] If the similarity is greater than or equal to the similarity threshold, the pedestrian re-identification algorithm is used to match the start and end points of the completed segment;

[0163] If the similarity is less than the similarity threshold, the Bezier curve interpolation algorithm is used to generate smooth trajectory completion segments to fill the gaps between the target person's trajectory segments.

[0164] Example:

[0165] In the traditional single-camera tracking process, input image quality issues (such as blur, out-of-focus, or motion blur) often lead to reduced target detection and tracking performance. To address this issue, this module designs a fuzzy dynamic decision mechanism combined with an image enhancement judgment model:

[0166] Quickly quantify the blur level of an image to determine whether additional enhancement processing is needed. Since modern mainstream surveillance cameras typically have high resolutions, gradient analysis (Laplace transform variance) and frequency domain analysis (Fast Fourier Transform) are used to ensure processing speed.

[0167] Based on the blur detection results, we determine whether image enhancement using super-resolution techniques is necessary and whether this enhancement is likely to significantly improve detection performance. This is achieved using a deep learning prediction model: a specialized prediction model is trained, which takes a blurry image as input and outputs predictions of detection performance for that image under both direct detection and super-resolution-based detection.

[0168] Clear images and slightly blurred images: directly sent to subsequent target detection, segmentation, and tracking modules without additional processing.

[0169] Moderately blurred images: Enter the super-resolution feasibility prediction process. For images where super-resolution is expected to significantly improve detection performance, full-image super-resolution processing is performed using a lightweight model (ESPCN, FSRCNN), and then sent to the subsequent process. Images with poor super-resolution predictions are discarded.

[0170] An object detector is used to obtain the bounding box location and confidence score of pedestrian targets in each frame. An instance segmentation algorithm is applied to each detected target box to generate a corresponding segmentation mask. The segmentation mask further refines the pixel-level region of the target, especially when there is overlap between pedestrian targets. The segmentation mask can effectively distinguish the area range of different targets.

[0171] Combine the detection bounding box with the segmentation mask and use an object tracking algorithm (such as multi-object tracking based on Kalman filtering and the Hungarian algorithm) to associate the target's trajectory in consecutive frames. This improves the integrity and accuracy of the trajectory through trajectory smoothing, re-association of lost targets, and dynamic updating of the segmentation mask.

[0172] The segmented pedestrian images are provided to the multimodal feature extraction module (visual features, facial features, language descriptions, etc.) for extracting the global visual features of pedestrians. In cases where there is overlap between pedestrian targets or the background is complex, the segmentation mask can refine the effective pixel area in the detection frame, improving the accuracy of visual feature extraction. It also provides higher-quality input to the pose estimation module, helping the system to more accurately extract local information in scenes with occlusion or pedestrian overlap. The spatiotemporal feature modeling module is provided with temporal and spatial information about pedestrian motion, which is used to analyze pedestrian motion patterns and the spatiotemporal constraints associated across cameras.

[0173] In the single-camera object tracking module, although the global blur level is checked, the target frame and segmentation mask may still be blurred due to issues such as resolution, motion blur, and lighting conditions. Direct feature extraction reduces the system's recognition accuracy, necessitating a secondary clarity assessment. Because the resolution of the target frame and segmentation mask has been significantly reduced, a convolutional network-based blur detector can be used to assess clarity. When clarity falls below a set threshold, the target frame is marked as "blurred," and the super-resolution processing phase begins.

[0174] For blurred or low-resolution object frames, super-resolution technology can restore the object's clarity, enabling the model to extract higher-quality features. Unlike the lightweight models used for full-image super-resolution, target super-resolution reconstruction uses high-quality models (such as EDSR and ESRGAN) to enhance the resolution of the detection frame and segmentation mask. For target frames combined with segmentation masks, super-resolution reconstruction is performed only on the area within the mask, avoiding inefficient calculations in background areas.

[0175] When the target bounding box or segmentation mask is incomplete due to occlusion or the target is located at the edge of the image, recovering the occluded target area or estimating the target's true boundary position is particularly important for improving the integrity of feature extraction. By analyzing the context area of the target bounding box (such as adjacent frames or trajectory information), combined with the target's motion direction and trajectory prediction, the actual boundary of the target can be inferred through spatiotemporal features.

[0176] Based on human structure modeling (such as keypoint detection) or contextual reasoning (such as GAN-based mask completion), mask completion is performed on occluded or truncated areas. The existing segmentation mask and estimated object bounding box are combined to generate a complete masked area. The restored complete bounding box and segmentation mask are used for subsequent feature extraction. For incompletely displayed objects, features of the occluded area are restored as much as possible.

[0177] Extract features from the overall appearance of pedestrians. Use deep learning models (such as ResNet or VisionTransformer) to extract global feature vectors containing information such as color, texture, and clothing style. Global features can provide information about the overall appearance of pedestrians and are suitable for situations where the targets are not occluded or have little overlap. Apply the segmentation mask generated by the instance segmentation model to extract features from the valid pixel area within the target box. Combined with human posture estimation (such as HRNet), extract component-level feature vectors of pedestrians (including head, torso, limbs, etc.) to enhance the ability to capture local differences. Local features are particularly important in scenarios with occlusion or target overlap, and can avoid interference from background noise on features.

[0178] The facial region of pedestrians is detected within the target frame, and high-precision facial feature vectors are extracted using facial recognition models (RetinaFace and ArcFace). Facial features are highly discriminative and provide highly reliable identity information when the target's frontal view is visible. Based on the facial region, pedestrian attribute information (including gender, age group, and expression) is further extracted. This attribute information serves as an auxiliary feature and is combined with other modal features to enhance system robustness.

[0179] A natural language description of the person (e.g., "male in a red coat") is generated from the target frame using an image-language model (BLIP). This language description provides high-level semantic information about visual features, particularly when the target's appearance is blurry or occluded. Language features can effectively supplement the system's feature expression capabilities. The generated language description is merged with the facial attributes generated in the previous step (this step prevents the image-language model from failing to extract information such as gender and age group) and input into a language embedding model (e.g., BERT or Transformer) to extract the corresponding semantic vector. The semantic vector is jointly optimized with visual and facial features to improve cross-modal matching performance.

[0180] The timestamps of the target's first appearance and last disappearance are extracted, and the spatial position of the target entering and leaving the field of view is extracted, expressed as the pixel coordinates of the center of the target frame, and converted into world coordinates by combining the camera monocular ranging algorithm.

[0181] The target tracking module obtains a sequence of trajectory points of the target within the camera's field of view, each containing both time and spatial location. The target's complete motion path is captured to support cross-camera association and behavioral analysis. The target's motion characteristics (such as speed, direction, acceleration, etc.) are calculated based on the trajectory point sequence. All feature vectors are flattened and stored in the corresponding vector database collection. Each vector is appended with trajectory metadata constructed using spatiotemporal features, fully leveraging the performance advantages of the vector database while retaining information about each trajectory segment to facilitate intra-trajectory clustering and cross-camera aggregation. The following is an example of storing the frame-level features of the target trajectory in the vector database collection:

[0182] {"trajectory_id":"T001",#Trajectory ID; "camera_id":"C001",#Camera ID; "time":"15:30:10.033",#Frame timestamp; "position":[100,200],#Detection box center position; "global_feature":[0.12,0.45,0.67,...,0.89]#Feature vector;}

[0183] {"trajectory_id":"T001",#Trajectory ID; "camera_id":"C001",#Camera ID; "time":"15:30:10.033",#Frame timestamp; "local_features":[f_head_1,...,f_head_256,f_torso_1,...,f_torso_256,f_legs_1,...,f_legs_256],#Feature vector

[0184] "body_info":{"head":[0,255],#Index range of head features; "torso":[256,511],#Index range of torso features; "legs":[512,767]#Index range of leg features;}}

[0185] {"trajectory_id":"T001",#Trajectory ID; "camera_id":"C001",#Camera ID; "time":"15:30:10.033",#Frame timestamp; "face_feature":[0.23,0.45,0.67,...,0.89],#Feature vector; "gender":"Male",#Gender; "age_group":"Youth",#Age group; "expression":"Smile"#Expression}

[0186] {"trajectory_id":"T001",#trajectory ID;"camera_id":"C001",#camera ID;"time":"15:30:10.033",#frame timestamp;"text":"A young man wearing red clothes and sunglasses, smiling",#text information;"text_vector":[0.11,0.22,0.33,...,0.99]#feature vector};

[0187] A sliding window moving average method is used to remove local noise from the feature sequence and enhance feature stability for each frame of the trajectory (e.g., facial features, overall visual features, and local features). Multi-frame smooth features in the trajectory are temporally clustered to capture the multi-segment dynamic characteristics of the trajectory. The clustering results can reflect changes in target features at different stages (e.g., when the target enters an occluded area, when the direction of motion changes, or when lighting conditions change). Each cluster is averaged and fused to generate cluster-level features. Cluster-level features are weighted and fused according to the number of frames in each cluster to generate a trajectory-level feature vector.

[0188] The cluster-level features after clustering and the final fused trajectory-level features are stored in the vector database in the same way as the single-frame features. It should be noted that due to the inconsistent feature fluctuations of different modalities, the lengths of each cluster in each modality may not be exactly the same.

[0189] Frame overlap constraint for the same camera: If two tracks have frame overlap in the field of view of the same camera (i.e., they are captured by the camera at the same time for a part of the time), then it can be determined that the two tracks are different targets and cannot be merged into the same track.

[0190] Frame overlap constraint for disjoint camera pairs: If the fields of view of two cameras do not intersect, then they cannot see the same target in the same time period. Therefore, tracks from different cameras that appear in the same time period are likely to belong to different targets and cannot be easily merged.

[0191] The position overlap constraint for intersecting camera pairs evaluates whether the spatial positions of different tracks from two cameras with intersecting fields of view are close enough during the same time period. This constraint determines whether two tracks belong to the same person by calculating the positions of the tracks in the world coordinate system.

[0192] Neighbor Constraint: Analyzes whether an object can possibly move from one camera’s field of view to another without being captured by other cameras in between. Using the adjacency table between cameras, the system can determine whether two trajectories are likely to belong to the same object.

[0193] Based on a multi-modal and multi-path recall strategy: Multi-modal features (such as global visual features, local features, facial features, language description features, etc.) are used for layer-by-layer retrieval and matching. Each modal feature is retrieved independently, and the total matching score is finally calculated through a fusion strategy and verified in combination with spatiotemporal constraints. The basic process is as follows:

[0194] New trajectories enter the database: Multimodal frame-level features of the new trajectories are extracted, and cluster-level features and trajectory-level features are generated through moving average and time series aggregation.

[0195] Multimodal hierarchical retrieval: Retrieve layer by layer according to the priority of trajectory-level features → cluster-level features → frame-level features.

[0196] Time and space constraint verification: Perform time and space constraint verification on the matching results to ensure the logical rationality of the matching.

[0197] Track-level feature retrieval rapidly discovers potential matching global tracks: Using the track-level features of the single-camera track to be matched, two categories of candidate tracks are retrieved from the vector database: a. All single-camera tracks that have not been matched to a cross-camera track; b. The last single-camera track in all matched cross-camera tracks. Track-level features are retrieved for each modality and their corresponding matching scores are obtained. The matching scores of each modality are weighted and fused. When facial features are present, they are given a higher weight. When no facial features are present, adjustments are made based on the scene (e.g., language features are weighted more heavily in occluded scenes). The fused total matching scores are sorted from high to low. Results with high confidence scores are then subjected to spatiotemporal constraint verification in this order. If a successful match is found with the aforementioned track in category a, the two single-camera tracks are merged into a new cross-camera track. If a successful match is found with the aforementioned track in category b, the current single-camera track is incorporated into the matched cross-camera track. Once a successful match is found, the process ends; if a match fails, the process proceeds to the next level of cluster-level feature retrieval.

[0198] Cluster-level feature retrieval models and matches local features in the trajectory and is suitable for situations where trajectory-level features cannot be matched: in the vector database, the cluster-level features of the single-camera trajectory to be matched are used to retrieve all cluster-level features in the two types of trajectories mentioned above, a and b. The cluster-level features of each modality are retrieved separately and the corresponding matching scores are obtained. As described above, the fusion and sorting are performed. The results with higher confidence scores are verified for spatiotemporal constraints in order. The rules are as described above. Once the match is successful, the process ends; if the match fails, three strategies are selected based on the current availability of computing resources: 1. Transition to the next layer of frame-level feature retrieval. 2. Save the existing results and perform frame-level feature retrieval when computing power is idle. 3. Mark the current single-camera trajectory as unmatched and wait for subsequent association.

[0199] Frame-level feature retrieval is the final retrieval layer, used to search for matching candidate trajectories frame by frame. In the vector database, the frame-level features of the single-camera trajectory to be matched are used to retrieve all frame-level features from the two aforementioned trajectories, a and b. Frame-level features are retrieved separately for each modality, and the corresponding matching scores are obtained. These features are then fused and sorted as described above. Results with high confidence scores are then sequentially validated against spatiotemporal constraints. The rules are as described above: once a match is successful, the process ends. If the frame-by-frame search still fails to match the current trajectory, the current single-camera trajectory is marked as an unmatched trajectory, pending subsequent association.

[0200] Spatiotemporal consistency check: Cross-camera trajectory associations must be consistent in both time and space. Spatiotemporal consistency check is the first step in spatiotemporal optimization. It verifies that the current trajectory association results are logical and eliminates unreasonable associations.

[0201] Temporal continuity check: Ensures that the target trajectory's temporal connection between different cameras is reasonable. The time interval between two trajectories is calculated. If the interval is too long (adjusted based on the scene), the temporal relationship between the two trajectories is considered unreasonable and the connection is disassociated. If the target has a long gap in time but is still plausible (for example, if the target is briefly occluded or leaves and then re-enters another camera), it is marked as a "pending trajectory" and the trajectory completion phase begins.

[0202] Spatial plausibility check: Ensures that the spatial movement of the target trajectory between different cameras complies with physical laws. The pixel coordinates of the trajectory are converted to a unified world coordinate system using a homography matrix. The actual spatial distance of the trajectory is calculated. If the distance is greater than a specified threshold (determined by the scene), the spatial movement of the two trajectories is considered unreasonable and the association is disassociated.

[0203] Spatiotemporal trajectory completion: Due to camera coverage limitations, object occlusion, or detection failure, there may be gaps in time or space between trajectories. Spatiotemporal trajectory completion aims to fill these gaps by generating smooth trajectory completion segments through interpolation methods. It searches for candidate trajectories that are temporally and spatially plausible for the missing trajectory end points. Trajectories with close temporal and spatial distances are retrieved from the database as candidates: a. The time interval between the candidate trajectory's start time and the current trajectory's end time must be within a threshold; b. The spatial distance between the candidate trajectory's start point and the current trajectory's end point must be less than a threshold. Feature consistency is then checked, and the start and end points of the completion segment are matched using the ReID feature vector. If feature consistency is low, completion is abandoned.

[0204] Trajectory interpolation: If the existing trajectory cannot be used to complete the trajectory, a smooth trajectory completion segment is generated to fill the gaps between the trajectories. Linear interpolation or Bezier curve interpolation can be used.

Claims

1. A multi-target person re-identification system based on a multimodal and vector database, characterized by: It includes a monocular tracking module, which is used to locate and track the target person in each frame of the original image based on a fuzzy dynamic decision mechanism based on quantization threshold and fuzzy threshold combined with an image enhancement judgment model, and generate multimodal information; The multimodal extraction module is used to process multimodal information according to the quality reconstruction judgment mechanism, extract and fuse multimodal features, and store them in the vector database; The trajectory generation module is used to extract trajectory-level feature vectors of multimodal features based on the temporal motion trajectory generation mechanism; The multi-camera matching module is used to match the trajectory-level feature vectors of different cameras according to the spatiotemporal constraint mechanism to generate the target person's trajectory; The global retrieval module is used to search and optimize in the vector database based on the multi-modal and multi-path recall strategy and the new target trajectory features.

2. The multi-target person re-identification system based on multimodal and vector database according to claim 1, characterized in that: The multimodal information includes target trajectory data, boundary detection boxes, and segmentation masks; the monocular tracking module converts the videos captured by all cameras into several frames of original images; the blurriness of each frame of the original image is detected and quantified according to the fuzzy dynamic decision mechanism, and the blurriness is compared with a preset blur threshold to obtain a clear image and a blurred image, and then the blurred image is enhanced to obtain an enhanced image; Use the target detector to detect each target person in the enhanced image and the clear image, and obtain the boundary detection box and confidence score; The boundary detection box is segmented to generate a segmentation mask; the boundary detection box is combined with the segmentation mask to associate and track the trajectory of the target person in the continuous frame original image to obtain the target trajectory data and assign a unique identifier.

3. The multi-target person re-identification system based on multimodal and vector database according to claim 1, characterized in that: The operation process of the fuzzy dynamic decision-making mechanism is as follows: The gradient analysis algorithm is combined with the frequency domain analysis algorithm to quantify the blur degree of each frame of the original image and obtain the image blur; Compare the image blur with the preset blur threshold: if the image blur is less than the blur threshold, the current image is determined to be a clear image, and target detection, target segmentation and target tracking operations are directly performed; If the image blur is greater than or equal to the blur threshold, the current image is judged to be a blurred image, and the feasibility prediction is input into the quantization judgment layer of the image enhancement judgment model to obtain the feasibility quantization value, which is then compared with the preset quantization threshold: If the feasibility quantization value is less than the quantization threshold, the image completion layer of the input image enhancement judgment model is combined with the unique identifier and timestamp, and the generative adversarial network is used to enhance the current image to obtain a clear image for target detection, target segmentation and target tracking operations, while the current image is discarded. If the feasibility quantization value is greater than or equal to the quantization threshold, it is input into the feature extraction layer of the image enhancement judgment model, and the fuzzy image is extracted with the principal component analysis algorithm to obtain the global fuzzy matrix; The enhanced feedback layer of the image enhancement decision model is combined with the PID algorithm to perform resolution adaptation on the global fuzzy matrix to generate a super-resolution feedback matrix; The iterative training layer of the image enhancement decision model combines the ESPCN network and the FSRCNN network to perform several iterative training on the super-resolution feedback matrix to obtain the full-image super-resolution weight matrix; The prediction decision layer of the image enhancement judgment model adopts the divergence-cross entropy loss function to update and optimize the full-image super-resolution weight matrix and output the enhanced image at the same time.

4. The multi-target person re-identification system based on multimodal and vector database according to claim 1, characterized in that: The multimodal extraction module preprocesses the multimodal information according to the quality reconstruction judgment mechanism: The clarity of the boundary detection box and segmentation mask is detected and compared with the preset clarity threshold: If the clarity of the boundary detection box or segmentation mask is less than the clarity threshold, local super-resolution reconstruction is performed on the current boundary detection box or segmentation mask; If the clarity of the bounding box or segmentation mask is greater than or equal to the clarity threshold, the integrity of the bounding box and segmentation mask is checked: If the boundary detection box is incomplete, the current boundary detection box is estimated according to the timestamp and background space based on the adjacent frame image or target trajectory data, combined with the target motion direction and trajectory prediction, to obtain the actual boundary box; If the segmentation mask is incomplete, the current segmentation mask is restored and completed based on the human body key points and the actual bounding box to obtain a complete mask, thereby completing the multimodal information preprocessing.

5. The multi-target person re-identification system based on multimodal and vector database according to claim 1, characterized in that: The multimodal extraction module uses a visual transformer combined with an actual bounding box to extract global visual features of the target person and obtain an overall appearance vector; At the same time, a human pose estimator is used in combination with a complete mask to extract local visual features of the target person and obtain a component-level feature vector; The face recognition algorithm is used in combination with the actual bounding box to detect the face of the target person and obtain the facial feature vector and target attribute information; Use the image-language model to describe the actual bounding box in language and obtain language feature information; The language embedding model is used to combine the language feature information with the target attribute information to obtain the semantic vector; A multimodal fusion algorithm is used to jointly optimize the semantic vector with the overall appearance vector, component-level feature vector, and facial feature vector to obtain a multimodal feature vector. The spatial position of the target person entering and leaving the camera's field of view is extracted to obtain the center coordinates of the target detection frame. Combined with the monocular ranging algorithm, it is converted into 3D world coordinates, and the appearance and disappearance timestamps are recorded at the same time. At the same time, the trajectory points of the target person in the field of view of the camera device are captured to obtain a trajectory point sequence, and then the motion characteristics of the target person are calculated; the motion characteristics include speed, direction and acceleration; The multimodal feature vectors are fused and appended with the trajectory point sequence respectively to be stored in the corresponding vector database.

6. The multi-target person re-identification system based on multimodal and vector database according to claim 1, characterized in that: The trajectory generation module uses a sliding window moving average algorithm to perform local denoising on the multimodal feature vector of each frame according to the temporal motion trajectory generation mechanism to obtain a frame-level smoothed feature vector; The frame-level smooth feature vectors are temporally clustered according to the appearance timestamp and disappearance timestamp to generate several modal clusters; each modal cluster is averaged and fused to generate a cluster-level feature vector; Perform weighted fusion on cluster-level feature vectors according to the number of cluster frames to generate trajectory-level feature vectors; The cluster-level feature vectors and trajectory-level feature vectors are stored in the vector database according to the trajectory point sequence.

7. The multi-target person re-identification system based on multimodal and vector database according to claim 1, characterized in that: The multi-camera matching module determines and counts the adjacent cameras of each camera based on the deployment position and field of view of different cameras to obtain an adjacency table for each camera. Based on the three-dimensional world coordinates of each camera, the field of view intersection between cameras is calculated. The trajectory-level feature vectors of each camera in the vector database are combined with the spatiotemporal constraint mechanism to perform trajectory matching and fusion of the target person. If the timestamps of N trajectory-level feature vectors are the same and they are located in the field of view of the same camera device, then each trajectory-level feature vector is determined to belong to a different target person and cannot be merged into the same target person trajectory; If N trajectory-level feature vectors have the same timestamp but are located in the field of view of different cameras and have no field of view intersection, then each trajectory-level feature vector is determined to belong to a different target person and cannot be merged into the same target person trajectory; If N trajectory-level feature vectors have the same timestamp but are located in the field of view of different cameras and have overlapping fields of view, the Euclidean distance between the trajectory-level feature vectors is calculated and compared with the preset trajectory distance threshold: If the Euclidean distance is less than the trajectory distance threshold, the current trajectory-level feature vectors are determined to belong to the same target person and are merged into the same target person trajectory; If the Euclidean distance is greater than or equal to the trajectory distance threshold, it is determined that the current trajectory-level feature vectors belong to different target persons and cannot be merged into the same target person trajectory; If the timestamps of N trajectory-level feature vectors are different and different camera devices are all in the adjacency table, the modal comprehensive matching degree of the multimodal feature vectors in different field of view spaces is calculated and compared with the preset matching threshold: If the modal comprehensive matching degree is greater than or equal to the matching threshold, the current trajectory-level feature vector is determined to belong to the same target person, and is merged into the same target person trajectory. The unmatched trajectory-level feature vector and the last segment of the target person trajectory are marked to obtain the candidate trajectory to be matched; If the modal comprehensive matching degree is less than the matching threshold, it is determined that the current trajectory-level feature vectors belong to different target persons and cannot be merged into the same target person trajectory.

8. The multi-target person re-identification system based on multimodal and vector database according to claim 1, characterized in that: The global retrieval module performs multimodal feature extraction on the new target trajectory to obtain a new frame-level smoothed feature vector, and generates a new cluster-level feature vector and a new trajectory-level feature vector. The process of searching the trajectory-level feature vector in the vector database according to the multimodal multi-path recall strategy is as follows: The new trajectory-level feature vector of each modality is decomposed to obtain several trajectory element points. At the same time, the candidate trajectories to be matched are analyzed, matched and sorted with several trajectory element points to obtain the trajectory matching scores of different modalities. The trajectory matching scores of each modality are weighted and fused according to the boundary detection box to obtain the total confidence trajectory matching score; The new trajectory-level feature vector is verified with spatiotemporal constraints based on the total confidence trajectory matching score: if the new trajectory-level feature vector successfully matches the unmatched trajectory-level feature vector, the two trajectory-level feature vectors are merged into the same target person trajectory; If the new trajectory-level feature vector successfully matches the last segment of the target person's trajectory, the new trajectory-level feature vector will be incorporated into the matched target person's trajectory; if neither is successfully matched, the next layer of cluster-level feature vector retrieval will be performed.

9. The multi-target person re-identification system based on multimodal and vector database according to claim 8, characterized in that: The process of the multi-mode multi-path recall strategy for hierarchical retrieval of cluster-level feature vectors and frame-level smooth feature vectors of the vector database is as follows: Locally aggregate the unmatched trajectory-level feature vector and the last segment of the target person's trajectory to obtain the cluster-level features to be matched; The new cluster-level feature vector of each modality is decomposed to obtain several cluster-level element points. At the same time, the cluster-level features to be matched are analyzed, matched and sorted with several cluster-level element points to obtain cluster-level matching scores of different modalities. The cluster-level matching scores of each modality are weighted and fused according to the boundary detection box to obtain the total confidence cluster-level matching score; The new cluster-level feature vector is verified with spatiotemporal constraints based on the total confidence cluster-level matching score: if the new cluster-level feature vector successfully matches the unmatched cluster-level feature vector, the two cluster-level feature vectors are merged into the same target person trajectory; If the new cluster-level feature vector successfully matches the last segment of the target person's cluster-level features, the new cluster-level feature vector is incorporated into the matched target person's cluster-level features; If no match is successful, the computing resource status is estimated. If the computing resource is idle, the next step is frame-level feature vector retrieval. If the computing resources are busy, the new cluster-level feature vector is saved and marked until the frame-level feature vector is retrieved when the computing resources are idle; The cluster-level features to be matched are retrieved frame by frame to obtain the frame-level features to be matched; The new frame-level feature vector of each modality is decomposed to obtain several frame-level element points. At the same time, the frame-level features to be matched are analyzed, matched and sorted with several frame-level element points to obtain the frame-level matching scores of different modalities. The frame-level matching scores of each modality are weighted and fused according to the boundary detection box to obtain the total confidence frame-level matching score; The new frame-level feature vector is verified for spatiotemporal constraints based on the total confidence frame-level matching score: if the new frame-level feature vector successfully matches the unmatched frame-level feature vector, the two frame-level feature vectors are merged into the same target person trajectory; If the new frame-level feature vector successfully matches the last segment of the target person's frame-level features, the new frame-level feature vector is incorporated into the matched target person's frame-level features; If no matching is successful, the new frame-level feature vector is saved and marked as an unmatched trajectory.

10. The multi-target person re-identification system based on multimodal and vector database according to claim 1, characterized in that: The operation process of the global search module further includes the following steps: The time intervals between frame-level feature vectors, cluster-level feature vectors, and trajectory-level feature vectors are calculated based on the timestamps, and compared with the preset time thresholds: If the time interval is greater than the time threshold, the time connection between the two trajectory feature vectors is judged to be unreasonable and the association is cancelled; If the time interval is less than or equal to the time threshold, the two trajectory feature vectors are marked as trajectories to be completed; According to the field of view space combined with the homography matrix, the spatial distance between frame-level feature vectors, cluster-level feature vectors, and trajectory-level feature vectors is calculated and compared with the preset distance threshold: If the spatial distance is greater than the distance threshold, the spatial movement of the two trajectory feature vectors is judged to be unreasonable and the association is cancelled; If the spatial distance is less than or equal to the distance threshold, the two trajectory feature vectors are marked as trajectories to be completed; The completed trajectory is analyzed according to time and space to obtain time feature parameters and space feature parameters; Based on the temporal and spatial feature parameters combined with the time threshold and distance threshold, the similarity of each segment of all target person trajectories in the vector database is calculated and compared with the preset similarity threshold: If the similarity is greater than or equal to the similarity threshold, the pedestrian re-identification algorithm is used to match the start and end points of the completed segment; If the similarity is less than the similarity threshold, a smooth trajectory completion segment is generated to fill the gaps between the target person's trajectory segments.

Citation Information

Patent Citations

  • Object recognition methods based on multimodal features

    CN117238045B

  • Method and system for realizing personalized recommendation based on multi-feature fusion face recognition

    CN119397072A

  • Object tracking method and device, storage medium and electronic device

    CN110443828A

  • Intelligent monitoring system and method based on identity recognition and cross-camera target tracking

    CN114693746A

  • Pedestrian re-identification method, system, medium and device based on top angle shooting

    CN116189236A

Cited By

  • YOLO-based sorting video identification processing method and system

    CN120913131A

  • Gait recognition method, device and equipment and storage medium

    CN121214552A

  • Cross-device multi-target tracking method and system

    CN121304723A

  • A multi-target tracking method and system across devices

    CN121304723B

  • Pedestrian cross-border head tracking method and system based on multi-modal dynamic feature fusion

    CN121330611A