A Real-Time Target Detection and Tracking System for Video Streams Based on Deep Learning
By employing a closed-loop optimization mechanism encompassing adversarial generation, feature alignment, adaptive detection, and spatiotemporal graph tracking modules, the problem of inaccurate feature mapping in deep learning video stream systems under dynamic scenarios is resolved, thereby enhancing the generalization ability and robustness of target detection and tracking.
Patent Information
- Application Number
- CN202511044863.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing deep learning-based real-time target detection and tracking systems for video streams suffer from insufficient model generalization ability and feature extraction bias in dynamic scenes due to limitations in training data distribution. This makes them unable to effectively identify damaged variants in tooling recognition and pedestrian detection.
Low-frequency samples are synthesized through an adversarial generation module, features are optimized by selecting either quantum compression or manifold projection algorithms through a feature alignment module, lightweight convolution processing is switched through an adaptive detection module, trajectory sequences are constructed through a spatiotemporal graph tracking module, and historical states and channel attention weights are fused through a trusted decision module to form a closed-loop optimization mechanism that dynamically calibrates the feature space mapping.
It improves the system's generalization ability and robustness in dynamic environments, effectively overcomes the feature mapping inaccuracy problem caused by the limitation of training data, and ensures the accuracy and stability of target detection and tracking.
Smart Images

Figure CN120564107B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a real-time target detection and tracking system for video streams based on deep learning. Background Technology
[0002] A deep learning-based real-time object detection and tracking system for video streams processes continuous video frame sequences using deep neural network models, achieving high-precision recognition and continuous object monitoring. In video stream applications, the input data is large and time-sensitive. Convolutional neural network architectures can directly extract semantic information from pixel-level features, quickly locating and classifying multiple target objects. Subsequently, the tracking module uses an association matching algorithm to correlate detection results across frames, maintaining object trajectories and supporting dynamic scene analysis. During training, deep learning methods utilize large-scale labeled data to optimize model parameters and enhance robustness, while hardware acceleration, such as GPU parallel computing, further improves processing speed to meet real-time requirements.
[0003] In practical applications, real-time target detection and tracking systems for video streams often suffer from insufficient generalization ability in dynamic scenarios due to limitations in the distribution of training data. This manifests as deep learning models failing to effectively learn the long-tail distribution characteristics of data in the actual deployment environment, leading to biased feature extraction. For example, in smart venue monitoring scenarios, variations in the color, texture, or damage of work clothes exceed the coverage of the model's training samples, causing pedestrian detection and workwear recognition models to fail in feature space mapping, resulting in missed detection of non-compliant attire and highlighting the generalization gap between the training set and real dynamic data. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a real-time target detection and tracking system for video streams based on deep learning, solving the technical problem of inaccurate feature space mapping caused by differences in the long-tail distribution of training data.
[0005] To solve the above-mentioned technical problems, the specific technical solution of the present invention is as follows:
[0006] This invention provides a real-time target detection and tracking system for video streams based on deep learning, comprising:
[0007] The data acquisition module is configured to capture the raw video stream and output a bimodal image sequence to the adversarial generation module;
[0008] The adversarial generation module is configured to receive the bimodal image sequence and specific spatial coordinates fed back by the adaptive detection module, synthesize low-frequency samples based on the specific spatial coordinates, and output the synthesized augmented sample library to the feature alignment module.
[0009] The feature alignment module is configured to receive the augmented sample library, extract the depth features of the augmented sample library, select a quantum compression algorithm or a manifold projection algorithm for feature optimization processing according to the real-time computing resource load, generate channel attention weights and fuse features based on the weights to obtain fused features, and output the fused features to the adaptive detection module.
[0010] The adaptive detection module is configured to process the fusion features using lightweight convolution through the main path based on the scene complexity index of the fusion features, generate target detection boxes and output the target detection boxes to the spatiotemporal graph tracking module, and activate the re-identification sub-network through the auxiliary path in response to the conflict node identifier, generate supplementary detection boxes and output the supplementary detection boxes to the spatiotemporal graph tracking module.
[0011] The spatiotemporal graph tracking module is configured to receive the target detection box and the supplementary detection box, embed the target detection box and the supplementary detection box into spatiotemporal nodes to construct a trajectory sequence, output the trajectory sequence and the conflict node identifier to the trusted decision module, and simultaneously feed back the conflict node identifier to the adaptive detection module as a re-identification trigger signal.
[0012] The trusted decision module is configured to receive the trajectory sequence and the channel attention weights generated by the feature alignment module, fuse the historical state of the trajectory sequence with the channel attention weights, generate a probabilistic behavior decision, and output the probabilistic behavior decision to the display terminal.
[0013] Furthermore, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the adversarial generation module is configured as follows:
[0014] Receive specific spatial coordinates from the adaptive detection module and channel attention weights from the feature alignment module;
[0015] Multiply the specific spatial coordinates by the channel attention weights to generate a sample rendering intensity factor;
[0016] The discriminator calculates the Euclidean distance between the generated sample and the source image in the feature space. The Euclidean distance is used to constrain the gradient descent direction of the generative adversarial network. The augmented sample is then output to the feature alignment module to update the feature distribution of the augmented sample library.
[0017] Furthermore, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the feature alignment module is configured as follows:
[0018] The loss gradient is received from the discriminator of the adversarial generation module, and the Riemannian manifold space mapping parameters are initialized using the loss gradient;
[0019] In response to the low-confidence region markers output by the adaptive detection module, the distribution adaptation phase is executed by calculating the difference in feature distribution between the source and target domains based on the Wasserstein distance. The channel attention weights are then dynamically adjusted according to the magnitude of the difference in feature distribution between the source and target domains, thereby increasing the channel attention weights of the regions corresponding to the low-confidence region markers by 20% to 50%.
[0020] The channel attention weight parameters are optimized through gradient backpropagation to generate optimized channel attention weights. The optimized channel attention weights are then output to the adaptive detection module to update the convolution kernel parameters of the adaptive detection module.
[0021] Furthermore, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the adaptive detection module is configured as follows:
[0022] Receive channel attention weights from the feature alignment module and conflict node identifiers from the spatiotemporal graph tracking module;
[0023] Calculate the scene complexity index based on the channel attention weights and conflict node identifiers;
[0024] If the scene complexity index exceeds a preset threshold for three consecutive frames, the redirection trigger condition is activated.
[0025] In response to the redirection triggering condition, a focus area redirection instruction is sent to the re-identification sub-network.
[0026] Furthermore, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the spatiotemporal graph tracking module is configured as follows:
[0027] The system receives channel attention weights from the feature alignment module as a weighting benchmark for appearance similarity calculation, and receives target detection boxes and supplementary detection boxes from the adaptive detection module, parsing target location information and feature descriptors.
[0028] The channel attention weight is applied to enhance the spatial flow appearance similarity calculation, generate weighted similarity data, extract the feature vector of the occluded node from the received detection box, and perform cross-modal matching between the extracted feature vector and the augmented sample library of the adversarial generation module. The matching process is performed in parallel in the HSV color space and the Lab color space. If the matching similarity exceeds 85%, the conflict node attributes are corrected and an updated trajectory sequence is generated.
[0029] The updated trajectory sequence is output to the trusted decision module for behavioral decision analysis. If the matching fails, a feature weight recalibration request is sent to the feature alignment module.
[0030] Furthermore, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the trusted decision module is configured as follows:
[0031] The updated trajectory sequence is received from the spatiotemporal map tracking module to obtain real-time target motion state data. Attenuation coefficient channel data is obtained from the feature alignment module. The attenuation coefficient is negatively correlated with the structural similarity index value output by the data acquisition module. The prior probability of historical state is read from the internal trajectory historical state database.
[0032] The trajectory continuity coefficient is calculated based on the attenuation coefficient channel data to quantify trajectory stability. The trajectory continuity coefficient is used to weight the prior probabilities of historical states. The dynamic threshold changes set by the adaptive detection module are monitored in real time. In response to the scene complexity index exceeding the dynamic threshold, the weight of the historical prior influence is compressed by 40% to 60%.
[0033] By integrating weighted probability and real-time status data, a dynamically calibrated posterior probability distribution is generated and output to the display terminal to drive the visual behavior analysis interface.
[0034] Furthermore, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the data acquisition module is configured as follows:
[0035] The original video stream is separated into visible light band data and near-infrared band data by a beam splitter prism. The structural similarity index of the visible light band data and near-infrared band data is calculated in real time to evaluate the alignment quality of the visible light band data and near-infrared band data.
[0036] In response to a structural similarity index below 0.8, a cross-spectral compensation request instruction is sent to the adversarial generation module. The cross-spectral compensation request instruction includes a real-time frame structural similarity index value and a spectral channel identifier. The cross-spectral compensation sample synthesized by the adversarial generation module is received. The cross-spectral compensation sample fuses texture features from visible light and infrared spectra.
[0037] The cross-spectral compensation samples are frame-aligned and fused with the original band data to generate a spatiotemporally aligned bimodal sequence. The frame-aligned bimodal sequence is then output to the feature alignment module for feature distribution optimization of the augmented sample library.
[0038] Furthermore, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the re-identification sub-network is configured as follows:
[0039] The cross-modal matching compensation features are received from the spatiotemporal graph tracking module. The cross-modal matching compensation features include occlusion region repair data and spatial coordinate identifiers in the Lab color space. The entangled state vector output by the quantum compression algorithm is called from the feature alignment module to obtain the entanglement association information in the feature space.
[0040] Real-time monitoring of scene complexity index changes; when the moving average of the index increases by 10% or more over 5 consecutive frames, switch to manifold projection algorithm to perform re-identification operation; under the manifold projection algorithm path, use the shortest path algorithm of Riemannian manifold space to realize geodesic distance measurement; perform feature matching based on geodesic distance to generate cross-modal feature similarity evaluation results.
[0041] Based on the feature matching results, an optimized supplementary detection box is generated and output to the corresponding node of the spatiotemporal graph tracking module to update the trajectory of the conflict node.
[0042] Furthermore, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the trusted decision module is configured as follows:
[0043] The structural similarity index value is obtained in real time from the data acquisition module. The feature confidence decay coefficient is calculated based on the structural similarity index value. The output range of the feature confidence decay coefficient is 0.2 to 1.0. The historical prior influence weight compression status is read from the internal weight compression record database.
[0044] Continuously monitor the compression magnitude of historical prior influence weights and the structural similarity index value. When the compression magnitude exceeds 40% and the structural similarity index value of two consecutive frames is lower than 0.8, calculate the average index value of the two consecutive frames. Based on the difference between 0.8 and the average index value of the two consecutive frames, multiply by 0.35 to calculate the probability cumulative threshold reduction ratio.
[0045] The probability accumulation threshold is lowered according to the probability accumulation threshold reduction ratio, an adaptive alarm trigger command including the target identifier and confidence deviation value is generated, the alarm trigger command is output to the display terminal, and the high priority alarm interface is driven.
[0046] Furthermore, the deep learning-based real-time target detection and tracking system for video streams described in this invention also includes:
[0047] The data acquisition module captures the internal fiber texture of the dark tooling in the near-infrared band and outputs a dual-modal sequence including fiber texture features to the adversarial generation module. The adversarial generation module synthesizes reflective samples of the tooling and outputs them to the feature alignment module.
[0048] The feature alignment module receives reflective samples and maps them to the Riemannian manifold space. Based on the Wasserstein distance distribution of fiber texture features in the augmented sample library, it filters out damaged feature clusters with a distance less than 0.3. The trustworthy decision module calls the cuff region trajectory continuity coefficient output by the spatiotemporal graph tracking module.
[0049] When the continuity coefficient of the cuff area trajectory is less than 0.5, a weighted constraint is applied to the wear probability. The weighted constraint amplitude increases with the wear probability value, generating a progressive wear probability report, including cuff area coordinates and historical wear trend curves. The progressive wear probability report is output to the display terminal to drive the tooling life assessment interface.
[0050] Beneficial effects of this invention;
[0051] The beneficial effects of this invention lie in the dynamic calibration of feature space mapping through a closed-loop optimization mechanism. After the data acquisition module captures a dual-modal image sequence, the adversarial generation module synthesizes low-frequency samples based on specific spatial coordinates from adaptive detection feedback to compensate for the lack of long-tail distribution in the training data. The feature alignment module selects quantum compression or manifold projection algorithms to optimize feature distribution based on real-time resource load and dynamically adjusts the difference between the source and target domains using channel attention weights. The adaptive detection module switches between primary and secondary paths based on the scene complexity index, processes conventional scenes through lightweight convolution, and activates the re-identification sub-network to handle conflict regions. The spatiotemporal graph tracking module constructs trajectory sequences and corrects node attributes based on channel weight-weighted cross-modal matching. The reliable decision-making module integrates trajectory history and feature reliability decay coefficients to generate probabilistic decisions. The entire system forms a collaborative closed loop of coordinate feedback-driven sample synthesis, channel weight adjustment and distribution calibration, and trajectory continuity constraint decision reliability, continuously compressing the difference between training and actual scene distribution, improving the generalization ability and robustness of target detection and tracking in dynamic environments, and effectively overcoming the feature mapping inaccuracy problem caused by the limitations of training data. Attached Figure Description
[0052] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on the drawings without creative effort.
[0053] Figure 1 This is a system architecture diagram of a real-time target detection and tracking system for video streams based on deep learning, provided as an embodiment of the present invention. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. The technical solutions provided by various embodiments of this invention will be described in detail below with reference to the accompanying drawings. To better understand the objectives of this invention, it will be described in further detail below.
[0055] Please see Figure 1 This invention provides a real-time target detection and tracking system for video streams based on deep learning, comprising:
[0056] The data acquisition module is configured to capture the raw video stream and output a bimodal image sequence to the adversarial generation module;
[0057] The data acquisition module is configured with a beam splitter prism assembly to separate the input video stream into visible and near-infrared spectral data streams. The module performs real-time dual-spectral frame alignment, calculating the structural similarity index parameter between corresponding visible and near-infrared frames to quantify the degree of matching between brightness contrast and structural features across spectral channels. When the structural similarity index falls below a preset quality benchmark, the module sends a cross-spectral compensation request command to the adversarial generation module. This command includes an abnormal spectral channel identifier and real-time frame deviation parameters.
[0058] After responding to the compensation request, the adversarial generation module uses a conditional generative adversarial network to synthesize cross-spectral compensated samples. The synthesis process integrates visible light channel color attributes with near-infrared channel texture details, achieving complementary reconstruction of multispectral information in the feature space. Once the generated compensated samples are returned to the data acquisition module, the module performs multi-source data spatial registration: a scale-invariant feature transformation algorithm is used to extract feature point sets from the compensated samples and the original spectral data; pixel-level coordinate mapping is achieved through an affine transformation matrix, completing the spatiotemporally synchronized construction of a dual-modal sequence.
[0059] The optimized bimodal sequence is transmitted to the adversarial generation module. The sequence includes enhanced texture features from the spectrally compensated region. Visible light data retains color information under natural lighting, while near-infrared data enhances the internal structural features of dark materials. The spatiotemporal alignment of the two channels provides the basic input for subsequent feature optimization. Through a closed-loop spectral quality monitoring and adaptive compensation mechanism, the system effectively suppresses spectral distortion caused by abrupt changes in ambient lighting.
[0060] The adversarial generation module is configured to receive the bimodal image sequence and specific spatial coordinates fed back by the adaptive detection module, synthesize low-frequency samples based on the specific spatial coordinates, and output the synthesized augmented sample library to the feature alignment module.
[0061] The adversarial generation module receives a dual-modal image sequence from the data acquisition module. This sequence contains frame data aligned to the visible and near-infrared bands. Simultaneously, the module acquires specific spatial coordinate data fed back by the adaptive detection module. This coordinate data identifies the locations of image regions with low detection confidence. The module performs a dot product operation between the specific spatial coordinate matrix and the channel attention weight matrix transmitted by the feature alignment module to generate a sample rendering intensity factor matrix. This matrix quantifies the priority weights for sample synthesis at different spatial locations.
[0062] The module inputs the sample rendering intensity factor into the generator branch of the generative adversarial network, guiding the network's computational resources to focus on the feature synthesis process in low-confidence regions. The discriminator branch calculates the Euclidean distance between the generated sample feature vector and the source image feature vector, using this distance as the core constraint of the loss function to control the direction of the backpropagation gradient update in the generative adversarial network.
[0063] During iterative training, the generator dynamically adjusts network parameters based on the gradient direction, synthesizing an enhanced sample set that covers low-frequency features. The newly synthesized sample set is integrated into the augmented sample library and output to the feature alignment module. By continuously injecting synthetic samples optimized for weak regions, the sample library gradually reduces the difference in feature distribution between the training data and the actual scene, forming an adaptive calibration mechanism for the feature space.
[0064] The feature alignment module is configured to receive the augmented sample library, extract the depth features of the augmented sample library, select a quantum compression algorithm or a manifold projection algorithm for feature optimization processing according to the real-time computing resource load, generate channel attention weights and fuse features based on the weights to obtain fused features, and output the fused features to the adaptive detection module.
[0065] The feature alignment module receives augmented sample library data from the adversarial generation module. This sample library contains a set of directionally synthesized low-frequency feature samples. The module extracts multi-level deep features from the sample library using a deep convolutional neural network, including shallow texture features and high-level semantic features. The feature extraction process employs a multi-scale convolutional kernel structure to capture feature representation patterns at different granularities.
[0066] Based on the real-time monitoring of computational resource load parameters, the module dynamically selects the feature optimization algorithm path: when the load is below a set threshold, the quantum compression algorithm is activated to achieve dimensionality reduction of high-dimensional features through quantum entangled state mapping; when the load exceeds the threshold, the module switches to the manifold projection algorithm to preserve the feature manifold structure using Riemannian geometry. This algorithm selection mechanism achieves a balance between computational efficiency and feature fidelity.
[0067] The module generates channel attention weight parameters, which are generated by a feature channel importance evaluation network. The evaluation network analyzes the contribution distribution of each feature channel in the object detection task and outputs a weight coefficient matrix. Based on these weight coefficients, multi-level deep features are weighted and fused to highlight the expression intensity of key feature channels and suppress redundant feature interference.
[0068] The fused feature data is output to the adaptive detection module. This feature data integrates the optimized distribution characteristics of the augmented sample library with discriminative features weighted by channel attention. Through a dynamic calibration mechanism between feature distribution and detection requirements, the generalization ability of subsequent detection modules in complex scenarios is improved, forming a closed-loop optimization path.
[0069] The adaptive detection module is configured to process the fusion features using lightweight convolution through the main path based on the scene complexity index of the fusion features, generate target detection boxes and output the target detection boxes to the spatiotemporal graph tracking module, and activate the re-identification sub-network through the auxiliary path in response to the conflict node identifier, generate supplementary detection boxes and output the supplementary detection boxes to the spatiotemporal graph tracking module.
[0070] The adaptive detection module receives fused feature data transmitted from the feature alignment module. This data contains multi-level feature representations weighted by channel attention weights. The module analyzes the fused features through a scene complexity index calculation unit, which calculates the entropy change rate of the fused channel attention weights and the spatial distribution density of conflict nodes fed back by the spatiotemporal graph tracking module. The calculation results quantify the intensity of environmental interference and the complexity of target interaction, forming the basis for path switching decisions.
[0071] The main path of the module design handles routine detection tasks, employing a lightweight convolutional neural network architecture to process fused features. Lightweight convolutional layers use depthwise separable kernels to reduce the number of parameters, and multi-scale semantic information is fused through a feature pyramid structure to output object detection box coordinates and classification confidence. Detection box data is transmitted in real-time to the spatiotemporal graph tracking module to construct basic trajectory nodes.
[0072] When the scene complexity index continuously exceeds a preset threshold, the module activates the auxiliary path processing mechanism. The auxiliary path responds to the conflict node identification signal transmitted by the spatiotemporal graph tracking module. This identification signal reflects the target ambiguity region during the trajectory association process. The module sends a focus region redirection instruction to the re-identification sub-network. The instruction contains the set of spatial coordinates of the conflict nodes and the feature channel identifier.
[0073] The re-identification sub-network focuses on the target region according to instructions and generates supplementary detection boxes through a high-resolution feature extraction network. The supplementary box data contains refined location information and feature descriptors of the occluded target, and is output to the spatiotemporal graph tracking module to correct the attributes of conflict nodes. A dual-path collaborative mechanism realizes dynamic scheduling of detection resources, with the main path ensuring processing efficiency in normal scenes and the auxiliary path enhancing the target capture capability in complex scenes.
[0074] The spatiotemporal graph tracking module is configured to receive the target detection box and the supplementary detection box, embed the target detection box and the supplementary detection box into spatiotemporal nodes to construct a trajectory sequence, output the trajectory sequence and the conflict node identifier to the trusted decision module, and simultaneously feed back the conflict node identifier to the adaptive detection module as a re-identification trigger signal.
[0075] The spatiotemporal graph tracking module receives target detection bounding box data and supplementary detection bounding box data transmitted from the adaptive detection module. Both types of bounding boxes contain target location coordinates and multi-dimensional feature descriptors. The module constructs a spatiotemporal graph structured network model, embedding the bounding boxes as nodes into the temporal and spatial dimensions of a continuous frame sequence. Node attributes include position vectors and feature vectors. The model processes the node attribute data through a graph neural network, establishing spatiotemporal connections between nodes in adjacent frames to generate a continuous trajectory sequence.
[0076] The module analyzes the node association status in the trajectory sequence in real time. When the feature similarity of the detection box is lower than a set threshold or the position prediction deviation exceeds the permissible range, a conflict node is marked. The marker includes the spatial location index and time frame number of the conflict node, used to identify trajectory breaks or target confusion areas. The conflict node marker data is synchronously output to the trusted decision module for behavior analysis and fed back to the adaptive detection module as a re-identification trigger signal.
[0077] The conflict node identifiers fed back to the adaptive detection module drive the activation of the re-identification sub-network, triggering refined detection of ambiguous target regions. The trajectory sequence output by the spatiotemporal graph model contains a complete record of the target displacement path and attribute evolution. This sequence data is transmitted to the trusted decision module to support probabilistic behavioral analysis. Through the bidirectional transmission mechanism of the conflict node identifiers, the system forms a closed-loop control link for trajectory correction.
[0078] The trusted decision module is configured to receive the trajectory sequence and the channel attention weights generated by the feature alignment module, fuse the historical state of the trajectory sequence with the channel attention weights, generate a probabilistic behavior decision, and output the probabilistic behavior decision to the display terminal.
[0079] The trusted decision-making module receives trajectory sequence data transmitted by the spatiotemporal graph tracking module. The sequence contains the target's position vector and feature descriptors in consecutive time frames. Simultaneously, the module acquires the channel attention weight matrix generated by the feature alignment module, which reflects the importance distribution of different feature channels in behavior discrimination. It retrieves historical state records of the target from the internal trajectory history database, including displacement pattern statistics and prior probability distributions of behavior categories.
[0080] The module calculates feature confidence coefficients based on channel attention weights and fuses real-time motion state parameters from the trajectory sequence using weighted methods. These motion state parameters include the rate of change of velocity vector direction and the amplitude of acceleration fluctuations. The feature confidence coefficients apply weighted calibration to the real-time parameters. Historical state prior probabilities are corrected based on trajectory continuity assessment values, which quantify the smoothness characteristics of the target displacement path.
[0081] When the scene complexity index issued by the adaptive detection module exceeds the dynamic threshold, the module initiates a historical prior weight compression mechanism. The compression ratio is positively correlated with the scene complexity index, reducing the influence of historical data on the current decision. The calibrated real-time motion parameters and corrected historical probabilities are fused to generate a posterior probability distribution of behavior within a probabilistic graphical model framework.
[0082] The posterior probability distribution data is output to the display terminal, driving the generation of a visual behavior analysis interface. The interface displays the probability intensity of different behavior categories through heatmap overlay, and combines spatial coordinate mapping of trajectory sequences to visualize behavioral situations in dynamic environments. The probability distribution data is simultaneously fed back to the data acquisition module, providing a decision-making basis for sample synthesis in the adversarial generation module, forming a closed-loop optimization chain. In the tooling inspection scenario, the module generates a progressive wear probability heatmap based on the trajectory continuity coefficient of the cuff area and the fiber damage characteristic distribution, supporting tooling life assessment decisions.
[0083] The data acquisition module captures the raw video stream and separates it into visible light and near-infrared band data using optical components such as a beam splitter. It calculates a structural similarity index between the two bands to assess alignment quality. When this index falls below a preset threshold, the module sends a cross-spectral compensation request command, including spectral channel identifiers, to the adversarial generation module and receives a cross-spectral compensation sample synthesized by the adversarial generation module. This sample fuses texture features from both visible and infrared spectra. The module then performs frame-aligned fusion of the compensation sample and the original band data to generate a spatiotemporally aligned bimodal image sequence, which is output to the adversarial generation module to provide input data for subsequent feature processing.
[0084] The adversarial generation module receives a bimodal image sequence output from the data acquisition module and specific spatial coordinates fed back from the adaptive detection module. The module multiplies these specific spatial coordinates by the channel attention weights generated by the feature alignment module to generate a sample rendering intensity factor. A discriminator calculates the Euclidean distance between the generated sample and the source image in the feature space, constraining the gradient descent direction of the generative adversarial network and synthesizing low-frequency samples in a targeted manner. The synthesized augmented sample library is output to the feature alignment module to update the feature distribution and enhance sample diversity to adapt to dynamic scene changes.
[0085] The feature alignment module receives the augmented sample library output by the adversarial generation module and extracts the deep features of the samples. Depending on the real-time computational resource load, the module selects either a quantum compression algorithm or a manifold projection algorithm for feature optimization. It receives the loss gradient from the discriminator of the adversarial generation module and applies this gradient to initialize the Riemannian manifold space mapping parameters. Responding to the low-confidence region markers output by the adaptive detection module, the module dynamically adjusts the channel attention weights during the distribution adaptation phase based on the Wasserstein distance to calculate the feature distribution difference between the source and target domains, increasing the weight values of low-confidence regions. It optimizes the channel attention weight parameters through gradient backpropagation, generating fused features that are output to the adaptive detection module, supporting the optimization and updating of the detection path.
[0086] The adaptive detection module receives fused features and channel attention weights from the feature alignment module, as well as conflict node identifiers from the spatiotemporal graph tracking module. The module calculates a scene complexity index based on the channel attention weights and conflict node identifiers. If this index exceeds a preset threshold for multiple consecutive frames, a retargeting trigger condition is activated. The fused features are processed using lightweight convolution on the main path to generate target detection boxes, which are then output to the spatiotemporal graph tracking module. Responding to the conflict node identifiers as a re-identification trigger signal, the re-identification sub-network on the auxiliary path is activated, generating supplementary detection boxes that are output to the spatiotemporal graph tracking module, thus implementing an adaptive detection strategy based on scene complexity.
[0087] The spatiotemporal graph tracking module receives the target detection boxes and supplementary detection boxes output by the adaptive detection module, and embeds the detection boxes into spatiotemporal nodes to construct a trajectory sequence. The module receives channel attention weights from the feature alignment module as a weighting benchmark for appearance similarity calculation, and applies weighted augmentation to spatial flow appearance similarity calculation to generate weighted similarity data. Feature vectors of occluded nodes are extracted from the detection boxes and subjected to cross-modal matching with the augmented sample library of the adversarial generation module, with the matching process executed in parallel in both the HSV and Lab color spaces. If the matching similarity exceeds a preset threshold, the conflicting node attributes are corrected, and an updated trajectory sequence is generated and output to the trusted decision module; otherwise, a feature weight recalibration request is sent to the feature alignment module to maintain trajectory continuity.
[0088] The trusted decision-making module receives the updated trajectory sequence output by the spatiotemporal graph tracking module and the channel attention weights generated by the feature alignment module. The module acquires real-time target motion state data and obtains attenuation coefficient channel data from the feature alignment module; this coefficient is negatively correlated with the structural similarity index value output by the data acquisition module. It reads historical state prior probabilities from the internal trajectory historical state database and calculates the trajectory continuity coefficient based on the attenuation coefficient channel data to quantify trajectory stability. The historical state prior probabilities are weighted using the trajectory continuity coefficient, and the module monitors real-time changes in the dynamic threshold set by the adaptive detection module. If the scenario complexity index exceeds the dynamic threshold, the influence weight of historical prior probabilities is compressed. The weighted probabilities are fused with real-time state data to generate a dynamically calibrated posterior probability distribution, which is output to the display terminal to drive the visual behavior analysis interface, forming a closed-loop decision optimization mechanism.
[0089] Specifically, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the adversarial generation module is configured as follows:
[0090] Receive specific spatial coordinates from the adaptive detection module and channel attention weights from the feature alignment module;
[0091] Multiply the specific spatial coordinates by the channel attention weights to generate a sample rendering intensity factor;
[0092] The discriminator calculates the Euclidean distance between the generated sample and the source image in the feature space. The Euclidean distance is used to constrain the gradient descent direction of the generative adversarial network. The augmented sample is then output to the feature alignment module to update the feature distribution of the augmented sample library.
[0093] The adversarial generation module obtains specific spatial coordinates from the adaptive detection module, which identify image regions with low detection confidence. Simultaneously, the module receives channel attention weights generated by the feature alignment module, reflecting the importance distribution of different feature channels in the object detection task. A matrix multiplication operation is performed between the specific spatial coordinates and the channel attention weights to generate a sample rendering intensity factor, which quantifies the priority of sample synthesis for each spatial location.
[0094] The module inputs the sample rendering intensity factor into the generator of the conditional generative adversarial network (GAN), guiding the network resources to focus on feature synthesis in low-confidence regions. The discriminator network calculates the Euclidean distance between the generated sample feature vector and the source image feature vector, serving as a feature similarity metric. This Euclidean distance, as a core component of the loss function, constrains the direction of the GAN's backpropagation gradient update, enabling the generated samples to gradually approximate the distribution of real samples in the feature space.
[0095] During training iterations, the generator adjusts network parameters according to the gradient descent direction, synthesizing augmented samples that cover low-frequency features. The newly synthesized samples are input into the feature alignment module to update the feature distribution of the augmented sample library. By continuously injecting synthetic samples optimized for low-confidence regions, the system gradually reduces the difference in feature distribution between the training data and the actual scene, forming an adaptive calibration mechanism for the feature space.
[0096] Specifically, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the feature alignment module is configured as follows:
[0097] The loss gradient is received from the discriminator of the adversarial generation module, and the Riemannian manifold space mapping parameters are initialized using the loss gradient;
[0098] In response to the low-confidence region markers output by the adaptive detection module, the distribution adaptation phase is executed by calculating the difference in feature distribution between the source and target domains based on the Wasserstein distance. The channel attention weights are then dynamically adjusted according to the magnitude of the difference in feature distribution between the source and target domains, thereby increasing the channel attention weights of the regions corresponding to the low-confidence region markers by 20% to 50%.
[0099] The channel attention weight parameters are optimized through gradient backpropagation to generate optimized channel attention weights. The optimized channel attention weights are then output to the adaptive detection module to update the convolution kernel parameters of the adaptive detection module.
[0100] The feature alignment module receives the loss gradient output by the discriminator of the adversarial generation module. This gradient contains information about the feature space differences between generated and real samples. The module applies this loss gradient to initialize the mapping parameters of the Riemannian manifold space, establishing the geometric representation basis of the feature distribution. This parameter initialization method aligns the manifold space mapping direction with the generative adversarial training objective, providing geometric constraints for feature distribution adaptation.
[0101] In response to the low-confidence region markers output by the adaptive detection module, the module performs feature alignment during the distribution adaptation phase. The Wasserstein distance is used to measure the magnitude of the difference between the feature distribution of the source domain and the actual scene feature distribution of the target domain; this distance calculation employs optimal transport theory to quantify the distribution difference. Based on the magnitude of the difference, the channel attention weight parameters are dynamically adjusted to strengthen the feature channels corresponding to the low-confidence region markers. This weight adjustment mechanism allocates system resources to the weaker parts of the model, enhancing the representation ability of difficult example features.
[0102] The module iteratively optimizes the channel attention weight parameters using the gradient backpropagation algorithm, calculating the gradient of the impact of weight adjustments on detection performance. The optimized channel attention weights are output to the adaptive detection module to update its convolution kernel parameters. This process forms a closed-loop feedback between detection feature requirements and feature generation supply, ensuring that the feature distribution continuously approximates the needs of the actual scenario. The transfer of weight parameters enables cross-module parameter linkage updates, improving the system's overall adaptability to dynamic environments.
[0103] Specifically, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the adaptive detection module is configured as follows:
[0104] Receive channel attention weights from the feature alignment module and conflict node identifiers from the spatiotemporal graph tracking module;
[0105] Calculate the scene complexity index based on the channel attention weights and conflict node identifiers;
[0106] If the scene complexity index exceeds a preset threshold for three consecutive frames, the redirection trigger condition is activated.
[0107] In response to the redirection triggering condition, a focus area redirection instruction is sent to the re-identification sub-network.
[0108] The adaptive detection module receives channel attention weights transmitted by the feature alignment module. These weights represent the contribution distribution of different feature channels in the detection task. Simultaneously, the module acquires conflict node identifiers fed back by the spatiotemporal graph tracking module. These identifiers reflect ambiguous regions in trajectory association during target tracking. Based on the entropy changes of the channel attention weights and the density distribution of conflict node identifiers, the module calculates a dynamic scene complexity index. This index comprehensively evaluates the intensity of environmental interference and the complexity of target interaction.
[0109] The module implements a scene complexity index threshold monitoring mechanism. When the index exceeds a preset threshold for multiple consecutive frames, the scene is determined to have entered a high-complexity state. At this point, the system activates the redirection trigger condition and generates a focus area redirection command. This command contains the set of spatial coordinates corresponding to the conflict node identifiers and the feature channel identifiers.
[0110] In response to the redirection trigger condition, the module sends a focus region redirection command to the re-identification sub-network. This command drives the re-identification sub-network to focus computational resources on the region corresponding to the conflict node identifier, performing high-resolution feature extraction. The redirection mechanism enables dynamic scheduling of detection resources, allowing the system to maintain continuous target tracking capabilities even in highly complex scenarios.
[0111] Specifically, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the spatiotemporal graph tracking module is configured as follows:
[0112] The system receives channel attention weights from the feature alignment module as a weighting benchmark for appearance similarity calculation, and receives target detection boxes and supplementary detection boxes from the adaptive detection module, parsing target location information and feature descriptors.
[0113] The channel attention weight is applied to enhance the spatial flow appearance similarity calculation, generate weighted similarity data, extract the feature vector of the occluded node from the received detection box, and perform cross-modal matching between the extracted feature vector and the augmented sample library of the adversarial generation module. The matching process is performed in parallel in the HSV color space and the Lab color space. If the matching similarity exceeds 85%, the conflict node attributes are corrected and an updated trajectory sequence is generated.
[0114] The updated trajectory sequence is output to the trusted decision module for behavioral decision analysis. If the matching fails, a feature weight recalibration request is sent to the feature alignment module.
[0115] The spatiotemporal graph tracking module receives channel attention weights transmitted from the feature alignment module, which serve as a weighted benchmark for appearance feature similarity measurement. The module synchronously acquires the target detection boxes and supplementary detection boxes output by the adaptive detection module, and parses the position coordinates and multi-dimensional feature descriptors of the targets within the boxes. Channel attention weights are applied to weight the spatial flow appearance similarity calculation, increasing the contribution of key feature channels in the similarity measurement, and generating weighted similarity data for trajectory association analysis.
[0116] The module identifies the feature vector of the occluded target from the detection bounding box and extracts this vector as a query sample for cross-modal matching. The query sample is then matched against an augmented sample library maintained by the adversarial generation module. Hue / saturation matching is performed in the HSV color space, while chroma / luminance matching is performed in parallel in the Lab color space. When the similarity between the two color spaces exceeds a preset threshold, the cross-modal matching is considered successful. Based on this, the attribute labels of conflicting nodes are corrected, and an ambiguous updated trajectory sequence is generated.
[0117] The updated trajectory sequence is output to the trusted decision module, providing spatiotemporal continuity data for behavioral decision analysis. If cross-modal matching fails, it indicates an adaptation bias in the current feature distribution. The module then sends a feature weight recalibration request to the feature alignment module. This request includes the spatial coordinates and feature channel identifiers of the failed matching nodes, driving the upstream module to re-optimize the feature distribution and forming a closed-loop calibration mechanism.
[0118] Specifically, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the trusted decision module is configured as follows:
[0119] The updated trajectory sequence is received from the spatiotemporal map tracking module to obtain real-time target motion state data. Attenuation coefficient channel data is obtained from the feature alignment module. The attenuation coefficient is negatively correlated with the structural similarity index value output by the data acquisition module. The prior probability of historical state is read from the internal trajectory historical state database.
[0120] The trajectory continuity coefficient is calculated based on the attenuation coefficient channel data to quantify trajectory stability. The trajectory continuity coefficient is used to weight the prior probabilities of historical states. The dynamic threshold changes set by the adaptive detection module are monitored in real time. In response to the scene complexity index exceeding the dynamic threshold, the weight of the historical prior influence is compressed by 40% to 60%.
[0121] By integrating weighted probability and real-time status data, a dynamically calibrated posterior probability distribution is generated and output to the display terminal to drive the visual behavior analysis interface.
[0122] The trusted decision-making module receives the updated trajectory sequence output by the spatiotemporal map tracking module, which contains the target's spatiotemporal position information in consecutive frames. The module synchronously acquires real-time target motion state data, including velocity vectors and acceleration variation characteristics. It extracts attenuation coefficient channel data from the feature alignment module; this coefficient is inversely proportional to the structural similarity index calculated by the data acquisition module, reflecting the degree of reliability attenuation in multimodal data alignment quality. The module retrieves historical state prior probabilities from its internal trajectory historical state database as a benchmark reference for behavior prediction.
[0123] Based on the attenuation coefficient channel data, the module calculates the trajectory continuity coefficient and quantifies trajectory stability through the smoothness analysis of the target displacement path. The trajectory continuity coefficient is used to weight the prior probabilities of historical states, giving higher historical weights to trajectories with high stability. The module monitors dynamic threshold parameters issued by the adaptive detection module in real time. When the scene complexity index exceeds this threshold, a historical prior influence weight compression mechanism is activated, significantly reducing the proportion of historical data's influence in the current decision-making process.
[0124] By fusing weighted historical probability data with real-time motion state data, a dynamically calibrated posterior probability distribution is generated within a Bayesian inference framework. This distribution considers both historical trajectory stability and real-time scene complexity, and is output to a display terminal to drive a visual behavior analysis interface. The probability distribution heatmap is overlaid with the trajectory, enabling a visual representation of behavioral decisions in dynamic environments.
[0125] Specifically, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the data acquisition module is configured as follows:
[0126] The original video stream is separated into visible light band data and near-infrared band data by a beam splitter prism. The structural similarity index of the visible light band data and near-infrared band data is calculated in real time to evaluate the alignment quality of the visible light band data and near-infrared band data.
[0127] In response to a structural similarity index below 0.8, a cross-spectral compensation request instruction is sent to the adversarial generation module. The cross-spectral compensation request instruction includes a real-time frame structural similarity index value and a spectral channel identifier. The cross-spectral compensation sample synthesized by the adversarial generation module is received. The cross-spectral compensation sample fuses texture features from visible light and infrared spectra.
[0128] The cross-spectral compensation samples are frame-aligned and fused with the original band data to generate a spatiotemporally aligned bimodal sequence. The frame-aligned bimodal sequence is then output to the feature alignment module for feature distribution optimization of the augmented sample library.
[0129] The data acquisition module separates the input video stream into visible and near-infrared spectral data using an optical beam splitter. The module calculates the structural similarity index of corresponding frames in both spectral bands in real time. This index comprehensively evaluates the consistency of brightness, contrast, and structural features, quantifying the alignment quality of the multispectral data. When the index value falls below a preset quality threshold, the module generates a cross-spectral compensation request command. This command includes an abnormal spectral channel identifier and deviation parameters, triggering the adversarial generation module to perform compensation operations.
[0130] After responding to the request, the adversarial generation module uses a generative adversarial network to synthesize a cross-spectral compensation sample. This sample fuses the color features of the visible light channel with the texture details of the near-infrared channel, achieving multispectral information complementarity through feature space mapping. After the synthesized sample is input into the data acquisition module, the module performs a multi-source data fusion operation: it uses a feature point matching algorithm to establish the spatial correspondence between the compensation sample and the original spectral data, achieves pixel-level frame alignment through affine transformation, and generates a spatiotemporally synchronized bimodal sequence.
[0131] The optimized bimodal sequence is output to the feature alignment module as an update source for the augmented sample library. This sequence contains enhanced feature representations of the spectral compensation region, providing high-quality input for feature distribution optimization. Through a closed-loop spectral quality monitoring and compensation mechanism, the system effectively overcomes the spectral distortion problem caused by changes in ambient illumination and maintains the stability of the multimodal feature space.
[0132] Specifically, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the re-identification sub-network is configured as follows:
[0133] The cross-modal matching compensation features are received from the spatiotemporal graph tracking module. The cross-modal matching compensation features include occlusion region repair data and spatial coordinate identifiers in the Lab color space. The entangled state vector output by the quantum compression algorithm is called from the feature alignment module to obtain the entanglement association information in the feature space.
[0134] Real-time monitoring of scene complexity index changes; when the moving average of the index increases by 10% or more over 5 consecutive frames, switch to manifold projection algorithm to perform re-identification operation; under the manifold projection algorithm path, use the shortest path algorithm of Riemannian manifold space to realize geodesic distance measurement; perform feature matching based on geodesic distance to generate cross-modal feature similarity evaluation results.
[0135] Based on the feature matching results, an optimized supplementary detection box is generated and output to the corresponding node of the spatiotemporal graph tracking module to update the trajectory of the conflict node.
[0136] The re-identification sub-network obtains cross-modal matching compensation features from the spatiotemporal graph tracking module. These features include occlusion region repair data in the Lab color space and corresponding spatial coordinate labels. The module synchronously calls the output of the quantum compression algorithm stored in the feature alignment module to extract entangled state vector data in the feature space. This vector encodes the quantum entanglement correlation characteristics between multi-channel features, providing high-order correlation information to support cross-modal matching.
[0137] The module monitors the scene complexity index changes in real time from the adaptive detection module. When the moving average of this index rises to a preset threshold over multiple consecutive frames, the system determines that the scene complexity has entered a phase of rapid change. At this point, the system automatically switches to the manifold projection algorithm path to perform re-identification, dynamically adapting to the rate of environmental change through algorithmic switching.
[0138] Within the manifold projection algorithm, the module employs a geodesic distance metric based on Riemannian manifold space. Leveraging the curvature characteristics of the manifold geometry, the module calculates the shortest path length of feature points in the Riemannian manifold space as a similarity metric. By comparing the manifold space projection positions of the current target feature and the features in the sample database with the geodesic distance, a cross-modal feature similarity metric evaluation result is generated.
[0139] Based on the feature matching similarity evaluation results, the module generates optimized supplementary detection boxes. These detection boxes correct the position coordinates and feature descriptors of the original conflict nodes and are output to the corresponding spatiotemporal nodes of the spatiotemporal graph tracking module. The updated detection data is used to reconstruct the target trajectory sequence, eliminating trajectory interruptions caused by occlusion or deformation and maintaining the spatiotemporal continuity of target tracking.
[0140] Specifically, in the deep learning-based real-time target detection and tracking system for video streams described in this invention, the trusted decision module is configured as follows:
[0141] The structural similarity index value is obtained in real time from the data acquisition module. The feature confidence decay coefficient is calculated based on the structural similarity index value. The output range of the feature confidence decay coefficient is 0.2 to 1.0. The historical prior influence weight compression status is read from the internal weight compression record database.
[0142] Continuously monitor the compression magnitude of historical prior influence weights and the structural similarity index value. When the compression magnitude exceeds 40% and the structural similarity index value of two consecutive frames is lower than 0.8, calculate the average index value of the two consecutive frames. Based on the difference between 0.8 and the average index value of the two consecutive frames, multiply by 0.35 to calculate the probability cumulative threshold reduction ratio.
[0143] The probability accumulation threshold is lowered according to the probability accumulation threshold reduction ratio, an adaptive alarm trigger command including the target identifier and confidence deviation value is generated, the alarm trigger command is output to the display terminal, and the high priority alarm interface is driven.
[0144] The trusted decision-making module receives the structural similarity index value transmitted by the data acquisition module in real time. This index value characterizes the spectral alignment quality of the multimodal data. Based on this index value, the module calculates the feature reliability decay coefficient, which is positively correlated with the spectral alignment quality and is used to quantify the degree of reliability decay of the current feature data. The module synchronously accesses the internal weight compression record database to extract historical prior influence parameters on the weight compression state. These parameters record the adjustment strength of the adaptive detection module's dependence on historical data.
[0145] The module establishes a multi-parameter joint monitoring mechanism to continuously track the coupling trend of historical prior influence weight compression magnitude and structural similarity index value. When the weight compression magnitude exceeds a preset intensity threshold and the structural similarity index value is lower than the quality baseline for multiple consecutive frames, an adaptive alarm analysis process is triggered. At this time, the module calculates the average index value for multiple consecutive frames, and based on the relative deviation of this average value from the quality baseline, combined with the sensitivity coefficient set by the system, calculates the dynamic adjustment ratio of the probability accumulation threshold.
[0146] The module lowers the cumulative probability threshold parameter according to the calculated adjustment ratio, generating an adaptive alarm trigger command containing a unique target identifier and confidence deviation value. This command drives the display terminal to activate the high-priority alarm interface, displaying a visually enhanced alert of key abnormal targets through a flashing target location box overlaid with a deviation heatmap. The threshold adjustment mechanism enables the system to maintain decision sensitivity when data quality fluctuates, avoiding the underreporting of important abnormal events.
[0147] Specifically, the deep learning-based real-time target detection and tracking system for video streams described in this invention further includes:
[0148] The data acquisition module captures the internal fiber texture of the dark tooling in the near-infrared band and outputs a dual-modal sequence including fiber texture features to the adversarial generation module. The adversarial generation module synthesizes reflective samples of the tooling and outputs them to the feature alignment module.
[0149] The feature alignment module receives reflective samples and maps them to the Riemannian manifold space. Based on the Wasserstein distance distribution of fiber texture features in the augmented sample library, it filters out damaged feature clusters with a distance less than 0.3. The trustworthy decision module calls the cuff region trajectory continuity coefficient output by the spatiotemporal graph tracking module.
[0150] When the continuity coefficient of the cuff area trajectory is less than 0.5, a weighted constraint is applied to the wear probability. The weighted constraint amplitude increases with the wear probability value, generating a progressive wear probability report, including cuff area coordinates and historical wear trend curves. The progressive wear probability report is output to the display terminal to drive the tooling life assessment interface.
[0151] The data acquisition module is equipped with a near-infrared imaging unit that penetrates the surface of dark-colored tooling to capture the internal fiber structure and texture features. This unit enhances the scattering properties of the fabric fibers using a specific wavelength light source, generating a near-infrared image sequence containing microscopic texture information. The module fuses the visible light appearance image and the near-infrared texture image into a dual-modal sequence, which is then output to the adversarial generation module for sample optimization processing.
[0152] After receiving the bimodal sequence, the adversarial generation module identifies the fiber distribution pattern in key areas of the tooling. The module then synthesizes reflective property samples of the tooling using a conditional adversarial network, focusing on enhancing the optical reflection characteristics of high-wear areas such as cuffs and elbows. The synthesized samples fuse visible light color attributes with near-infrared structural features and are output to the feature alignment module to establish a wear analysis sample library.
[0153] The feature alignment module maps reflective samples to a Riemannian manifold space, preserving the topological structure of the fiber texture using manifold geometry. The module calculates the Wasserstein distance distribution of fiber feature vectors in the augmented sample library and filters potential damage feature clusters based on the distance distribution density. Using a geodesic distance metric in the manifold space, it identifies sets of fiber features with similar damage patterns and generates a spatial coordinate mapping table of potential wear areas.
[0154] The trusted decision-making module calls the spatiotemporal graph tracking module to calculate the continuity coefficient of the cuff region trajectory. This coefficient quantifies the displacement stability of the target cuff area in consecutive frames. When the coefficient is lower than a preset benchmark value, the module activates the wear probability enhancement mechanism: based on the spatial distribution density of the damage feature clusters, it dynamically increases the weight ratio of the corresponding region in the probability model. The weight enhancement magnitude adopts a non-linear growth strategy, increasing with the initial wear probability value.
[0155] The module integrates fiber breakage characteristic distribution, trajectory stability coefficient, and weighted reinforcement parameters to generate a progressive wear probability report. The report includes spatial coordinate encoding of the wear area, real-time probability values, and historical trend curves. The curve data originates from periodic detection results in a historical database, and time-series analysis is used to demonstrate the wear evolution trend. The complete report is output to the display terminal, driving the 3D model rendering and risk heatmap overlay display on the tooling life assessment interface.
[0156] This invention addresses the feature space mapping inaccuracy caused by the long-tail distribution differences in training data by constructing a closed-loop optimization mechanism. The data acquisition module captures bimodal image sequences, and the generative adversarial network (GAN) synthesizes low-frequency samples based on specific spatial coordinates fed back by the adaptive detection module, compensating for the missing long-tail distribution in the training set. By multiplying the spatial coordinates with the channel attention weights provided by the feature alignment module to generate a sample rendering intensity factor, the GAN is constrained to gradient descent towards low-frequency feature regions, generating an augmented sample library to update the feature distribution and narrowing the difference between the training data and the actual scene distribution.
[0157] After receiving the augmented sample library, the feature alignment module selects either quantum compression or manifold projection algorithm to optimize the feature distribution based on real-time resource load. It initializes the Riemannian manifold mapping parameters using the loss gradient of the discriminator in the adversarial generation module, and dynamically increases the attention weights of channels corresponding to low-confidence regions based on the Wasserstein distance quantizing the difference in feature distribution between the source and target domains. The weight parameters are optimized through gradient backpropagation and output to the adaptive detection module to update the convolutional kernel, achieving alignment and calibration between the feature distribution and detection requirements.
[0158] When constructing trajectory sequences, the spatiotemporal graph tracking module uses channel attention weights to calculate appearance similarity and performs cross-modal matching with the augmented sample library to correct conflicting nodes. The reliable decision-making module integrates trajectory continuity coefficients and scene complexity indices, dynamically compressing historical prior weights to generate probabilistic decisions. The decision results are fed back to the data acquisition and sample generation stages through a display terminal, forming a closed-loop optimization of "coordinate feedback driving sample synthesis, channel weight adjustment and distribution calibration, and reliable trajectory continuity constraint decisions," continuously compressing feature mapping bias.
[0159] The key parameters and application logic of each algorithm module in the technical solution of this invention are as follows. The sample rendering intensity factor parameter of the adversarial generation module is generated by the dot product of a specific spatial coordinate matrix and a channel attention weight matrix, and is input to the low-confidence region coordinate data fed back by the adaptive detection module and the channel weight data transmitted by the feature alignment module. The module processes the intensity factor through a conditional generative adversarial network, uses the Euclidean distance parameter between the generated sample and the source image feature space calculated by the discriminator to constrain the gradient descent direction, and outputs an augmented sample library to the feature alignment module to achieve low-frequency feature sample synthesis.
[0160] The Wasserstein distance parameter of the feature alignment module measures the difference in feature distribution between the source and target domains. It is input to the loss gradient data of the discriminator in the adversarial generation module and the low-confidence region markers output by the adaptive detection module. The module initializes the Riemannian manifold mapping parameters using the loss gradient, dynamically adjusts the channel attention weight enhancement ratio based on the magnitude of the distribution difference, optimizes the weight parameters through the backpropagation algorithm, and outputs the optimized channel attention weights to the adaptive detection module to drive the update of the detection model parameters.
[0161] The adaptive detection module integrates the scene complexity index parameter with the channel weight entropy change rate and the spatial density of conflict nodes. It inputs the channel weight data transmitted by the feature alignment module and the conflict node identifiers fed back by the spatiotemporal map tracking module. The module monitors consecutive frames exceeding the index limit to activate redirection trigger conditions and outputs focus region redirection instruction data to the re-identification sub-network, including the conflict node coordinate set and feature channel identifiers, enabling dynamic switching between dual detection paths.
[0162] The cross-modal matching similarity threshold parameter of the spatiotemporal graph tracking module serves as the correction benchmark for trajectory conflict nodes. It inputs the channel weight data, target detection box feature descriptors, and occluded node feature vectors into the feature alignment module. The module applies channel weights to enhance dual-color space appearance similarity calculation, using parallel processing of HSV hue / saturation matching and Lab chroma / brightness matching. The updated trajectory sequence is then output to the trusted decision module or a feature weight recalibration request.
[0163] The reliable decision-making module quantifies the smoothness of the displacement path using the trajectory continuity coefficient parameter. It inputs trajectory sequence data from the spatiotemporal graph tracking module, attenuation coefficient channel data from the feature alignment module, and prior probabilities from the historical state database. Based on the attenuation coefficient, the module calculates trajectory stability weights, monitors the scene complexity index to dynamically compress the proportion of historical influences, fuses real-time motion state data to generate a posterior probability distribution, and outputs probabilistic decision commands to drive the display terminal.
[0164] In tooling inspection scenarios, the Wasserstein distance parameter for fiber texture is used to filter manifold spatial damage feature clusters, which are then input into the fiber texture feature vector from near-infrared imaging. The cuff region trajectory continuity coefficient parameter is input into the motion trajectory data of the spatiotemporal mapping tracking module. When the coefficient falls below a set benchmark, a wear probability weighting enhancement mechanism is triggered, outputting a progressive wear report that includes historical trend curves and spatial coordinate mapping. These parameters work together through closed-loop feedback to form a system-level synergy, supporting feature mapping optimization in dynamic environments.
[0165] This invention addresses the target detection bias caused by discrepancies between training data and actual scene distribution through a multi-module collaborative mechanism. In a smart venue monitoring scenario, the data acquisition module utilizes a near-infrared imaging unit to penetrate the surface of dark workwear, capturing the internal fiber structure texture features and generating a dual-modal sequence of visible and near-infrared light. After receiving this sequence, the adversarial generation module, based on the cuff wear area coordinates fed back by the adaptive detection module and combined with the channel attention weights provided by the feature alignment module, synthesizes reflective samples containing fiber damage features to compensate for missing workwear wear sample types in the training set.
[0166] The feature alignment module maps the synthesized samples to a Riemannian manifold space, and selects either quantum compression or manifold projection algorithms to optimize the feature distribution based on real-time computing resources. The manifold mapping parameters are initialized using the loss gradient provided by the discriminator in the adversarial generation module, and the Wasserstein distance distribution of fiber texture features in the augmented sample library is calculated to dynamically increase the channel attention weights in low-confidence wear regions. The optimized weight parameters update the convolutional kernels of the adaptive detection module through gradient backpropagation, enhancing the model's ability to extract fiber damage features.
[0167] The adaptive detection module calculates the scene complexity index based on the conflict node identifiers fed back by the spatiotemporal graph tracking module and the channel weights. When the index continuously exceeds the threshold, the re-identification sub-network is activated to perform high-resolution feature extraction in the cuff region. The spatiotemporal graph tracking module applies channel weights to calculate appearance similarity, performs cross-modal matching of the occluded workwear node features with the augmented sample library, and generates an ambiguous trajectory sequence. The reliable decision module integrates the trajectory continuity coefficient and the scene complexity index. When the coefficient in the cuff region is below the threshold, it applies nonlinear weight enhancement to the wear probability, generating a progressive probability report containing historical wear trend curves. The system continuously optimizes the feature space mapping accuracy through a triple closed loop of coordinate feedback-driven directional sample synthesis, channel weight adjustment feature distribution calibration, and trajectory continuity constraint decision reliability, overcoming the missed detection problem caused by the difficulty in recognizing dark workwear textures.
[0168] This invention achieves dynamic calibration of the difference between training data and the actual scene distribution through a modular architecture design. The data acquisition module uses optical beam splitting technology to separate visible light and near-infrared spectra, generating a dual-modal image sequence. When the structural similarity index is lower than a preset threshold, a cross-spectral compensation mechanism is triggered, and the adversarial generation module fuses dual-spectral texture features to synthesize compensated samples, solving the spectral distortion problem caused by changes in ambient lighting.
[0169] The feature alignment module employs a distribution adaptation strategy, initializing the Riemannian manifold mapping parameters based on the loss gradient of the adversarial generation module. The difference in feature distribution between the source and target domains is quantified using the Wasserstein distance, and channel attention weights are applied to strengthen low-confidence regions identified by the adaptive detection module. Dynamic adjustment of these weights updates the convolutional kernel parameters of the detection module through gradient backpropagation, forming a closed-loop optimization of feature demand and generation supply.
[0170] The adaptive detection module innovatively designs a dual-path switching mechanism. It calculates a scene complexity index based on channel attention weights and conflict node identifiers, and activates a redirection instruction when the index continuously exceeds a threshold. The main path uses lightweight convolution to handle normal scenes, while the auxiliary path focuses on conflict regions through a re-identification sub-network, maintaining detection accuracy under resource-constrained conditions.
[0171] The spatiotemporal graph tracking module establishes a cross-modal feature matching mechanism. It applies channel weighting to calculate appearance similarity and performs parallel matching of occluded target features with an augmented sample library in a dual-color space. Successfully matched feature vectors are used to correct the trajectories of conflicting nodes, while failed cases trigger feature weight recalibration requests, forming a mutual feedback mechanism between trajectory continuity and feature calibration.
[0172] The credible decision-making module innovatively integrates multi-source decision factors. It calculates the feature credibility decay coefficient based on the structural similarity index and dynamically weights historical prior probabilities by combining them with the trajectory continuity coefficient. When it detects that weight compression exceeds limits and data quality continues to deteriorate, the system automatically lowers the probability accumulation threshold and generates an adaptive alarm command with a target identifier, balancing the conflict between false alarm rate and false negative rate.
[0173] In tooling inspection scenarios, the system captures fiber texture by penetrating dark fabrics using near-infrared imaging. The feature alignment module filters for damage feature clusters with proximity to the Wasserstein distance in the manifold space, while the reliable decision module utilizes the trajectory continuity coefficient of the cuff region. When the coefficient falls below a threshold, a nonlinear weighting enhancement strategy is implemented, generating a progressive wear probability report containing historical trend curves, providing a quantitative basis for tooling life assessment. Each module continuously compresses feature mapping bias through a triple closed loop of coordinate feedback, weight adjustment, and trajectory constraints.
Claims
1. A real-time target detection and tracking system for video streams based on deep learning, characterized in that, include: The data acquisition module is configured to capture the raw video stream and output a bimodal image sequence to the adversarial generation module; The adversarial generation module is configured to receive the bimodal image sequence and specific spatial coordinates fed back by the adaptive detection module, synthesize low-frequency samples based on the specific spatial coordinates, and output the synthesized augmented sample library to the feature alignment module. The feature alignment module is configured to receive the augmented sample library, extract the depth features of the augmented sample library, select a quantum compression algorithm or a manifold projection algorithm for feature optimization processing according to the real-time computing resource load, generate channel attention weights and fuse features based on the weights to obtain fused features, and output the fused features to the adaptive detection module. The adaptive detection module is configured to process the fusion features using lightweight convolution through the main path based on the scene complexity index of the fusion features, generate target detection boxes and output the target detection boxes to the spatiotemporal graph tracking module, and activate the re-identification sub-network through the secondary path in response to the conflict node identifier, generate supplementary detection boxes and output the supplementary detection boxes to the spatiotemporal graph tracking module. The spatiotemporal graph tracking module is configured to receive the target detection box and the supplementary detection box, embed the target detection box and the supplementary detection box into spatiotemporal nodes to construct a trajectory sequence, output the trajectory sequence and the conflict node identifier to the trusted decision module, and simultaneously feed back the conflict node identifier to the adaptive detection module as a re-identification trigger signal. The trusted decision-making module is configured to receive the channel attention weights generated by the trajectory sequence and feature alignment module, fuse the historical state of the trajectory sequence with the channel attention weights, generate probabilistic behavior decisions, and output the probabilistic behavior decisions to the display terminal. The adversary generation module is configured as follows: Receive specific spatial coordinates from the adaptive detection module and channel attention weights from the feature alignment module; Multiply the specific spatial coordinates by the channel attention weights to generate a sample rendering intensity factor; The discriminator calculates the Euclidean distance between the generated sample and the source image in the feature space. The Euclidean distance is used to constrain the gradient descent direction of the generative adversarial network. The augmented sample is then output to the feature alignment module to update the feature distribution of the augmented sample library. The feature alignment module is configured as follows: The loss gradient is received from the discriminator of the adversarial generation module, and the Riemannian manifold space mapping parameters are initialized using the loss gradient; In response to the low-confidence region markers output by the adaptive detection module, the distribution adaptation phase is executed by calculating the difference in feature distribution between the source and target domains based on the Wasserstein distance. The channel attention weights are dynamically adjusted according to the magnitude of the difference in feature distribution between the source and target domains, thereby increasing the channel attention weights of the regions corresponding to the low-confidence region markers by 20% to 50%. The channel attention weight parameters are optimized through gradient backpropagation to generate optimized channel attention weights. The optimized channel attention weights are then output to the adaptive detection module to update the convolution kernel parameters of the adaptive detection module. The adaptive detection module is configured as follows: Receive channel attention weights from the feature alignment module and conflict node identifiers from the spatiotemporal graph tracking module; Calculate the scene complexity index based on the channel attention weights and conflict node identifiers; If the scene complexity index exceeds a preset threshold for three consecutive frames, the redirection trigger condition is activated. In response to the redirection triggering condition, a focus area redirection instruction is sent to the re-identification sub-network; The spatiotemporal graph tracking module is configured as follows: The system receives channel attention weights from the feature alignment module as a weighting benchmark for appearance similarity calculation, and receives target detection boxes and supplementary detection boxes from the adaptive detection module, parsing target location information and feature descriptors. The channel attention weight is applied to enhance the spatial flow appearance similarity calculation, generate weighted similarity data, extract the feature vector of the occluded node from the received detection box, and perform cross-modal matching between the extracted feature vector and the augmented sample library of the adversarial generation module. The matching process is performed in parallel in the HSV color space and the Lab color space. If the matching similarity exceeds 85%, the conflict node attributes are corrected and an updated trajectory sequence is generated. The updated trajectory sequence is output to the trusted decision module for behavioral decision analysis. If the matching fails, a feature weight recalibration request is sent to the feature alignment module.
2. The real-time target detection and tracking system for video streams based on deep learning according to claim 1, characterized in that, The trusted decision module is configured as follows: The updated trajectory sequence is received from the spatiotemporal map tracking module to obtain real-time target motion state data. Attenuation coefficient channel data is obtained from the feature alignment module. The attenuation coefficient is negatively correlated with the structural similarity index value output by the data acquisition module. The prior probability of historical state is read from the internal trajectory historical state database. The trajectory continuity coefficient is calculated based on the attenuation coefficient channel data to quantify trajectory stability. The trajectory continuity coefficient is used to weight the prior probabilities of historical states. The dynamic threshold changes set by the adaptive detection module are monitored in real time. In response to the scene complexity index exceeding the dynamic threshold, the weight of the historical prior influence is compressed by 40% to 60%. By integrating weighted probability and real-time status data, a dynamically calibrated posterior probability distribution is generated and output to the display terminal.
3. The real-time target detection and tracking system for video streams based on deep learning according to claim 2, characterized in that, The data acquisition module is configured as follows: The original video stream is separated into visible light band data and near-infrared band data by a beam splitter prism. The structural similarity index of the visible light band data and near-infrared band data is calculated in real time to evaluate the alignment quality of the visible light band data and near-infrared band data. In response to a structural similarity index below 0.8, a cross-spectral compensation request instruction is sent to the adversarial generation module. The cross-spectral compensation request instruction includes a real-time frame structural similarity index value and a spectral channel identifier. The cross-spectral compensation sample synthesized by the adversarial generation module is received. The cross-spectral compensation sample fuses texture features from visible light and infrared spectra. The cross-spectral compensation samples are frame-aligned and fused with the original band data to generate a spatiotemporally aligned bimodal sequence. The frame-aligned bimodal sequence is then output to the feature alignment module for feature distribution optimization of the augmented sample library.
4. The real-time target detection and tracking system for video streams based on deep learning according to claim 1, characterized in that, The re-identification sub-network is configured as follows: The cross-modal matching compensation features are received from the spatiotemporal graph tracking module. The cross-modal matching compensation features include occlusion region repair data and spatial coordinate identifiers in the Lab color space. The entangled state vector output by the quantum compression algorithm is called from the feature alignment module to obtain the entanglement association information in the feature space. Real-time monitoring of scene complexity index changes; when the moving average of the index increases by 10% or more over 5 consecutive frames, switch to manifold projection algorithm to perform re-identification operation; under the manifold projection algorithm path, use the shortest path algorithm of Riemannian manifold space to realize geodesic distance measurement; perform feature matching based on geodesic distance to generate cross-modal feature similarity evaluation results. Based on the feature matching results, an optimized supplementary detection box is generated and output to the corresponding node of the spatiotemporal graph tracking module to update the trajectory of the conflict node.
5. The real-time target detection and tracking system for video streams based on deep learning according to claim 4, characterized in that, The trusted decision module is configured as follows: The structural similarity index value is obtained in real time from the data acquisition module. The feature confidence decay coefficient is calculated based on the structural similarity index value. The output range of the feature confidence decay coefficient is 0.2 to 1.
0. The historical prior influence weight compression status is read from the internal weight compression record database.
6. The real-time target detection and tracking system for video streams based on deep learning according to claim 5, characterized in that, Also includes: Monitor the compression magnitude of historical prior influence weights and the structural similarity index value. When the compression magnitude exceeds 40% and the structural similarity index value of two consecutive frames is lower than 0.8, calculate the average index value of the two consecutive frames. Based on the difference between 0.8 and the average index value of the two consecutive frames, multiply by 0.35 to calculate the probability cumulative threshold reduction ratio. The probability accumulation threshold is lowered according to the probability accumulation threshold reduction ratio, an adaptive alarm trigger command including the target identifier and confidence deviation value is generated, the alarm trigger command is output to the display terminal, and the high priority alarm interface is driven.
Citation Information
Patent Citations
Target tracking method and device, electronic equipment and readable storage medium
CN113610895A
Improved YOLOv11s safety helmet wearing detection model and optimization method thereof
CN120356237A