Dynamic intelligent target detection system and method based on reasoning, namely training normal form
Through the dynamic intelligent object detection system of inference, training paradigm, real-time online learning and dynamic resource adjustment, the adaptability and real-time nature of the computer vision object detection system in a dynamic environment is solved, and efficient operation and accurate detection on different hardware platforms are achieved.
Patent Information
- Application Number
- CN202510340792.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-08-12
AI Technical Summary
The existing computer vision object detection system is not adaptable enough when facing a dynamic environment, and it is difficult to meet the real-time and accuracy requirements of industrial quality inspection and other scenarios, and it is difficult to operate efficiently on different hardware platforms.
A dynamic intelligent object detection system based on the inference-based training paradigm is adopted, including a real-time learning module, a parameter fusion engine, a multi-level verification subsystem and a resource perception controller. Through real-time online learning, gradient importance sorting and dynamic resource adjustment, the model is realized in real-time update and efficient operation.
It improves the real-time and accuracy of target detection, enhances environmental adaptability, and achieves efficient operation on different hardware platforms, reducing the rate of catastrophic forgetting and resource utilization.
Smart Images

Figure CN120471138A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the intersection of computer vision and edge computing, and in particular to a dynamic intelligent target detection system and method based on a reasoning-and-training paradigm. Background Art
[0002] In the field of object detection in computer vision, traditional methods have many limitations when faced with dynamically changing environments, such as insufficient adaptability and generalization capabilities. To overcome these limitations, deep learning methods have gradually become mainstream in recent years, such as Faster R-CNN, YOLO, and SSD. These methods automatically learn features through convolutional neural networks (CNNs), can better adapt to the diverse changes of targets, and have achieved certain improvements in detection speed and accuracy.
[0003] However, with the increasing demands for real-time, accurate, and environmentally adaptable target detection in scenarios like industrial quality inspection, existing detection systems struggle to meet these demands. For example, in industrial quality inspection scenarios, the emergence of new products or defects requires detection systems to rapidly update models to ensure detection accuracy. Traditional models require large amounts of labeled data and long offline training periods, making them unable to meet real-time requirements. Furthermore, existing technologies also have shortcomings in computing resource utilization and the stability of model updates, making them difficult to implement efficiently on different hardware platforms. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a dynamic intelligent target detection system and method based on the reasoning-and-training paradigm. Through innovative architecture design and training mechanism, the real-time performance, accuracy and environmental adaptability of target detection are significantly improved, and the model can be efficiently run on different hardware platforms.
[0005] First aspect
[0006] The present invention provides a dynamic intelligent target detection system based on the inference-as-training paradigm, including a real-time learning module, a parameter fusion engine, a multi-level verification subsystem, and a resource-aware controller:
[0007] The real-time learning module is set in the edge computing node, and uses the YOLOX model to perform real-time reasoning on the input data. At the same time, an uncertainty-driven update trigger is formulated to monitor the model reasoning results in real time. When the entropy value of the detection box is greater than a preset value, the online learning is triggered to activate;
[0008] The parameter fusion engine is set up in the cloud coordination center to train models that require online learning. It uses the gradient importance ranking algorithm to calculate parameter sensitivity based on the Hessian matrix approximation, uses the difference perception aggregation protocol, and uses cosine similarity to filter abnormal node gradients. It adopts progressive knowledge fusion and achieves smooth parameter transition through momentum update.
[0009] The multi-level verification subsystem automatically verifies the trained model and sends it to the edge computing node for real-time reasoning after passing the verification;
[0010] The resource-aware controller dynamically adjusts resource allocation and gradient sparse transmission by acquiring hardware resource conditions in real time.
[0011] Furthermore, in the real-time learning module, a dynamic parameter isolation unit is set to divide the YOLOX model into a frozen backbone layer and a plasticity detection head.
[0012] Furthermore, the real-time learning module is also integrated with a deformable feature pyramid network.
[0013] Furthermore, the real-time learning module also includes a short-term memory cache that uses an LRU strategy to maintain the most recent 512 frames of high-value samples.
[0014] Furthermore, the preset value of the detection box entropy value in the real-time learning module is set to 1.2.
[0015] Furthermore, the multi-level verification subsystem specifically includes: generating multiple perturbation patterns to test the robustness of the model through a real-time adversarial verifier, comparing the spatiotemporal consistency of the detection results for 24 consecutive hours through a delayed consistency checker, and using a crowdsourcing annotation interface to reduce the end-to-end delay from annotation results to model updates.
[0016] Furthermore, the resource-aware controller specifically includes: automatically adjusting the batch size according to the GPU memory occupancy through a dynamic batch processing regulator, setting an energy efficiency optimization module for automatically switching to 4-bit quantized inference mode when powered by battery, and using DeepSpeedZero-3 technology through a communication compression unit to reduce the gradient transmission bandwidth.
[0017] Second aspect
[0018] The present invention provides a dynamic intelligent target detection method based on the inference-training paradigm, the method comprising the following steps:
[0019] Step S1: Set up a real-time learning module on the edge computing node, use the YOLOX model to perform real-time inference on the input data, and formulate an uncertainty-driven update trigger to monitor the model inference results in real time. When the entropy value of the detection box is greater than a preset value, trigger the activation of online learning;
[0020] Step S2: Set up a parameter fusion engine in the cloud coordination center to perform online training on models that require online learning. Use the gradient importance ranking algorithm to calculate parameter sensitivity based on the Hessian matrix approximation. Through the difference-aware aggregation protocol, use cosine similarity to filter abnormal node gradients, adopt progressive knowledge fusion, and achieve stable parameter transitions through momentum updates.
[0021] Step S3: The trained model is automatically verified through a multi-level verification subsystem and sent to the edge computing node after passing the verification.
[0022] Furthermore, in the real-time learning module, a dynamic parameter isolation unit is set to divide the YOLOX model into a frozen backbone layer and a plasticity detection head.
[0023] Furthermore, the real-time learning module is also integrated with a deformable feature pyramid network.
[0024] Furthermore, the real-time learning module also includes a short-term memory cache that uses an LRU strategy to maintain the most recent 512 frames of high-value samples.
[0025] Furthermore, the preset value of the detection box entropy value in the real-time learning module is set to 1.2.
[0026] Furthermore, the method also includes: obtaining hardware resource status in real time through a resource-aware controller to dynamically adjust resource allocation and perform gradient sparse transmission, specifically as follows: automatically adjusting the batch size according to the GPU memory occupancy through a dynamic batch processing regulator, setting the energy efficiency optimization module to automatically switch to 4-bit quantized inference mode when powered by battery, and using DeepSpeed Zero-3 technology through a communication compression unit to reduce the gradient transmission bandwidth.
[0027] Furthermore, step S3 specifically includes: generating multiple perturbation patterns to test the robustness of the model through a real-time adversarial verifier, comparing the spatiotemporal consistency of the detection results for 24 consecutive hours through a delay consistency checker, and using a crowdsourcing annotation interface to reduce the end-to-end delay from the annotation results to the model update.
[0028] The advantages of the present invention are:
[0029] 1. Improve real-time performance and accuracy: Online learning enables real-time model updates at the single-sample level. The environment-aware elastic learning strategy enables continuous improvement of the average precision mean indicator during continuous operation.
[0030] 2. Optimize resource utilization: Build a resource-precision self-balancing mechanism, use a dynamic batch regulator to automatically adjust the batch size based on GPU memory occupancy, and combine the energy efficiency optimization module and communication compression unit to improve resource utilization.
[0031] 3. Enhance model stability: Effectively suppress catastrophic forgetting through the gradient importance ranking algorithm, difference-aware aggregation protocol, and progressive knowledge fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0033] Figure 1 This is a schematic diagram of the architecture of a dynamic intelligent target detection system based on the inference-and-training paradigm of the present invention.
[0034] Figure 2 This is an execution flow chart of a dynamic intelligent target detection method based on the inference-and-training paradigm of the present invention.
[0035] Figure 3 This is a schematic diagram of dynamic parameter isolation of the present invention.
[0036] Figure 4 Flowchart of online learning of the present invention. DETAILED DESCRIPTION
[0037] The embodiments of the present application provide a dynamic intelligent target detection system and method based on the reasoning-and-training paradigm. Through innovative architecture design and training mechanisms, combined with offline and online model updates and multi-level verification, the system achieves the goal of both accurate and real-time target detection. It can adapt well to environmental changes and realize efficient operation of the model on different hardware platforms.
[0038] The technical solution in the embodiments of the present application has the following overall idea: providing a dynamic intelligent target detection system based on the reasoning-and-training paradigm and its implementation method, through innovative architecture design and training mechanism, designing a real-time learning module at the edge computing node to process and learn the collected data in real time, while using uncertainty-driven update triggers to determine whether online learning is needed, and saving high-value samples for model updates through short-term memory cache; designing a parameter fusion engine at the cloud coordination center to receive parameters from each edge node, using the gradient importance ranking algorithm and the difference-aware aggregation protocol for parameter fusion, and achieving smooth parameter transition through progressive knowledge fusion; the cloud-based online learning results are verified in multiple dimensions through a multi-level verification subsystem to ensure the validity of the model, and designing a resource-aware controller to dynamically adjust the computing resource allocation and model parameter transmission strategy according to the hardware resource situation.
[0039] To make the present invention more clearly understood, preferred embodiments are now described in detail below with reference to the accompanying drawings.
[0040] Example 1
[0041] like Figure 1 、 Figure 3 and Figure 4 As shown in FIG, a dynamic intelligent target detection system based on the inference-training paradigm of the present invention includes a real-time learning module, a parameter fusion engine, a multi-level verification subsystem, and a resource-aware controller:
[0042] The real-time learning module is set in the edge computing node, and uses the YOLOX model to perform real-time reasoning on the input data. At the same time, an uncertainty-driven update trigger is formulated to monitor the model reasoning results in real time. When the entropy value of the detection box is greater than a preset value, the online learning is triggered to activate;
[0043] The edge computing node is also equipped with core components, including: computing core, network connection module, and storage system. The computing core includes multi-core CPU, GPU, etc., which is used to process high-concurrency data. The network connection module supports hybrid networking of multiple network protocols, as well as edge gateways with protocol conversion and data preprocessing functions. The storage system adopts a layered storage architecture, including memory and SSD cache, to meet different data storage needs.
[0044] The parameter fusion engine is set up in the cloud coordination center to train the model that needs online learning, and uses the gradient importance ranking algorithm (by calculating the gradient information of the feature to evaluate its importance, specifically, by calculating the absolute value or square value of the gradient of each feature and using it as the importance score) to calculate the parameter sensitivity based on the Hessian matrix approximation. It uses the difference-aware aggregation protocol (i.e., a protocol based on differential privacy (DP) for protecting user privacy in distributed data collection and analysis. Other randomization mechanisms are introduced to prevent data collectors from accurately inferring the original data of individual users, thereby protecting user privacy during statistical analysis). It also uses cosine similarity to filter abnormal node gradients (for example, by setting the filtering threshold to θ = 0.78 to filter out node gradients less than the threshold). It adopts progressive knowledge fusion (by gradually increasing the complexity and depth of the fusion, integrating knowledge from different sources in stages) to smooth parameter transitions through momentum updates. That is, by introducing the concept of "momentum" and setting the momentum decay factor β = 0.95, historical gradient information is accumulated into the current update, thereby smoothing the direction of parameter updates.
[0045] The multi-level verification subsystem automatically verifies the trained model and sends it to the edge computing node for real-time reasoning after passing the verification;
[0046] The resource-aware controller dynamically adjusts resource allocation and gradient sparse transmission by acquiring hardware resource conditions in real time.
[0047] Preferably, the real-time learning module divides the YOLOX model into a frozen backbone layer and a plasticity detection head by setting a dynamic parameter isolation unit.
[0048] Preferably, the real-time learning module also integrates a deformable feature pyramid network to enhance the feature extraction capabilities of the YOLOX backbone network, further improving the model's detection accuracy and robustness, particularly when dealing with small objects and object detection in complex environments. This improvement not only improves the model's performance but also enhances its applicability and reliability in practical applications.
[0049] Preferably, the real-time learning module further comprises maintaining a short-term memory cache of the most recent 512 frames of high-value samples using an LRU strategy for use in local fine-tuning.
[0050] Preferably, the default value of the detection box entropy value in the real-time learning module is set to 1.2. The detection box entropy value is calculated by calculating the loss function and then normalizing the loss function value to a range of 0 to 2. The normalization formula here is:
[0051]
[0052] in:
[0053] x is the loss function value, x min is the minimum loss, x max is the maximum loss x scaled It is the normalized value, that is, the detection box entropy value, which ranges from 0 to 2.
[0054] Preferably, the multi-level verification subsystem specifically includes: generating multiple perturbation patterns (in this embodiment, FGSM, PGD, and general perturbations are mainly used for verification to enhance cross-sample generalization) through a real-time adversarial verifier to test model robustness; comparing the spatiotemporal consistency of continuous 24-hour test results through a delayed consistency checker; and utilizing a crowdsourcing annotation interface to reduce the end-to-end delay from annotation results to model updates to less than 5 minutes. Data with unsatisfactory verification results is re-uploaded to the cloud for re-learning.
[0055] Preferably, the resource-aware controller specifically includes: automatically adjusting the batch size (1-256) according to the GPU memory occupancy through a dynamic batch processing regulator, setting an energy efficiency optimization module for automatically switching to 4-bit quantized inference mode when powered by battery, and using DeepSpeed Zero-3 technology through a communication compression unit to reduce the gradient transmission bandwidth by 90%.
[0056] The present invention can be applied to similar vehicle-mounted scenarios: image data is collected in real time through vehicle sensors, and the real-time learning module of the edge computing node quickly processes the data. When an obstacle is detected ahead, a decision is made quickly based on the confidence threshold (such as 0.7) to control the vehicle to brake or avoid, that is, the reasoning process of the real-time learning module. When model updating and optimization are required, the data during driving is used to activate online learning through uncertainty-driven update triggers. The parameter fusion engine of the cloud coordination center regularly fuses the parameters uploaded by each vehicle and optimizes the model. When online learning is required, the optimized model is sent to the edge side for model updating, thereby improving the accuracy and safety of detection.
[0057] Example 2
[0058] like Figures 2 to 4 As shown, the present invention provides a dynamic intelligent target detection method based on the inference-training paradigm, the method comprising the following steps:
[0059] Step S1: Set up a real-time learning module on the edge computing node, use the YOLOX model to perform real-time inference on the input data, and formulate an uncertainty-driven update trigger to monitor the model inference results in real time. When the entropy value of the detection box is greater than a preset value, trigger the activation of online learning;
[0060] Step S2: Set up a parameter fusion engine in the cloud coordination center to perform online training on models that require online learning. Use the gradient importance ranking algorithm to calculate parameter sensitivity based on the Hessian matrix approximation. Through the difference-aware aggregation protocol, use cosine similarity to filter abnormal node gradients, adopt progressive knowledge fusion, and achieve stable parameter transitions through momentum updates.
[0061] Step S3: The trained model is automatically verified through a multi-level verification subsystem and sent to the edge computing node after passing the verification.
[0062] Preferably, the real-time learning module divides the YOLOX model into a frozen backbone layer and a plasticity detection head by setting a dynamic parameter isolation unit.
[0063] Preferably, the real-time learning module also integrates a deformable feature pyramid network to enhance the feature extraction capabilities of the YOLOX backbone network, further improving the model's detection accuracy and robustness, particularly when dealing with small objects and object detection in complex environments. This improvement not only improves the model's performance but also enhances its applicability and reliability in practical applications.
[0064] Preferably, the real-time learning module further comprises maintaining a short-term memory cache of the most recent 512 frames of high-value samples using an LRU strategy for use in local fine-tuning.
[0065] Preferably, the default value of the detection box entropy value in the real-time learning module is set to 1.2. The detection box entropy value is calculated by calculating the loss function and then normalizing the loss function value to a range of 0 to 2. The normalization formula here is:
[0066]
[0067] in:
[0068] x is the loss function value, x min is the minimum loss, x max is the maximum loss x scaled It is the normalized value, that is, the detection box entropy value, which ranges from 0 to 2.
[0069] Preferably, the method further includes: obtaining hardware resource status in real time through a resource-aware controller to dynamically adjust resource allocation and perform gradient sparse transmission, specifically as follows: automatically adjusting the batch size (1-256) according to the GPU memory occupancy through a dynamic batch regulator, setting the energy efficiency optimization module to automatically switch to 4-bit quantized inference mode when powered by battery, and using DeepSpeedZero-3 technology through a communication compression unit to reduce the gradient transmission bandwidth by 90%.
[0070] Preferably, step S3 specifically includes: generating multiple perturbation patterns (in this embodiment, FGSM, PGD, and general perturbations are mainly used for verification to enhance cross-sample generalization) through a real-time adversarial verifier to test model robustness, comparing the spatiotemporal consistency of 24-hour test results through a delayed consistency checker, and utilizing a crowdsourcing annotation interface to reduce the end-to-end delay from annotation results to model updates to less than 5 minutes. Data with unsatisfactory verification results is re-uploaded to the cloud for re-learning.
[0071] The technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0072] Improved real-time performance and accuracy: This system enables real-time model updates at the single-sample level. Through its flexible, context-aware learning strategy, it continuously improves the mean average precision (mAP) metric over time. On the COCO 2017 dataset, the detection accuracy of new object categories improved by 37.5% within 48 hours compared to the traditional YOLOX model.
[0073] Optimized resource utilization: A resource-precision self-balancing mechanism is built. The dynamic batch controller automatically adjusts the batch size based on GPU memory usage. The energy efficiency optimization module switches to 4-bit quantized inference mode when powered by battery. The communication compression unit uses DeepSpeed Zero-3 technology to reduce gradient transmission bandwidth by 90%, ensuring availability across different hardware platforms while increasing energy consumption by only 15%.
[0074] Enhanced model stability: Through the gradient importance ranking algorithm, difference-aware aggregation protocol and progressive knowledge fusion, catastrophic forgetting is effectively suppressed. Compared with the system before improvement, the catastrophic forgetting rate of the improved system is reduced from 23% to 4.7% in 30 days, and the model volume growth rate is reduced from 2.1MB / day to 0.3MB / day.
[0075] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A dynamic intelligent target detection system based on the inference-and-train paradigm, characterized by: Includes a real-time learning module, a parameter fusion engine, a multi-level verification subsystem, and a resource-aware controller: The real-time learning module is set in the edge computing node, and uses the YOLOX model to perform real-time reasoning on the input data. At the same time, an uncertainty-driven update trigger is formulated to monitor the model reasoning results in real time. When the entropy value of the detection box is greater than a preset value, the online learning is triggered to activate; The parameter fusion engine is set up in the cloud coordination center to train models that require online learning. It uses the gradient importance ranking algorithm to calculate parameter sensitivity based on the Hessian matrix approximation, uses the difference perception aggregation protocol, and uses cosine similarity to filter abnormal node gradients. It adopts progressive knowledge fusion and achieves smooth parameter transition through momentum update. The multi-level verification subsystem automatically verifies the trained model and sends it to the edge computing node for real-time reasoning after passing the verification; The resource-aware controller dynamically adjusts resource allocation and gradient sparse transmission by acquiring hardware resource conditions in real time.
2. The dynamic intelligent target detection system based on the inference-and-training paradigm according to claim 1, characterized in that: In the real-time learning module, a dynamic parameter isolation unit is set to divide the YOLOX model into a frozen backbone layer and a plasticity detection head.
3. The dynamic intelligent target detection system based on the inference-and-training paradigm according to claim 1, characterized in that: The real-time learning module is also integrated with a deformable feature pyramid network.
4. The dynamic intelligent target detection system based on the inference-and-training paradigm according to claim 1, characterized in that: The real-time learning module also includes a short-term memory cache that uses an LRU strategy to maintain the most recent 512 frames of high-value samples.
5. The dynamic intelligent target detection system based on the inference-and-training paradigm according to claim 1, characterized in that: The preset value of the detection box entropy value in the real-time learning module is set to 1.
2.
6. The dynamic intelligent target detection system based on the inference-and-training paradigm according to claim 1, characterized in that: The multi-level verification subsystem specifically includes: generating multiple perturbation patterns to test the robustness of the model through a real-time adversarial verifier, comparing the spatiotemporal consistency of 24-hour continuous detection results through a delayed consistency checker, and using a crowdsourcing annotation interface to reduce the end-to-end delay from annotation results to model updates.
7. The dynamic intelligent target detection system based on the inference-and-training paradigm according to claim 1, characterized in that: The resource-aware controller specifically includes: automatically adjusting batch size according to GPU memory occupancy through a dynamic batch processing regulator, setting an energy efficiency optimization module for automatically switching to 4-bit quantized inference mode when powered by battery, and using DeepSpeedZero-3 technology through a communication compression unit to reduce gradient transmission bandwidth.
8. A dynamic intelligent target detection method based on the inference-and-training paradigm, characterized by: The method comprises the following steps: Step S1: Set up a real-time learning module on the edge computing node, use the YOLOX model to perform real-time inference on the input data, and formulate an uncertainty-driven update trigger to monitor the model inference results in real time. When the entropy value of the detection box is greater than a preset value, trigger the activation of online learning; Step S2: Set up a parameter fusion engine in the cloud coordination center to perform online training on models that require online learning. Use the gradient importance ranking algorithm to calculate parameter sensitivity based on the Hessian matrix approximation. Through the difference-aware aggregation protocol, use cosine similarity to filter abnormal node gradients, adopt progressive knowledge fusion, and achieve stable parameter transitions through momentum updates. Step S3: The trained model is automatically verified through a multi-level verification subsystem and sent to the edge computing node after passing the verification.
9. The dynamic intelligent target detection system based on the inference-and-training paradigm according to claim 1, characterized in that: The method also includes: obtaining hardware resource status in real time through a resource-aware controller to dynamically adjust resource allocation and perform gradient sparse transmission, specifically as follows: automatically adjusting batch size according to GPU memory occupancy through a dynamic batch regulator, setting an energy efficiency optimization module to automatically switch to 4-bit quantized inference mode when powered by battery, and using DeepSpeed Zero-3 technology through a communication compression unit to reduce gradient transmission bandwidth.
10. The dynamic intelligent target detection system based on the inference-and-training paradigm according to claim 1, characterized in that: The step S3 specifically includes: generating multiple perturbation patterns to test the robustness of the model through a real-time adversarial verifier, comparing the spatiotemporal consistency of the detection results for 24 consecutive hours through a delay consistency checker, and using a crowdsourcing annotation interface to reduce the end-to-end delay from annotation results to model updates.
Citation Information
Cited By
Standardized dynamic information intelligent detection system
CN122286375A