A closed-loop self-optimization design system and method for a video reconnaissance system based on multiple domestic hardware platforms

CN122824875APending Publication Date: 2026-09-25江苏和正特种装备有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611093581.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

这些技术虽然能够解决单一环节的问题,但在多国产化硬件平台的视频侦察系统中仍存在明显不足

Benefits of technology

[0198]1.通过ai-compute-caps算力特征字段将驱动适配结果转化为算法选择可直接使用的结构化输入,减少静态配置和人工适配导致的偏差。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824875A_ABST
    Figure CN122824875A_ABST
Patent Text Reader

Abstract

The application provides a kind of video reconnaissance system closed loop self-optimizing design system and method based on multiple domestic hardware platforms, the system includes: hardware adaptation unit, for scanning hardware, matching drive, registering uniform interface and generating computing power characteristic field;Video pre-processing unit, for video source access, protocol resolution, decoding and format conversion, output quality index;Adaptive decision unit, combined with computing power, task demand and quality index, establish state and action space, generate scheduling strategy based on dynamic probability correction Markov decision process;Efficiency evaluation unit, collect performance, accuracy and resource index, calculate efficiency by entropy weight-TOPSIS method;Decision optimization unit, according to the evaluation result, correct algorithm parameters and state transition matrix, and write back strategy.The application effectively reduces the multi-platform adaptation cost by constructing a whole-process closed loop, improves the strategy consistency and running stability of domestic platform deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of video image reconnaissance system design and edge intelligent computing, specifically involving a closed-loop self-optimization design system and method for video reconnaissance systems based on multiple domestically produced hardware platforms. Background Technology

[0002] Video reconnaissance systems typically need to perform tasks such as target reconnaissance, target tracking, abnormal behavior early warning, situational awareness stitching, video fusion, video compression, and intelligence output in complex environments. With the increasing application of domestically produced hardware platforms in reconnaissance equipment, the same video processing system often needs to be deployed on devices with different processor architectures and computing power configurations. For example, some platforms are primarily CPU-based, suitable for lightweight control and traditional image processing; some are equipped with GPUs, suitable for parallel image enhancement, deep learning detection, or multi-channel video processing; some are equipped with NPUs, suitable for fixed-point quantization model inference; and some platforms need to connect to different types of video sources via Serial High-Speed ​​Input / Output Interface (SRIO), Real-Time Streaming Protocol (RTSP), Real-Time Messaging Protocol (RTMP), User Datagram Protocol (UDP), Linux Video Acquisition Interface Version 2 (V4L2), or a local bus.

[0003] Existing solutions typically accomplish several individual functions. For example, hardware adaptation is achieved through device trees, driver plugins, or hardware abstraction layers; video preprocessing is performed through wavelet denoising, filtering enhancement, or format conversion; task scheduling is accomplished through fixed rules, reinforcement learning, or heuristic algorithms; and performance evaluation is achieved through statistics on processing latency, frame rate, recognition accuracy, and resource utilization. While these technologies can solve problems in individual stages, they still have significant shortcomings in video reconnaissance systems using multiple domestically produced hardware platforms.

[0004] First, there is a lack of coordination between hardware adaptation and algorithm selection. After the driver is loaded, the upper-layer algorithm module usually only knows whether a certain type of hardware is available, and it is difficult to obtain information such as integer inference computing power, floating-point inference computing power, batch processing latency, maximum concurrent inference paths, operator subset version, video memory capacity, storage bandwidth, and current available load in a timely manner. If algorithm selection still relies on static configuration files, high-precision algorithms may be assigned to unsuitable hardware, or adjustments may not be made in a timely manner when hardware availability changes.

[0005] Second, there is a lack of coordination between video preprocessing and the task scenario. Video denoising, enhancement, and format conversion are often set as fixed pre-processing steps, making it difficult to adjust preprocessing parameters according to changes in scenarios such as wartime, peacetime, fault conditions, low light, interference, and rapid maneuvering. For example, in target tracking tasks, excessive denoising may weaken target edge features; in low-light reconnaissance, insufficient denoising may lead to inadequate input quality for the target detection model. If the preprocessing results cannot be included in the performance evaluation, the system cannot determine whether the current denoising intensity is truly beneficial to task completion.

[0006] Third, task scheduling models are not sufficiently adaptable to environmental fluctuations. General cloud-edge collaborative or video processing scheduling methods often assume relatively stable state transition patterns. However, reconnaissance platforms may be affected by factors such as electromagnetic interference, platform maneuvering, load switching, temperature changes, power consumption limitations, and sudden task interruptions, causing continuous fluctuations in hardware load and processing capacity. Fixed state transition probabilities or fixed weight scheduling strategies can easily cause the model to gradually deviate from the actual environment.

[0007] Fourth, there is a lack of a closed loop between performance evaluation and strategy optimization. Many systems can generate performance reports, but these reports are only for human review and cannot automatically correct the "hardware-task-algorithm" mapping, nor can they update the probability matrix or reward function parameters of the scheduling model. If the evaluation results cannot influence the next round of scheduling, continuous self-optimization capabilities cannot be formed.

[0008] Therefore, a closed-loop approach is needed that integrates hardware adaptation, video preprocessing, adaptive decision-making, performance evaluation, and decision optimization, enabling the system to not only run on different domestic platforms but also continuously adjust its strategies based on computing power, video quality, task effectiveness, and resource consumption during operation. Summary of the Invention

[0009] Purpose of the invention: To address the following problems, this invention provides a closed-loop self-optimization design system and method for video reconnaissance systems based on multiple domestically produced hardware platforms:

[0010] 1. How to provide the real-time computing power of domestically produced CPUs, GPUs, NPUs and other hardware to the upper-layer algorithm selection module in the form of structured fields.

[0011] 2. How to make video preprocessing parameters change with task type and scene status, and how to feed back the preprocessing quality results to the performance evaluation unit.

[0012] 3. How to adjust the scheduling model when hardware performance fluctuates, task queues change, and video quality changes, so as to avoid the failure of fixed probability or fixed weight strategies.

[0013] 4. How to automatically convert performance evaluation results into mapping relationship correction, parameter correction, and scheduling probability matrix update, so that the system can form a closed-loop self-optimization during operation.

[0014] The system includes:

[0015] The hardware adaptation unit is used to scan domestic hardware platforms, match drivers, register unified hardware interfaces, and generate the computing power feature description field ai-compute-caps.

[0016] The video preprocessing unit is used to access different video sources, complete protocol decryption, decapsulation, decoding, format conversion and scene adaptive preprocessing, and output video quality indicators;

[0017] The adaptive decision unit is used to parse task instructions, read computing power characteristics, task requirements and video quality indicators, establish state space, action space and reward function, and generate scheduling strategies based on dynamic probability correction of Markov Decision Process (MDP).

[0018] The performance evaluation unit is used to evaluate acquisition and processing performance, task accuracy, video quality, and resource utilization. It calculates system performance results using the TOPSIS (Topology-Approximation Ideal Solution Ranking Method) or an equivalent multi-index evaluation method.

[0019] The decision optimization unit is used to correct the mapping relationship library, algorithm parameter range and state transition probability matrix based on the performance evaluation results, and write the corrected strategy back to the adaptive decision unit.

[0020] This invention also provides a closed-loop self-optimization design method for a video reconnaissance system based on multiple domestically produced hardware platforms, comprising the following steps:

[0021] Step S1: After the system starts up, the hardware adaptation unit performs a hardware scan on the current platform to identify the central processing unit (CPU), graphics processing unit (GPU), neural network processor (NPU), video encoding / decoding unit, image processing unit, storage unit, and video access interface, and matches and loads the corresponding drivers.

[0022] Step S2: After the driver is successfully loaded, the hardware adaptation unit registers the unified hardware interface and generates the computing power feature description field ai-compute-caps. The computing power feature description field ai-compute-caps is written to the computing power feature cache or the unified hardware interface directory for the adaptive decision unit to read.

[0023] Step S3: The video preprocessing unit receives video data from the serial high-speed input / output interface SRIO, real-time streaming protocol RTSP, real-time message transmission protocol RTMP, Linux video capture interface version 2 V4L2, user datagram protocol UDP, or local video interface, and performs protocol de-protocol, de-encapsulation, decoding, format conversion, and scene adaptive preprocessing on the video stream.

[0024] Step S4: The adaptive decision unit receives the task instruction, parses the task type, priority, latency requirements, accuracy requirements and output format, reads the computing power characteristics, video quality indicators and current hardware load, and generates candidate algorithms and parameter combinations.

[0025] Step S5: The adaptive decision unit models the current state as a state vector and generates scheduling actions based on the dynamic probability correction MDP. The scheduling actions include resource allocation, task migration, algorithm switching, parameter adjustment or priority adjustment.

[0026] Step S6: The performance evaluation unit collects and processes latency, frame rate, accuracy, tracking success rate, video quality, CPU utilization, GPU utilization, NPU utilization, video memory usage, memory usage, and storage bandwidth usage, and calculates the system performance evaluation results.

[0027] Step S7: When the performance evaluation result is lower than the preset condition, or when a downward trend occurs for two or more consecutive periods, the decision optimization unit corrects the mapping relationship library, the algorithm parameter range and the state transition probability matrix, and writes the correction result back to the adaptive decision unit.

[0028] In step S8, the adaptive decision-making unit uses the updated strategy in the next round of task scheduling, thereby forming a closed loop of hardware recognition, video preprocessing, task scheduling, performance evaluation, strategy correction, and rescheduling.

[0029] In step S1, the hardware capabilities are represented as a vector:

[0030] ,

[0031] in, Indicates the first Capability vector of each hardware unit; This indicates 8-bit integer (INT8) reasoning capability; This indicates 16-bit floating-point FP16 computing power; Indicates the minimum latency for batch processing; Indicates the maximum number of concurrent inference paths; Indicates storage bandwidth; Indicates the operator set version or the operator's supported encoding; This indicates the current load.

[0032] In step S1, during driver matching, the system first performs an exact match: assuming the hardware identifier obtained from the scan is... The first in the device tree The compatibility identifier for each node is Then the exact matching function Defined as:

[0033] ,

[0034] If no exact matching node is found, a general matching is performed. The general matching calculates a matching score based on the vendor, architecture, bus type, and functional category. :

[0035] ,

[0036] The system selects the highest-scoring system that exceeds the threshold. The node serves as a general-purpose driver node:

[0037] ,

[0038] Among them I vendor Indicates whether the manufacturer is compatible, I arch Indicates whether the architecture matches, I bus Indicates whether the bus type matches, I func Indicates whether the functional categories match; a value of 1 indicates a match, and a value of 0 indicates a mismatch. ω1 represents the weighting coefficient for vendor matching, ω2 represents the weighting coefficient for architecture matching, ω3 represents the weighting coefficient for bus matching, and ω4 represents the weighting coefficient for functional category matching. min S represents the scoring threshold for general matching. n This represents the general matching score of the nth node, where n is the nth node. * This indicates the node number with the highest score;

[0039] The system creates a hash index using the compatible field:

[0040] ,

[0041] Where compatible represents the hardware compatibility identifier field, Hash represents the hash function, and Index represents the hash index value calculated from the compatibility identifier using the hash function;

[0042] Once the hardware scan obtains a compatibility identifier, the system first locates candidate nodes using a hash index, and then performs an exact match or a general match.

[0043] In step S3, let the input video frame be... , Frame attributes for:

[0044] ,

[0045] in, and These represent the width and height of the resolution, respectively. Indicates frame format, Indicates frame rate, Represents a timestamp. Indicates the video source type;

[0046] The video preprocessing unit converts the input frames into the luminance / chrominance color format YUV, the red-green-blue color format RGB, or the blue-green-red color format BGR.

[0047] For video frames requiring noise reduction, the system employs wavelet transform for noise reduction; for YUV420 video frames, the system extracts the luminance component. Perform second-level wavelet decomposition:

[0048] ,

[0049] Among them, Y t This represents the luminance component of the t-th frame of the video; YUV, RGB, and BGR are all color encoding formats for video images, among which YUV420 is a YUV chroma subsampling format;

[0050] In the above formula, the symbol → indicates the brightness component Y to the left of the arrow. t Performing a second-level wavelet decomposition operation yields the wavelet coefficients within the curly braces to the right of the arrow, signifying "decomposition into". This arrow symbol differs in meaning from the arrow symbol representing "data flow direction" in the closed-loop data flow formula of this invention, depending on the formula in which they appear.

[0051] LL2 represents the low-frequency approximation coefficients of the second-order wavelet decomposition; LH2 represents the second-order horizontal high-frequency coefficients, HL2 represents the second-order vertical high-frequency coefficients, HH2 represents the second-order diagonal high-frequency coefficients; LH1 represents the first-order horizontal high-frequency coefficients, HL1 represents the first-order vertical high-frequency coefficients, and HH1 represents the first-order diagonal high-frequency coefficients.

[0052] The noise intensity is estimated using the following formula:

[0053] ,

[0054] in denoted as the noise intensity estimate at time t, median indicates median operation, HH1 represents the first-order diagonal high-frequency coefficient, and the constant 0.6745 is the coefficient for converting the absolute deviation of the median to the standard deviation under the Gaussian distribution.

[0055] soft threshold Defined as:

[0056] ,

[0057] Among them, coefficient Defined as:

[0058] ,

[0059] in, The basic threshold coefficient; Indicates the scene intensity factor; Indicates the current degree of video quality degradation; Indicates the strength of the task's real-time constraint; , , This is the adjustment coefficient;

[0060] For any high-frequency coefficient Soft thresholding function for:

[0061] ,

[0062] Where sign represents the sign function;

[0063] After processing the high-frequency coefficients, the system performs wavelet reconstruction to obtain the denoised luminance components. Then, with color components , Synthesized output frames :

[0064] , in U represents the luminance component after noise reduction and reconstruction. t V represents the blue chromaticity component of the t-th frame. t This represents the red chroma component of the t-th frame. Indicates the luminance component after noise reduction With chromaticity component U t V t The output video frames are recombined using the synthesis function Combine;

[0065] Calculate the peak signal-to-noise ratio (PSNR): , , Among them, MSE t W represents the mean square error between the original brightness and the brightness after noise reduction in frame t. t H represents the width of the frame (in pixels).t Y represents the height of the frame (in pixels). t (x,y) represents the pixel value of the original luminance component at coordinates (x,y). This represents the pixel value of the luminance component at coordinates (x, y) after noise reduction, where x represents the horizontal coordinate of the pixel and y represents the vertical coordinate of the pixel; PSNR t MAX represents the peak signal-to-noise ratio of frame t. Y This indicates the maximum value of the luminance component; for 8-bit video, it is typically 255.

[0066] In step S4, the system establishes a mapping relationship database for hardware characteristics, task types, algorithms, and parameter ranges. The m-th record r in the mapping relationship database... m Represented as:

[0067] ,

[0068] in, Indicates hardware conditions. Indicates the task type. This refers to a recommendation algorithm. Indicates the parameter range. This represents historical performance statistics;

[0069] When selecting an algorithm, the system first filters candidate records based on task type, and then calculates the fit score based on hardware capability vectors and task constraints.

[0070] ,

[0071] The system selects the candidate algorithm and hardware combination with the highest score that meets the task constraints; the score represents the overall suitability of the candidate algorithm and hardware combination for the current task, with a higher value indicating a higher degree of suitability; Fit compute Indicates the degree of matching in computing power, Fit latency Indicates the degree of delay satisfaction, Fit accuracy Indicates the degree of precision achieved, Fit resource μ1 represents the weight of the degree of resource utilization suitability; μ2 represents the weight of the degree of computing power matching; μ3 represents the weight of the degree of latency satisfaction; and μ4 represents the weight of the degree of accuracy satisfaction.

[0072] The confidence threshold of the object detection algorithm is expressed as: The input size is represented as The tracking algorithm search window is represented as The parameter selection target is:

[0073] ,

[0074] Where Ω represents the candidate set of parameters, J(Θ) represents the comprehensive performance function under the current task and hardware conditions, and Θ * This represents the optimal combination of parameters that maximizes the overall performance function.

[0075] In step S5, the adaptive decision-making unit models the task scheduling process as a Markov decision process (MDP):

[0076] ,

[0077] in, For state space, For the action space, Let be the state transition probability. For the reward function, Discount factor;

[0078] The state vector can be defined as a 12-dimensional continuous vector s t :

[0079] ,

[0080] in, Indicates CPU utilization. Indicates GPU load. Indicates NPU utilization. This indicates the amount of remaining GPU memory. Indicates the amount of memory remaining. Indicates storage bandwidth. , , These represent the number of high, medium, and low priority tasks, respectively. This indicates the average latency of the current task. Indicates the length of the task queue. This indicates the timeout rate for unfinished tasks;

[0081] The system introduces a scene adaptation coefficient matrix. Let the first The basic weights for each state dimension are: The current scenario's influencing factor is , No. The sensitivity of each dimension in the current scenario is: The dynamic weights are then:

[0082] ,

[0083] Where i represents the index of the state dimension, w i 0 φ represents the basic weight of the i-th state dimension. s κ represents the influencing factor of the current scenario.i,s w represents the sensitivity of the i-th state dimension in the current scenario. i t This represents the dynamic weight of the i-th state dimension after adjustment in the current scenario at time t;

[0084] The state fusion vector is defined as:

[0085] ,

[0086] in, Represents the normalization function;

[0087] in, This represents the state fusion value of the i-th state dimension at time t, which is equal to the dynamic weight w. i t With normalized state components Norm(s) i,t The product of ); s i,t This represents the original value of the i-th state dimension at time t, and Norm(·) represents the normalization function;

[0088] Action space is defined as a set of discrete actions, which are divided into resource allocation actions, task migration actions, algorithm adjustment actions, and parameter adjustment actions.

[0089] Resource allocation actions include adjusting the resource ratios of CPU, GPU, and NPU:

[0090] ,

[0091] Where A resource This represents a set of resource allocation actions. In this set, CPU, GPU, and NPU represent the corresponding processors, and their subscript numbers (20, 40, 60) indicate that the resource usage of that processor is adjusted to the corresponding percentage. For example, CPU... 20 This indicates that the CPU resource usage ratio will be adjusted to 20%, and the GPU... 20 This indicates that the GPU resource usage ratio will be adjusted to 20%, NPU 20 This indicates that the NPU resource usage ratio will be adjusted to 20%.

[0092] Task migration operations include migrating high-priority or medium-priority tasks from one hardware unit to another;

[0093] Algorithm adjustments include switching object detection models, object tracking algorithms, image segmentation algorithms, or video compression strategies;

[0094] Parameter adjustment actions include adjusting the confidence threshold, input resolution, tracking search window, noise reduction threshold, inter-frame sampling interval, and batch size;

[0095] Introducing a two-factor penalty coefficient based on scenario and task:

[0096] ,

[0097] in, Indicates time The dynamic penalty coefficient; Basic penalty coefficient; As a scene penalty factor; Based on the urgency of the situation; As a task penalty factor; This represents the percentage of current task types or the percentage of critical tasks.

[0098] Reward function R t Defined as:

[0099] ,

[0100] in, Rewards are given based on task completion rate. Incentives for resource utilization As a reward for low latency, As a reward for video quality, This is a timeout penalty item. , , , Dynamic weights;

[0101] Dynamic weights are determined by the task queue and resource status:

[0102] ,

[0103] ,

[0104] ,

[0105] ,

[0106] in, This indicates the percentage of high-priority tasks. Indicates the percentage of real-time tasks. and These represent GPU resource utilization and NPU resource utilization, respectively. Indicates storage bandwidth usage. Indicates the average delay. Indicates task latency constraints. Indicates the current video quality. This indicates the video quality requirement; f1 represents the percentage of high-priority tasks ρ. high and the proportion of real-time tasks ρ realCalculate the task completion rate reward weight α t The mapping function, f2, represents the mapping function based on GPU utilization U. gpu NPU utilization rate U npu and storage bandwidth usage B store Calculate the reward weight β for resource utilization t The mapping function, f3, represents the average time delay. and task delay constraint L bound Calculate the low-latency reward weight γ t The mapping function, f4, represents the mapping function based on the current video quality Q. t And video quality requirements Q bound Calculate the video quality reward weight ζ t The mapping functions mentioned above are all monotonically normalized functions whose output values ​​are between 0 and 1.

[0107] Task completion rate reward R comp Represented as:

[0108] ,

[0109] in, Number of task types As a weight for task type, For the first Number of tasks completed by class For the first Total number of tasks of each type;

[0110] Low latency reward R latency Represented as:

[0111] ,

[0112] Among them, l t L represents the current task delay. bound Indicates task latency constraints;

[0113] Video quality bonus R quality Represented as:

[0114] ,

[0115] Where R quality Indicates a video quality reward, Q t Q indicates the current video quality. bound Indicates video quality requirements, Q min This represents the minimum acceptable video quality, i.e., the lower limit of video quality when the video quality bonus is 0.

[0116] Resource utilization rate reward R util Represented as:

[0117] ,

[0118] in, For the number of hardware units, For the first Utilization of each hardware unit Use the center of the interval to target the objective;

[0119] The system constructs the initial state transition probability based on historical data. :

[0120] ,

[0121] in, Represents the state in historical data Execute action After transitioning to state Number of times, Indicates the state in historical data Execute action Total number of times;

[0122] Introducing environmental fluctuation coefficient :

[0123] ,

[0124] in, Represents the environment vector at time t; These are the normalization coefficients; x is the maximum correction factor; t-1 Let represent the environment vector at the previous moment; ||·||2 represents the L2 norm of the vector, which is the square root of the sum of the squares of the vector's components (i.e., the Euclidean modulus).

[0125] Corrected state transition probabilities for:

[0126] ,

[0127] in, This represents the compensation probability distribution under the current environment;

[0128] For the sliding window length Recent probability Represented as:

[0129] ,

[0130] in, Indicates recent Statistical frequency within a period For smoothing coefficients, Let the number of states be defined as follows:

[0131] .

[0132] In step S6, let there be a total of Each evaluation object, There are 1 evaluation index, and the original index matrix X is:

[0133] ,

[0134] in, Indicates the first The evaluation object is in the first The original values ​​of each indicator; m represents the number of evaluation objects, and n represents the number of evaluation indicators;

[0135] For positive indicators, the standardization is as follows:

[0136] ,

[0137] Among them, z ij This represents the standardized value of the i-th evaluation object on the j-th indicator;

[0138] For negative indicators, the standardization is as follows:

[0139] ,

[0140] For interval-type indicators, the definition is:

[0141] ,

[0142] Where, x j target D represents the center of the target interval for the j-th indicator. j This represents the maximum allowable deviation of the j-th indicator. Positive indicators indicate that the larger the value, the better; negative indicators indicate that the smaller the value, the better; and interval indicators indicate that the optimal indicator is when the value falls within the target interval.

[0143] When calculating entropy weights, first calculate the first... The first indicator The proportion of each evaluation object :

[0144] ,

[0145] No. Information entropy of each indicator e j for:

[0146] ,

[0147] Coefficient of difference gj for:

[0148] ,

[0149] Indicator weights for:

[0150] ,

[0151] After obtaining the weights, construct the weighted normalization matrix:

[0152] ,

[0153] Where v ij This represents the weighted standardized value of the i-th evaluation object on the j-th indicator, which is equal to the weight ω of the j-th indicator. j With standardized value z ij The product;

[0154] Positive Ideal Solution and negative ideal solution They are respectively:

[0155] ,

[0156] No. Distance from each evaluation object to the ideal solution Distance to the negative ideal solution for:

[0157] ,

[0158] ,

[0159] Relative closeness for:

[0160] ,

[0161] The system will As the effectiveness score of the current strategy, if Below the threshold , or continuous If the evaluation period decreases, optimization is triggered:

[0162] ,

[0163] Where C i C represents the relative proximity of the i-th evaluated object and serves as the effectiveness score of the current strategy; min Indicates the performance score threshold; r represents the number of evaluation periods with consecutive declines; Trigger tThis represents the optimization trigger flag at time t. A value of 1 triggers optimization, while a value of 0 does not trigger optimization.

[0164] The trigger result enters the decision optimization unit.

[0165] In step S7, the decision optimization unit determines the problem type based on the performance evaluation results, and the mapping relationship library is corrected as follows:

[0166] ,

[0167] in, Indicates the first The mapping relationship at time Historical performance estimates Historical retention factor;

[0168] If the mapping relationship is continuously below the threshold :

[0169] ,

[0170] The system lowers the recommendation priority of the mapping relationship or narrows the parameter range; if another candidate mapping relationship is more efficient in similar tasks, the recommendation priority is increased.

[0171] The parameter range correction is performed as follows: Let the algorithm parameters be... The current recommendation range is The current optimal parameters are If the current strategy scores high, then focus on... Narrowing the search scope:

[0172] ,

[0173] ,

[0174] Where, θ min θ represents the lower limit of the recommended range for the parameter. max Δ represents the upper limit of the recommended range for the parameter. θ This indicates a narrowing step size for the parameter range. This indicates the lower limit of the corrected parameter range. Indicates the upper limit of the corrected parameter range;

[0175] If the current strategy score is low, expand the search scope or switch candidate algorithms;

[0176] For long-term strategy prediction, a regression model is established based on historical performance data, with the feature vector z denoted as z. t for:

[0177] ,

[0178] in, Let be the hardware capability vector at time t. Let be the state vector at time t. Let be the algorithm parameters at time t, and Θ be the parameter range represented in the mapping database. m The subscripts 't' and 'm' are used to distinguish them; they have different meanings. The task code at time t is the same as the task type T. m The subscripts 't' and 'm' are used to distinguish them; they have different meanings. Let be the current video quality metric; and let t be the actual performance score of the t-th sample. Establish a linear regression model with the variable as the dependent variable:

[0179] ,

[0180] Where z t This represents the eigenvector at time t. a represents the performance score predicted by the regression model. reg Let λ represent the regression coefficient vector, b represent the bias term of the regression model, and λ represent the regression coefficient vector. r Represents the regularization coefficient;

[0181] The model training objective is:

[0182] ,

[0183] in, N is the regularization coefficient. s This indicates the number of historical samples used to train the regression model. In the regression model, the superscript T indicates the transpose of the vector.

[0184] During policy write-back, the decision optimization unit writes the updated mapping relationship, parameter range, and probability matrix into the policy cache and notifies the adaptive decision unit via the policy push service. The adaptive decision unit reads the latest policy in the next round of task scheduling. The system sets a minimum update interval or a policy stability window.

[0185] ,

[0186] in, This is the time since the last strategy update. For the minimum update interval, Update t This represents the policy update flag at time t. A value of 1 indicates that a policy update is performed, and a value of 0 indicates that no update is performed; t represents the current time. last τ represents the time of the last policy update. min This indicates the minimum update interval.

[0187] In step S8, the closed-loop feedback relationship is defined by data type. The hardware adaptation unit outputs computing power parameters to the adaptive decision-making unit; the video preprocessing unit outputs video quality indicators to the adaptive decision-making unit and the performance evaluation unit; the adaptive decision-making unit outputs scheduling actions, hardware load status, and task execution results to the performance evaluation unit; the performance evaluation unit outputs performance scores and optimization trigger signals to the decision optimization unit; and the decision optimization unit writes back the mapping relationship correction results, parameter range correction results, and probability matrix update results to the adaptive decision-making unit.

[0188] This data stream is represented as:

[0189] ,

[0190] ,

[0191] ,

[0192] ,

[0193] ,

[0194] The symbol → indicates the data flow direction, meaning the data on the left side of the arrow is output from the corresponding unit and flows to the unit on the right side of the arrow; Caps t Representing computing power characteristics, Decision t This represents the scheduling decision generated by the adaptive decision unit, Quality. t Indicates video quality metrics, Action t Indicates a scheduling action, Load t Indicates load status, Result t Evaluation indicates the result of task execution. t Indicates the performance evaluation results, Optimization t Represents the policy optimization process. t+1 This indicates the scheduling strategy for the next round.

[0195] This invention enables hardware computing power characteristics, video quality indicators, task scheduling status, performance evaluation results, and strategy correction results to form a unified data link within the system, and automatically corrects algorithm selection, parameter configuration, and task scheduling strategies accordingly. This method is applicable to reconnaissance platforms that include CPUs, GPUs, NPUs, video encoding / decoding units, image processing units, SRIO interfaces, Ethernet interfaces, or other video access interfaces, and can be used in manned vehicle-mounted platforms, unmanned vehicle-mounted platforms, UAV platforms, portable reconnaissance terminals, and fixed reconnaissance nodes.

[0196] This invention particularly relates to the following technical aspects: computing power feature modeling and driver adaptation of domestic hardware platforms, access, decoding, format conversion and scene adaptive preprocessing of reconnaissance video streams, algorithm selection and task scheduling based on hardware computing power status and task requirements, system performance evaluation based on multi-index evaluation, and closed-loop strategy optimization of mapping relationship library and scheduling model based on performance results.

[0197] Compared with the prior art, the present invention has the following beneficial effects:

[0198] 1. By using the ai-compute-caps computing power feature field, the driver adaptation results are transformed into structured inputs that can be directly used by the algorithm selection, reducing the deviation caused by static configuration and manual adaptation.

[0199] 2. Through scene-adaptive preprocessing and video quality feedback, noise reduction, format conversion, and enhancement processing can be adjusted according to the task and scene, and then incorporated into the performance evaluation process.

[0200] 3. By correcting the state transition probability of the Markov Decision Process (MDP) through environmental fluctuation coefficients, the task scheduling strategy can adapt to real-time changes in hardware load, task queue, and video quality.

[0201] 4. By using comprehensive evaluation methods such as the entropy weight-approximation ideal solution ranking method TOPSIS and strategy write-back mechanism, the system can correct the mapping relation library, parameter range and probability matrix according to the performance changes, forming a continuous self-optimization capability.

[0202] 5. By organizing hardware adaptation, video preprocessing, scheduling decisions, performance evaluation, and strategy optimization into a closed loop, the adaptation cost of "one platform, one development" can be reduced, and the strategy consistency and operational stability when deploying multiple domestic platforms can be improved. Attached Figure Description

[0203] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0204] Figure 1 This is a closed-loop self-optimization structure diagram of a video reconnaissance system based on multiple domestically produced hardware platforms, illustrating the data links and strategy feedback relationships between hardware adaptation, video preprocessing, adaptive decision-making, performance evaluation, and decision optimization.

[0205] Figure 2 This is a flowchart of the method of the present invention, showing the execution process of the system from startup, hardware adaptation, video preprocessing, task scheduling, performance evaluation to strategy optimization and writeback.

[0206] Figure 3The diagram shows the feedback path between computing power parameters, video quality indicators, load status, optimization trigger signals, mapping relationship correction, and probability matrix update in this invention.

[0207] Figure 4 This is a schematic diagram comparing the closed-loop self-optimization results of Embodiment 4 of the present invention, showing the comparison results of four indicators: target tracking success rate, average processing latency, neural network processor (NPU) utilization rate, and peak signal-to-noise ratio (PSNR) when using the existing static configuration scheme and the scheme of the present invention.

[0208] Figure 5 This is a schematic diagram comparing the results of multi-task concurrent scheduling in Embodiment 5 of the present invention. It shows the comparison results of four indicators: target tracking success rate, target tracking latency, situation stitching frame rate, and neural network processor (NPU) utilization rate when using the existing static configuration scheme (single-task optimal strategy) and the scheme of the present invention (global equilibrium strategy) under the concurrent task of target tracking and situation stitching. Detailed Implementation

[0209] like Figure 1 As shown, this embodiment of the invention provides a closed-loop self-optimization design system for a video reconnaissance system based on multiple domestically produced hardware platforms, including:

[0210] The hardware adaptation unit is used to scan domestic hardware platforms, match drivers, register unified hardware interfaces, and generate the computing power feature description field ai-compute-caps.

[0211] The video preprocessing unit is used to access different video sources, complete protocol decryption, decapsulation, decoding, format conversion and scene adaptive preprocessing, and output video quality indicators.

[0212] The adaptive decision unit is used to parse task instructions, read computing power characteristics, task requirements and video quality indicators, establish state space, action space and reward function, and generate scheduling strategy based on dynamic probability correction MDP.

[0213] The performance evaluation unit is used to evaluate acquisition and processing performance, task accuracy, video quality, and resource utilization, and calculates system performance results using entropy weight-TOPSIS or equivalent multi-index evaluation methods.

[0214] The decision optimization unit is used to correct the mapping relationship library, algorithm parameter range and state transition probability matrix based on the performance evaluation results, and write the corrected strategy back to the adaptive decision unit.

[0215] The aforementioned units form a closed loop through data flow: the hardware adaptation unit outputs computing power parameters, the video preprocessing unit outputs video quality indicators, the adaptive decision-making unit outputs scheduling actions and load status, the performance evaluation unit outputs optimization trigger signals, and the decision optimization unit outputs mapping relationship correction results and probability matrix update results. This closed loop ensures that the evaluation results are not merely reports, but directly influence the next round of task scheduling.

[0216] like Figure 2 As shown, this embodiment of the invention also provides a closed-loop self-optimization design method for a video reconnaissance system based on multiple domestically produced hardware platforms, including the following steps.

[0217] Step S1: After the system starts up, the hardware adaptation unit performs a hardware scan on the current platform to identify the CPU, GPU, NPU, video encoding and decoding unit, image processing unit, storage unit and video access interface, and matches and loads the corresponding drivers.

[0218] In step S2, after the driver is successfully loaded, the hardware adaptation unit registers the unified hardware interface and generates the ai-compute-caps computing power feature field. This field is written to the computing power feature cache or the unified hardware interface directory for the adaptive decision-making unit to read.

[0219] Step S3: The video preprocessing unit accesses video data from SRIO, RTSP, RTMP, V4L2, UDP, or local video interfaces, and performs protocol de-protocol, de-encapsulation, decoding, format conversion, and scene adaptive preprocessing on the video stream.

[0220] In step S4, the adaptive decision unit receives the task instruction, parses the task type, priority, latency requirements, accuracy requirements and output format, reads the computing power characteristics, video quality indicators and current hardware load, and generates candidate algorithms and parameter combinations.

[0221] In step S5, the adaptive decision-making unit models the current state as a state vector and generates scheduling actions based on the dynamically probabilistic modified MDP. Scheduling actions include resource allocation, task migration, algorithm switching, parameter adjustment, or priority adjustment.

[0222] Step S6: The performance evaluation unit collects and processes latency, frame rate, accuracy, tracking success rate, video quality, CPU / GPU / NPU utilization, video memory usage, memory usage, and storage bandwidth usage, and calculates the system performance evaluation results.

[0223] Step S7: When the performance evaluation result is lower than the preset condition, or when a downward trend occurs for several consecutive periods, the decision optimization unit corrects the mapping relationship library, the algorithm parameter range and the state transition probability matrix, and writes the correction result back to the adaptive decision unit.

[0224] In step S8, the adaptive decision-making unit uses the updated strategy in the next round of task scheduling, thereby forming a closed loop of "hardware recognition - video preprocessing - task scheduling - performance evaluation - strategy correction - rescheduling".

[0225] The hardware adaptation unit first establishes a tree model of domestically produced hardware devices. This model can adopt a four-level structure:

[0226] Root node -> Hardware category node -> Specific model node -> Attribute node.

[0227] In the hardware device tree structure described above, the symbol "->" indicates the hierarchical inclusion relationship between nodes, that is, the node to the left of "->" contains the lower-level node to its right.

[0228] The hardware category nodes include CPU, GPU, NPU, video encoding / decoding unit, image processing unit, storage unit, and video access interface. The specific model node records the processor model, device identifier, bus type, and driver information. The attribute node records hardware parameters and capability descriptions.

[0229] The fields included in each hardware node are shown in Table 1.

[0230] Table 1

[0231]

[0232] The contents of the ai-compute-caps field are shown in Table 2.

[0233] Table 2

[0234]

[0235] To facilitate the selection of upper-level algorithms, hardware capabilities are represented as a vector:

[0236] ,

[0237] in, Indicates the first Capability vector of each hardware unit; This indicates 8-bit integer (INT8) reasoning capability; This indicates 16-bit floating-point FP16 computing power; Indicates the minimum latency for batch processing; Indicates the maximum number of concurrent inference paths; Indicates storage bandwidth; Indicates the operator set version or the operator's supported encoding; This indicates the current load.

[0238] During driver matching, the system first performs an exact match. Let the hardware identifier obtained from the scan be... The first in the device tree The compatibility identifier for each node is The exact matching function is then defined as:

[0239] ,

[0240] If no exact matching node is found, a general match is performed. The general match calculates a match score based on vendor, architecture, bus type, and functional category.

[0241] ,

[0242] in, , , , These indicate whether the vendor, architecture, bus, and function category match; to This is the weighting coefficient. The system selects the highest score that exceeds the threshold. The node serves as a general-purpose driver node:

[0243] ,

[0244] To improve matching efficiency, the system creates a hash index on the compatible field:

[0245] ,

[0246] Once the hardware scan obtains a compatibility identifier, the system first locates candidate nodes using a hash index, and then performs an exact match or a general match. This avoids linear traversal of all nodes.

[0247] After the driver is successfully loaded, the hardware adaptation unit calls the unified interface registration function to register the CPU's computing interface, GPU's parallel processing interface, NPU's inference interface, and video encoding / decoding interface as unified hardware access interfaces. Simultaneously, the computing power feature generation module writes the hardware capability vector into ai-compute-caps. After reading this field, the adaptive decision-making unit can directly determine whether a task is suitable to be executed by a particular hardware unit.

[0248] The video preprocessing unit is used to access various reconnaissance video sources. For SRIO or local bus input, the system reads the raw frames or payload transmission frames; for RTSP, RTMP, or UDP video streams, the system first performs protocol decoding and decapsulation; for advanced video coding standards such as H.264, high-efficiency video coding standards such as H.265, motion still image format MJPEG, transport stream format TS, or Moving Picture Experts Group 4 format MP4, the system performs decoding to obtain the raw video frames.

[0249] Let the input video frame be , Frame attributes for:

[0250] ,

[0251] in, and Indicates resolution, Indicates frame format, Indicates frame rate, Represents a timestamp. Indicates the video source type.

[0252] Depending on the requirements of subsequent tasks, the video preprocessing unit converts the input frames into YUV (luminance / chrominance), RGB (red / green / blue), or BGR (blue / green / red) color formats. If the object detection model requires RGB input, it is converted to RGB; if the hardware display or encoding requires YUV input, it is converted to YUV; if the algorithm library uses BGR arrangement, it is converted to BGR.

[0253] For video frames requiring noise reduction, the system can employ wavelet transform for noise reduction. Taking a YUV420 video frame as an example, the system extracts the luminance component. Perform second-level wavelet decomposition:

[0254] ,

[0255] in, It is a second-order low-frequency coefficient. , , , , , These are high-frequency coefficients at different scales and orientations. High-frequency coefficients mainly contain edge, texture, and noise information.

[0256] In the above second-order wavelet decomposition formula, the symbol → indicates the brightness component Y to the left of the arrow. t The second-level wavelet decomposition operation is performed to obtain the wavelet coefficients within the curly braces to the right of the arrow, signifying "decomposition into". This arrow symbol differs in meaning from the arrow symbol indicating "data flow direction" in the closed-loop data flow formula of this manual, depending on the formula in which they appear.

[0257] Noise intensity It can be estimated using local variance or median absolute deviation. Taking median absolute deviation as an example:

[0258] ,

[0259] soft threshold Defined as:

[0260] ,

[0261] in, It is not a fixed constant, but is determined by the scene state, task type, and real-time constraints:

[0262] ,

[0263] In the formula, The basic threshold coefficient; This represents the scene intensity factor; for example, a lower value is used for normal scenes, and a higher value is used for fault or interference scenes. Indicates the current degree of video quality degradation; Indicates the strength of the task's real-time constraint; , , This is the adjustment coefficient. This formula increases the noise reduction threshold when the noise level is high, and reduces the complexity of processing when real-time requirements are high.

[0264] For any high-frequency coefficient The soft thresholding function is:

[0265] ,

[0266] After processing the high-frequency coefficients, the system performs wavelet reconstruction to obtain the denoised luminance components. Then, with color components , Composite output frame:

[0267] ,

[0268] To ensure the preprocessing results are included in the closed-loop process, the system calculates video quality metrics. For example, it can calculate the peak signal-to-noise ratio (PSNR).

[0269] ,

[0270] ,

[0271] in, For maximum brightness, 255 is typically used for 8-bit video. If an ideal reference frame is unavailable, referenceless sharpness, local variance, edge strength, or noise estimation metrics can be used as substitutes. The preprocessing unit will... , , Or other quality indicators are written into the performance data channel so that the performance evaluation unit can determine whether the current preprocessing parameters are conducive to task execution.

[0272] The adaptive decision-making unit receives task instructions. These instructions can be in JavaScript object representation, JSON, or other structured formats, and their contents are shown in Table 3.

[0273] Table 3

[0274]

[0275] Mission types include, but are not limited to, target reconnaissance, target detection, target tracking, abnormal behavior early warning, battlefield situation stitching, reconnaissance video fusion, image segmentation, video compression, and target localization.

[0276] The system establishes a mapping database of hardware characteristics, task types, algorithms, and parameter ranges. A record in the mapping database is represented as:

[0277] ,

[0278] in, Indicates hardware conditions. Indicates the task type. This refers to a recommendation algorithm. Indicates the parameter range. This indicates historical performance statistics.

[0279] For example, for target tracking tasks, if the CPU is available but the GPU and NPU are under high load, the Kernel Correlation Filter (KCF) tracking algorithm, the Channel Spatial Reliability Tracking (CSRT) algorithm, or other lightweight tracking algorithms can be selected; for target detection tasks, if the NPU supports the operators required by the target model and has sufficient concurrent inference paths, quantized deep learning detection algorithms can be selected; for situation stitching tasks, if the GPU has sufficient video memory and parallel capabilities, GPU-accelerated stitching algorithms can be selected; for video compression tasks, processing units with hardware encoding and decoding capabilities can be prioritized.

[0280] When selecting an algorithm, the system first filters candidate records based on task type, and then calculates the fit score based on hardware capability vectors and task constraints.

[0281] ,

[0282] in, Indicates the degree of matching of computing power. Indicates the degree of delay satisfaction. Indicates the degree to which the precision is satisfied. Indicates the degree of resource utilization adaptation; to The weights are used to determine the candidate algorithm and hardware combination that have the highest score and meet the task constraints.

[0283] For algorithm parameters, the system can perform grid search, local search, or parameter selection based on historical performance within a given range of the mapping database. For example, the confidence threshold of an object detection algorithm can be expressed as... The input size can be expressed as The search window of the tracking algorithm can be represented as The parameter selection target is:

[0284] ,

[0285] in, Represents the candidate set of parameters. This represents the overall performance function under the current task and hardware conditions.

[0286] The adaptive decision unit models the task scheduling process as a Markov decision process. A standard MDP is represented as:

[0287] ,

[0288] in, For state space, For the action space, Let be the state transition probability. For the reward function, This invention introduces an environmental fluctuation coefficient as a discount factor. This allows the state transition probability to be dynamically adjusted as the environment changes.

[0289] The state vector can be defined as a 12-dimensional continuous vector:

[0290] ,

[0291] in, Indicates CPU utilization. Indicates GPU load. Indicates NPU utilization. This indicates the amount of remaining GPU memory. Indicates the amount of memory remaining. Indicates storage bandwidth. , , These represent the number of high, medium, and low priority tasks, respectively. This indicates the average latency of the current task. Indicates the length of the task queue. This indicates the timeout rate for unfinished tasks.

[0292] To adapt to different reconnaissance scenarios, the system introduces a scenario adaptation coefficient matrix. Let the first... The basic weights for each state dimension are: The current scenario's influencing factor is , No. The sensitivity of each dimension in the current scenario is: The dynamic weights are then:

[0293] ,

[0294] The above normalization process makes the sum of all weights equal to 1. For wartime scenarios, the weights of high-priority task count, GPU / NPU load, average latency, and timeout rate can be increased; for peacetime scenarios, the weights of resource utilization and stability can be increased; for fault scenarios, the weights of hardware load fluctuation, storage bandwidth, and task migration-related dimensions can be increased.

[0295] The state fusion vector is defined as:

[0296] ,

[0297] in, This represents the normalization function. Normalization can be handled differently depending on the indicator type. For indicators where larger values ​​are better:

[0298] ,

[0299] For indicators where smaller is better:

[0300] ,

[0301] The action space can be defined as a set of discrete actions. Actions can be categorized into resource allocation actions, task migration actions, algorithm adjustment actions, and parameter adjustment actions.

[0302] Resource allocation actions include adjusting the resource ratio of CPU, GPU, and NPU, for example:

[0303] ,

[0304] Task migration actions include migrating high-priority or medium-priority tasks from one hardware unit to another, such as from CPU to GPU, GPU to NPU, or NPU to CPU.

[0305] Algorithm adjustments include switching object detection models, object tracking algorithms, image segmentation algorithms, or video compression strategies. For example, when real-time performance is insufficient, a high-complexity detection model can be switched to a lightweight detection model; when video quality degrades, algorithms more robust to low-light conditions can be considered as candidates.

[0306] Parameter adjustment actions include adjusting the confidence threshold, input resolution, tracking search window, noise reduction threshold, inter-frame sampling interval, and batch size.

[0307] Reconnaissance missions exhibit significant differences in scenario and mission priority. To ensure that the reward function can express the cost of mission success or failure, this invention introduces a scenario-mission dual-factor penalty coefficient:

[0308] ,

[0309] in, Indicates time The dynamic penalty coefficient; Basic penalty coefficient; As a scene penalty factor; Based on the urgency of the situation; As a task penalty factor; This represents the percentage of current task types or the percentage of critical tasks.

[0310] For example, in wartime or fault scenarios, Higher priority tasks; higher priority tasks such as target tracking and target reconnaissance correspond to higher priority. Ordinary video compression or low-priority storage tasks correspond to lower... This design enables the system to impose greater penalties on task timeouts in high-urgency scenarios, preventing scheduling strategies from sacrificing critical tasks to improve average resource utilization.

[0311] The reward function is defined as:

[0312] ,

[0313] in, Rewards are given based on task completion rate. Incentives for resource utilization As a reward for low latency, As a reward for video quality, This is a timeout penalty item. , , , For dynamic weights.

[0314] Dynamic weights can be determined by the task queue and resource status:

[0315] ,

[0316] ,

[0317] ,

[0318] ,

[0319] in, This indicates the percentage of high-priority tasks. Indicates the percentage of real-time tasks. and Indicates GPU / NPU resource utilization. Indicates storage bandwidth usage. Indicates the average delay. Indicates task latency constraints. Indicates the current video quality. This indicates the video quality requirements.

[0320] Task completion rate reward is represented as follows:

[0321] ,

[0322] in, Number of task types As a weight for task type, For the first Number of tasks completed by class For the first Total number of tasks of each type.

[0323] The low-latency reward can be represented as:

[0324] ,

[0325] Among them, l t L represents the current task delay. bound This indicates the task delay constraint.

[0326] Video quality bonuses can be represented as:

[0327] ,

[0328] Resource utilization bonuses can be reduced when resources are too low or too high:

[0329] ,

[0330] in, For the number of hardware units, For the first Utilization of each hardware unit Use the center of the interval to target the objective.

[0331] The system constructs the initial state transition probabilities based on historical data:

[0332] ,

[0333] in, Represents the state in historical data Execute action After transitioning to state Number of times, Indicates the state in historical data Execute action The total number of times.

[0334] In actual operation, hardware load, video quality, and task queues may vary. This invention introduces an environmental fluctuation coefficient. :

[0335] ,

[0336] in, This represents the current environment vector, which can consist of hardware load, task queue length, video quality, and communication status. These are the normalization coefficients; This is the maximum correction factor. The greater the environmental change, the greater the correction factor. The larger.

[0337] In the above formula, x t-1 represents the environment vector at the previous moment; ||·|2 represents the second norm of the vector, which is the square root of the sum of the squares of the vector's components (i.e., the Euclidean modulus).

[0338] The corrected state transition probability is:

[0339] ,

[0340] in, This represents the compensation probability distribution under the current environment. If there is no reliable prior, a uniform distribution can be used. If recent operational data exists, it can be obtained through a sliding window.

[0341] For the sliding window length The recent probability can be expressed as:

[0342] ,

[0343] in, Indicates recent Statistical frequency within a period For smoothing coefficients, Let be the number of states. At this point, we can let:

[0344] ,

[0345] This design allows the system to retain historical experience while also being able to adjust scheduling strategies as the current environment changes.

[0346] The performance evaluation unit collects data from four categories of indicators: processing performance, task accuracy, video quality, and resource utilization.

[0347] Processing performance metrics include video processing latency, task processing latency, frame rate, and task queue waiting time. Task accuracy metrics include object detection accuracy, object tracking success rate, image segmentation Intersection over Union (IoU) value, and abnormal behavior recognition accuracy. Video quality metrics include PSNR, sharpness rating, noise intensity estimate, and edge preservation. Resource utilization metrics include CPU utilization, GPU utilization, NPU utilization, video memory utilization, system memory utilization, and storage bandwidth utilization.

[0348] Assume there is a total Each evaluation object, There are 10 evaluation indicators, and the original indicator matrix is ​​as follows:

[0349] ,

[0350] in, Indicates the first The evaluation object is in the first The value of each indicator. The evaluation object can be a combination of hardware, algorithm, and parameters, or the system's operating status within a certain time window.

[0351] For positive indicators, i.e., indicators where a larger value is better, the standardization is as follows:

[0352] ,

[0353] For negative indicators, i.e., indicators where smaller values ​​are better, the standardization is as follows:

[0354]

[0355] For interval-based metrics, such as optimal hardware utilization within a certain range, the following definition can be used:

[0356] ,

[0357] in, Center of the target interval This is the maximum permissible deviation. When... Take 0 at that time.

[0358] When calculating entropy weights, first calculate the first... The first indicator The proportion of each evaluation object:

[0359] ,

[0360] in, To avoid smooth terms with a denominator of 0. The information entropy of each indicator is:

[0361] ,

[0362] The coefficient of difference is:

[0363] ,

[0364] The indicator weights are:

[0365] ,

[0366] After obtaining the weights, construct the weighted normalization matrix:

[0367] ,

[0368] The positive ideal solution and the negative ideal solution are as follows:

[0369] ,

[0370] No. The distances from each evaluation object to the positive and negative ideal solutions are:

[0371] ,

[0372] ,

[0373] The relative similarity is:

[0374] ,

[0375] The larger the value, the closer the evaluated object is to the ideal state. The system can... This serves as the effectiveness score for the current strategy. If Below the threshold , or continuous If the evaluation period decreases, optimization is triggered:

[0376] ,

[0377] The trigger result enters the decision optimization unit.

[0378] The decision optimization unit determines the problem type based on the performance evaluation results. If resource utilization is too high and task latency increases, the task migration strategy is adjusted first. If video quality indicators decline and task accuracy decreases, the preprocessing parameters are adjusted first, or an algorithm that is more robust to low-quality videos is switched. If an algorithm consistently scores low on certain hardware, the recommended order and parameter range in the mapping relation library are corrected. If hardware load fluctuations cause unstable scheduling actions, the MDP state transition probability matrix is ​​corrected.

[0379] The mapping relation library correction can be represented as:

[0380] ,

[0381] in, Indicates the first The mapping relationship at time Historical performance estimates This indicates the current performance score. Historical retention coefficient.

[0382] If a certain mapping relationship is continuously below a threshold:

[0383] ,

[0384] The system lowers the recommendation priority of this mapping relationship or narrows its parameter range. If another candidate mapping relationship performs better in similar tasks, its recommendation priority is increased.

[0385] Parameter range correction can be performed in the following way. Let the parameters of a certain algorithm be... The current recommendation range is The current optimal parameters are If the current strategy scores high, then focus on... Narrowing the search scope:

[0386] ,

[0387] ,

[0388] If the current strategy score is low, expand the search scope or switch candidate algorithms.

[0389] For long-term strategy prediction, a regression model can be built based on historical performance data. Let the eigenvector be:

[0390] ,

[0391] in, For hardware capability vectors, For state vectors, For algorithm parameters, Code the task. This is a video quality metric, scored based on performance. As the dependent variable, a linear regression model can be established:

[0392] ,

[0393] The model training objective is:

[0394] ,

[0395] in, This is the regularization coefficient. For nonlinear relationships, random forests, gradient boosting trees, or lightweight neural network models can also be used, but this invention does not limit the specific prediction model. The prediction results are used to adjust the mapping relation library and the range of candidate parameters.

[0396] In the above formula, N s C represents the number of historical samples used to train the regression model. t Let represent the actual performance score of the t-th sample, and let represent the predicted value of the regression model. t To distinguish them; in the regression model, the superscript T indicates the transpose of the vector.

[0397] During policy write-back, the decision optimization unit writes the updated mapping relationships, parameter ranges, and probability matrices to the policy cache and notifies the adaptive decision unit via the policy push service. The adaptive decision unit reads the latest policy in the next round of task scheduling. To avoid frequent oscillations, the system can set a minimum update interval or a policy stability window.

[0398] ,

[0399] in, This is the time since the last strategy update. This is the minimum update interval.

[0400] like Figure 3 As shown, the closed-loop feedback relationship of this invention is bounded by data type. The hardware adaptation unit outputs computing power parameters to the adaptive decision-making unit; the video preprocessing unit outputs video quality indicators to the adaptive decision-making unit and the performance evaluation unit; the adaptive decision-making unit outputs scheduling actions, hardware load status, and task execution results to the performance evaluation unit; the performance evaluation unit outputs performance scores and optimization trigger signals to the decision optimization unit; and the decision optimization unit writes back the mapping relationship correction results, parameter range correction results, and probability matrix update results to the adaptive decision-making unit.

[0401] This data stream can be represented as:

[0402] ,

[0403] ,

[0404] ,

[0405] ,

[0406] ,

[0407] in, Indicates computing power characteristics, Indicates video quality metrics, Indicates a scheduling action. Indicates the load status. Indicates the result of task execution. Indicates the performance evaluation results. This represents the strategy optimization process. This indicates the scheduling strategy for the next round.

[0408] The closed loop described above differs from a simple linear process. In a linear process, hardware adaptation, video preprocessing, task scheduling, and performance statistics are executed sequentially, and subsequent statistical results do not change the front-end module. In this invention, performance evaluation results are written back to the mapping relation library and probability matrix, so that the next round of task scheduling is affected by the results of the previous round.

[0409] Example 1: Closed-loop self-optimization process of a single high-priority target tracking task on a vehicle-mounted domestically produced CPU / GPU / NPU platform. The video reconnaissance system is deployed on a vehicle-mounted platform. This platform includes a domestically produced CPU, a domestically produced GPU, a domestically produced NPU, a video encoding / decoding unit, a storage unit, and an SRIO video access interface. After the system starts, the hardware adaptation unit scans each hardware unit, identifies the number of CPU cores, GPU memory capacity, NPU operator subset version, and maximum concurrent inference paths, and generates the ai-compute-caps field.

[0410] Hardware capability vector h npu Example:

[0411] ,

[0412] in, This indicates the current integer inference capability of the NPU. Indicates the number of video channels that can be processed concurrently. This indicates the current NPU load. After reading this vector, the adaptive decision unit determines whether the target tracking task is suitable to be executed by the NPU, GPU, or CPU.

[0413] h npu Represents the hardware capability vector of the NPU; C int8 C represents the integer reasoning capability of the NPU. fp16 L represents the floating-point inference capability of the NPU. batch M represents the minimum latency for batch processing. stream The maximum number of concurrent inference paths is represented by B, storage bandwidth is represented by O, operator set version or operator supported encoding is represented by U, and the current NPU load is represented by U.

[0414] The system receives one forward-looking reconnaissance video stream and one lateral video stream. The video preprocessing unit decodes the video stream and preserves target edge information according to the target tracking task. If the video quality is good, the noise reduction threshold coefficient is reduced to avoid weakening the edges; if the video quality deteriorates, the noise reduction intensity is increased and the quality indicators are written into the performance evaluation unit.

[0415] When a high-priority target tracking task appears in the task instruction, the adaptive decision unit generates a state vector. If the state indicates that the NPU load is increasing, the GPU load is low, and the target tracking latency is close to the upper limit, the scheduling model calculates the environmental fluctuation coefficient. The system can also adjust the state transition probabilities. Low-priority video compression tasks can be migrated to the CPU or hardware encoding / decoding unit, while target tracking tasks can remain executed on the GPU or NPU.

[0416] After a period of execution, the performance evaluation unit collects target tracking success rate, average processing latency, GPU / NPU utilization, and video quality metrics, and calculates proximity. .like If the load falls below a threshold, the decision optimization unit determines the primary cause. If the cause is persistently high NPU load, the recommendation weight for "prioritizing NPU allocation for target tracking tasks" is reduced, and the candidate priority of GPU tracking algorithms is increased. The updated mapping relationship is written back to the adaptive decision unit, and the new recommendation relationship is used directly in the next round of task scheduling.

[0417] Example 2: Closed-loop self-optimization process of video preprocessing and adaptive algorithm switching in low-light or interference scenarios. The platform is in a low-light or interference scenario. The video preprocessing unit detects an increase in the noise intensity of the luminance component and calculates it to be larger. The system calculates a dynamic threshold based on the scene state and task real-time constraints. :

[0418] ,

[0419] λ t k represents the wavelet soft threshold at time t. t This represents the dynamic threshold coefficient, determined by scene intensity, video quality degradation, and latency constraints, σ. t This represents the estimated noise intensity at time t.

[0420] in, The system's performance is determined based on scene intensity, video quality degradation, and latency constraints. In low-light reconnaissance, if target detection accuracy decreases while real-time performance requirements are still met, the system's performance should be appropriately increased. Enhance noise reduction; if real-time performance is close to the upper limit, the system reduces the wavelet decomposition scale or selects low-complexity preprocessing.

[0421] After preprocessing, the system will Edge preservation and noise intensity estimation are incorporated into the performance evaluation unit. When the performance evaluation unit finds that video quality metrics have improved but processing latency has increased, it will weigh this change through a comprehensive evaluation. If the object detection task is high priority, the improvement in video quality may result in a higher overall score; if the task is low priority storage, the system may choose a lower complexity preprocessing.

[0422] If insufficient video quality leads to a decline in recognition results across multiple consecutive cycles, the decision optimization unit corrects the mapping database, switching the recommendation algorithm for the current scene to one more robust to low-light conditions, or adjusting the input size, confidence threshold, and preprocessing threshold range of the detection model. After the correction results are written back, the system's next scheduling round will simultaneously change the preprocessing parameters and algorithm selection.

[0423] Example 3: Formation process of global scheduling strategy in multi-task concurrent scenarios. The system simultaneously processes target reconnaissance, video fusion, situational awareness stitching, and video compression tasks. As the number of high-priority tasks in the task queue increases, the weight of the state vector increases, and the scene adaptation weight matrix increases the weight of the number of high-priority tasks and average latency. The sum of the values ​​in the reward function increases accordingly, making task completion rate rewards and low latency rewards account for a higher proportion in scheduling.

[0424] If the system reduces video compression resources to ensure high-priority target reconnaissance tasks, the performance evaluation unit will record the completion rate and resource consumption of different task types. The decision optimization unit updates the mapping database based on the evaluation results, so that high-priority tasks can obtain GPU or NPU resources in similar scenarios, while low-priority compression tasks can use hardware encoding / decoding units more often or be executed later.

[0425] This embodiment illustrates that the present invention does not only adjust the algorithm in a single task, but also forms a scheduling strategy by combining task type, scenario state, resource state and performance evaluation under multi-task concurrency conditions.

[0426] Example 4: A complete closed-loop self-optimization process for concurrent target detection and tracking tasks on a domestically produced vehicle-mounted CPU / GPU / NPU platform. This example provides a set of specific hardware models, parameter values, and execution results for each step to illustrate the feasibility of the invention and the complete execution process of the closed-loop self-optimization process in a single high-priority task scenario. The execution metrics given for each step in this example (such as tracking success rate, processing latency, resource utilization, peak signal-to-noise ratio, etc.) are exemplary values ​​used to illustrate the execution process of the closed-loop self-optimization process and the calculation method of each step, and do not constitute a limitation on the actual performance of the system.

[0427] The hardware platform uses domestically produced processors, with the following specific models and key parameters: The CPU is a Phytium FT-2000 / 4 (quad-core, 2.2GHz); the GPU is a Jingjia Micro JM7200 (4 gigabytes of video memory); the NPU is a Rockchip RK3588 (6 TOPS for integer operations per second); the video encoding / decoding unit supports H.265 hardware decoding; the storage unit uses domestically produced solid-state storage (sequential bandwidth approximately 1.5GB / s); and the video access interface is an SRIO interface. All of the above hardware components are existing mass-produced devices, connected according to their publicly available datasheets, without any changes to the hardware structure.

[0428] The NPU computing power feature field ai-compute-caps generated after scanning by the hardware adaptation unit has the following values: int8-tops=6 (indicating INT8 integer inference capability of 6 TOPS), fp16-tflops=1.5 (indicating FP16 floating-point inference capability of 1.5 trillion floating-point operations per second TFLOPS), batch-latency-ms=8 (indicating minimum batch processing latency of 8 milliseconds), max-concurrent-streams=4 (indicating maximum concurrent inference streams of 4 streams), memory-bandwidth=1.5 (indicating storage bandwidth of 1.5 GB / s), op-set=onnx-opset13 (indicating that the NPU supports the ONNX opset-13 operator set version, and the corresponding operator in the capability vector supports coded component O), and current-load=0.95 (indicating that the current NPU load is 95%). Based on this, the adaptive decision unit judges that the NPU is close to full load, and all operators required by the object detection model are within the range supported by op-set.

[0429] The system receives two channels of 1920×1080 H.265 reconnaissance video at 25 frames per second. After decoding by the video preprocessing unit, two-stage wavelet denoising is performed on the luminance component. The noise intensity σ of a given frame is estimated using median absolute deviation. t =6.0 (unit is 8-bit gray level, value range 0-255); take the basic threshold coefficient k0=0.6, scene intensity factor S t =0.5, video quality degradation degree ΔQ t =0.4, Real-time constraint strength ξ t =0.3, Adjustment coefficient η s =0.3、η q =0.2、η l =0.15, then the dynamic threshold coefficient k t=0.6×(1+0.3×0.5+0.2×0.4-0.15×0.3)=0.6×1.185≈0.711, soft threshold λ t =k t ×σ t ≈0.711×6.0≈4.27. After noise reduction, the peak signal-to-noise ratio (PSNR) of this frame increased from 34 dB to 36 dB.

[0430] The adaptive decision unit receives high-priority target tracking tasks, with a time delay constraint L. bound =40 milliseconds, with an accuracy constraint of a tracking success rate of no less than 0.85. Candidate strategies include: Strategy 1, prioritizing target tracking tasks to the NPU; Strategy 2, assigning target tracking tasks to the GPU and video compression tasks to the hardware encoding / decoding unit; Strategy 3, having the target tracking task executed collaboratively by the CPU and GPU. After performance evaluation and strategy selection, the selected strategy is issued and written to the strategy cache with the following execution parameters: target detection confidence threshold θ conf =0.45, Detect input size R in =640×640, Tracking algorithm search window W search =64×64.

[0431] The adaptive decision unit models the system state as a 12-dimensional state vector, where CPU utilization u t cpu =0.40, GPU utilization u t gpu =0.35, NPU utilization rate u t npu =0.95, average delay l̄ t =45 milliseconds (the other dimensions take the default value when the device is idle in this example). Since the NPU utilization is close to 1 and the latency is close to the constraint limit, the environmental fluctuation coefficient Δ t The calculation result is relatively large (taking the normalization coefficient ψ=1 and the maximum correction coefficient Δ). max When Δ = 0.3, t (≈0.27), based on this, the system dynamically adjusts the MDP state transition probability to increase the state transition probability from strategy one to strategy two.

[0432] The performance evaluation unit collected four indicators for the three strategies to form the original indicator matrix: tracking success rate (positive indicator) was 0.82, 0.91, and 0.85 respectively; average latency (negative indicator, in milliseconds) was 45, 32, and 40 respectively; NPU utilization (interval indicator, target center 0.70, maximum allowable deviation 0.30) was 0.95, 0.70, and 0.55 respectively; peak signal-to-noise ratio (positive indicator, in decibels) was 34, 36, and 33 respectively (Peak signal-to-noise ratio PSNR here represents the output image quality of the video reconnaissance system: different scheduling strategies have different occupancy and allocation of the three types of computing power (CPU, GPU, and NPU), and the computing power margin that the video preprocessing unit can use to continuously perform complete second-level wavelet denoising is different accordingly. Therefore, the output image PSNR that the system can stably maintain under each strategy is different. Since the video reconnaissance system not only needs to complete the tracking task, but also needs to ensure the quality of the reconnaissance image, PSNR and tracking indicators are included in the multi-attribute comprehensive evaluation. The values ​​here are illustrative values). Calculate using the following steps (assuming 0·ln0=0): (a) Directional normalization: positive index r=(x-min) / (max-min), negative index r=(max-x) / (max-min), interval index r=1-|xc| / d (c=0.70, d=0.30). The normalized matrices are obtained (in the order of the above indicators): Strategy 1 (0.000, 0.000, 0.167, 0.333), Strategy 2 (1.000, 1.000, 1.000, 1.000), Strategy 3 (0.333, 0.385, 0.500, 0.000), where the NPU utilization column is obtained by interval formula: Strategy 1 1 - |0.95 - 0.70| / 0.30 ≈ 0.167, Strategy 2 1 - 0 / 0.30 = 1.000, Strategy 3 1 - |0.55 - 0.70| / 0.30 = 0.500. (b) Entropy weight: from p ij =r ij / Σ i r ij E j =-(1 / ln3)·Σ i p ij ·ln p ij The entropy values ​​E of the four indicators are obtained. j The coefficients of difference g were 0.512, 0.538, 0.817, and 0.512, respectively. j =1-E j The values ​​are 0.488, 0.462, 0.183, and 0.488 respectively. After normalization, the entropy weight ω is obtained. jThe values ​​are 0.301, 0.285, 0.113, and 0.301 (summing up to 1). (c) TOPSIS: After weighting the normalized matrix with entropy weights, the maximum value in each column constitutes the positive ideal solution, and the minimum value constitutes the negative ideal solution. Calculate the Euclidean distance d from each strategy to the positive and negative ideal solutions. + d - Strategy 1 + ≈0.470, d - ≈0.100, Strategy 2 d + ≈0.000, d - ≈0.521, Strategy 3d + ≈0.406, d - ≈0.153; by C i =d - / (d + +d - The relative similarity of the three strategies is C. i The values ​​are 0.176, 1.000, and 0.274, respectively.

[0433] Since Strategy 2 has the highest relative closeness C2=1.000 and is higher than the effectiveness score threshold C min (Take 0.6), the system selects strategy two. The decision optimization unit lowers the recommended weight of "prioritizing the allocation of target tracking tasks to the NPU" and increases the recommended priority of "allocating target tracking tasks to the GPU"; the updated mapping relationship is written to the strategy cache through the strategy write-back channel, and the adaptive decision unit directly adopts it in the next round of scheduling.

[0434] After executing this closed-loop procedure, the target tracking success rate improved from 0.82 to 0.91, the average processing latency decreased from 45 milliseconds to 32 milliseconds, the NPU utilization rate decreased from 0.95 to 0.70, and the peak signal-to-noise ratio improved from 34 dB to 36 dB. In contrast, with existing static configuration schemes, the fixed configuration file allocates target tracking tasks to the NPU, leading to tracking latency exceeding constraints (45 milliseconds > 40 milliseconds) and insufficient success rate (0.82 < 0.85) when the NPU is near full load, and these issues cannot be automatically adjusted. This embodiment demonstrates that the present invention solves the technical problem of static configuration easily allocating high-precision tasks to overloaded hardware units when hardware load fluctuates, resulting in task timeouts and decreased accuracy, by writing back the performance evaluation results to correct the scheduling strategy. Compared with existing technologies, it improves processing latency, task accuracy, and resource balancing.

[0435] Example 5: A complete closed-loop self-optimization process of concurrent target tracking and situation stitching tasks on a domestically produced vehicle platform, illustrating the global superiority of the present invention over a single-task locally optimal scheduling strategy. This example uses the same domestically produced hardware platform as Example 4 (CPU Phytium FT-2000 / 4, GPU Jingjia Micro JM7200, NPU Rockchip RK3588 with built-in NPU, H.265 hardware decoding unit, and domestic solid-state storage). The difference lies in that the system needs to concurrently execute high-priority target tracking and situation stitching tasks, with the two competing for GPU and NPU resources. Similar to Example 4, the operational metrics given in each step of this example are exemplary values ​​used to illustrate the formation process and calculation method of the global equilibrium strategy under multi-task concurrency, and do not constitute a limitation on the actual performance of the system.

[0436] Step 1: Tasks and Constraints. The system receives two channels of 1920×1080 H.265 reconnaissance video at 25 frames per second. The hard constraints for the high-priority target tracking task are a tracking success rate of no less than 0.80 and a tracking latency of no more than 50 milliseconds; the concurrent situation stitching task requires a frame rate of no less than 10 frames per second. The adaptive decision unit generates three candidate strategies: Strategy 1 (single-task optimal) allocates the target tracking task exclusively to the NPU and GPU, while the situation stitching task only receives the remaining resources; Strategy 2 (global balance) allocates the target tracking task to the NPU and reserves GPU quota for the situation stitching task; Strategy 3 (cooperative) executes the target tracking task collaboratively by the CPU and GPU, while the NPU yields to other tasks. All three strategies satisfy the aforementioned real-time and accuracy hard constraints, therefore the final selection is determined by the multi-attribute decision-making of the performance evaluation unit.

[0437] Step two: State modeling and dynamic correction of the Markov decision process. The adaptive decision unit models the system state as a 12-dimensional state vector, where the NPU utilization rate approaches 1 and the situation stitching task queue experiences backlog. Due to the environmental fluctuation coefficient Δ t Larger (taking normalization coefficient ψ=1, maximum correction coefficient Δ) max When Δ = 0.3, t (≈0.28) Based on this, the decision optimization unit lowers the state transition probability from the current state to strategy one (single task optimal) in order to prevent the system from falling into the local optimal trap that makes the tracking task locally optimal but starves the resources of the situation splicing task; the three candidate strategies are then submitted to the performance evaluation unit for decision.

[0438] Step 3: Construct the original indicator matrix. The performance evaluation unit collected four indicators for the three strategies: tracking success rate (positive indicator) was 0.93, 0.88, and 0.80 respectively; tracking latency (negative indicator, in milliseconds) was 28, 35, and 46 respectively; situation stitching frame rate (positive indicator, in frames per second) was 12, 23, and 25 respectively; and NPU utilization (interval indicator, target center 0.70, maximum allowable deviation 0.30) was 0.98, 0.78, and 0.72 respectively. The three strategies each have their advantages and disadvantages, and no single strategy dominates: Strategy 1 is optimal in tracking success rate and latency, but has the lowest situation stitching frame rate and the highest NPU utilization (most unbalanced resources); Strategy 3 is the opposite.

[0439] Step 4, Direction Normalization. For positive indicators, use r=(x-min) / (max-min); for negative indicators, use r=(max-x) / (max-min); for interval indicators, use r=1-|xc| / d (where c is the target center 0.70 and d is the maximum allowable deviation 0.30). The resulting normalized matrices (in the order of the indicators above) are: Strategy 1 (1.000, 1.000, 0.000, 0.067), Strategy 2 (0.615, 0.611, 0.846, 0.733), Strategy 3 (0.000, 0.000, 1.000, 0.933).

[0440] Step 5, Entropy weight calculation. Press p ij =r ij / Σ i r ij Calculate the weight of each indicator (by convention 0·ln0=0), using the entropy value E. j =-(1 / ln3)·Σ i p ij ·ln p ij The entropy values ​​of the four indicators were 0.605, 0.604, 0.628, and 0.749, respectively, and the coefficient of difference g was... j =1-E j The values ​​are 0.395, 0.396, 0.372, and 0.251 respectively. After normalization, the entropy weight ω is obtained. j The values ​​are 0.279, 0.280, 0.263, and 0.178 (which sum to 1).

[0441] Step 6: TOPSIS proximity calculation. After weighting the normalized matrix with entropy weights, the maximum value in each column constitutes the positive ideal solution, and the minimum value constitutes the negative ideal solution. Calculate the Euclidean distance d from each strategy to the positive and negative ideal solutions. + d - Strategy 1 + ≈0.305, d - ≈0.395, Strategy 2 d +≈0.162, d - ≈0.350, Strategy 3d + ≈0.395, d - ≈0.305; by C i =d - / (d + +d - The relative similarity of the three strategies is C. i The values ​​are 0.565, 0.683, and 0.435, respectively.

[0442] Step 7, Strategy Selection. Strategy 2 has the highest relative closeness (C² = 0.683) and is above the effectiveness score threshold C. min (Taking 0.6), the system selects Strategy 2. It's worth noting that Strategy 2 is not the best in any of the four metrics: its tracking success rate and tracking latency are lower than Strategy 1, and its situation stitching frame rate and NPU utilization are lower than Strategy 3, only ranking first in overall proximity. If the rule of selecting the highest tracking success rate or the lowest tracking latency is used, Strategy 1 will be selected; if the rule of selecting the highest situation stitching frame rate is used, Strategy 3 will be selected. None of these naive rules can select the compromise Strategy 2 that balances various tasks and global resource allocation.

[0443] Step 8: Policy Write-back and Next Round Scheduling. Based on this, the decision optimization unit lowers the recommended weight for the target tracking task to exclusively utilize the NPU and GPU, and increases the recommended priority for reserving GPU quotas for the situation stitching task; the updated mapping relationship is written to the policy cache via the policy write-back channel, and the adaptive decision unit directly uses it in the next round of scheduling.

[0444] Step 9: Compare with existing static configuration schemes. When using existing static configuration schemes, the configuration file always allocates high-priority tracking tasks optimally per task (i.e., Strategy 1). The result is a tracking success rate of 0.93 and a tracking latency of 28 milliseconds. Although the tracking task's own metrics are excellent, the situation stitching frame rate drops to 12 frames per second, and the NPU utilization rate is as high as 0.98. Concurrent situation stitching tasks experience significant frame drops, severe system resource imbalance, and cannot be automatically adjusted. After adopting this invention, the system selects Strategy 2: a tracking success rate of 0.88 (meeting the constraint of not less than 0.80), a tracking latency of 35 milliseconds, a situation stitching frame rate rebounding to 23 frames per second, and an NPU utilization rate falling back to 0.78 (close to the target center of 0.70). This embodiment demonstrates that the present invention, through the synergy of entropy weight TOPSIS performance evaluation and dynamic correction of Markov decision process state transition, can select a compromise strategy that balances the resources of each task and the global resource balance when there is resource competition among tasks. It solves the technical problem that static configuration and any naive scheduling rule based on a single index can easily lead to local optima for a single task, but cause resource starvation and global imbalance for concurrent tasks. Compared with the prior art, it improves the global resource balance and overall availability of multiple tasks.

[0445] This invention provides a closed-loop self-optimization design system and method for video reconnaissance systems based on multiple domestically produced hardware platforms. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A closed-loop self-optimizing design system for a video reconnaissance system based on multiple domestically produced hardware platforms, characterized in that, include: The hardware adaptation unit is used to scan domestic hardware platforms, match drivers, register unified hardware interfaces, and generate the computing power feature description field ai-compute-caps. The video preprocessing unit is used to access different video sources, complete protocol decryption, decapsulation, decoding, format conversion and scene adaptive preprocessing, and output video quality indicators; The adaptive decision unit is used to parse task instructions, read computing power characteristics, task requirements and video quality indicators, establish state space, action space and reward function, and generate scheduling strategies based on dynamic probability correction of Markov Decision Process (MDP). The performance evaluation unit is used to evaluate the acquisition and processing performance, task accuracy, video quality, and resource utilization. It calculates the system performance results using the TOPSIS (Topology-Approximation Ideal Solution Ranking Method) or an equivalent multi-index evaluation method. The decision optimization unit is used to correct the mapping relationship library, algorithm parameter range and state transition probability matrix based on the performance evaluation results, and write the corrected strategy back to the adaptive decision unit.

2. A closed-loop self-optimization design method for a video reconnaissance system based on multiple domestically produced hardware platforms for the system, characterized in that, The steps include the following: Step S1: After the system starts up, the hardware adaptation unit performs a hardware scan on the current platform to identify the central processing unit (CPU), graphics processing unit (GPU), neural network processor (NPU), video encoding / decoding unit, image processing unit, storage unit, and video access interface, and matches and loads the corresponding drivers. Step S2: After the driver is successfully loaded, the hardware adaptation unit registers the unified hardware interface and generates the computing power feature description field ai-compute-caps. The computing power feature description field ai-compute-caps is written to the computing power feature cache or the unified hardware interface directory for the adaptive decision unit to read. Step S3: The video preprocessing unit receives video data from the serial high-speed input / output interface SRIO, real-time streaming protocol RTSP, real-time message transmission protocol RTMP, Linux video capture interface version 2 V4L2, user datagram protocol UDP, or local video interface, and performs protocol de-protocol, de-encapsulation, decoding, format conversion, and scene adaptive preprocessing on the video stream. Step S4: The adaptive decision unit receives the task instruction, parses the task type, priority, latency requirements, accuracy requirements and output format, reads the computing power characteristics, video quality indicators and current hardware load, and generates candidate algorithms and parameter combinations. Step S5: The adaptive decision unit models the current state as a state vector and generates scheduling actions based on the dynamic probability correction MDP. The scheduling actions include resource allocation, task migration, algorithm switching, parameter adjustment or priority adjustment. Step S6: The performance evaluation unit collects and processes latency, frame rate, accuracy, tracking success rate, video quality, CPU utilization, GPU utilization, NPU utilization, video memory usage, memory usage, and storage bandwidth usage, and calculates the system performance evaluation results. Step S7: When the performance evaluation result is lower than the preset condition, or when a downward trend occurs for two or more consecutive periods, the decision optimization unit corrects the mapping relationship library, the algorithm parameter range and the state transition probability matrix, and writes the correction result back to the adaptive decision unit. In step S8, the adaptive decision-making unit uses the updated strategy in the next round of task scheduling to form a closed loop of hardware recognition, video preprocessing, task scheduling, performance evaluation, strategy correction, and rescheduling.

3. The method according to claim 2, characterized in that, In step S1, the hardware capabilities are represented as a vector: , in, Indicates the first Capability vector of each hardware unit; This indicates 8-bit integer (INT8) reasoning capability; This indicates 16-bit floating-point FP16 computing power; Indicates the minimum latency for batch processing; Indicates the maximum number of concurrent inference paths; Indicates storage bandwidth; Indicates the operator set version or the operator's supported encoding; This indicates the current load.

4. The method according to claim 3, characterized in that, In step S1, during driver matching, the system first performs an exact match: assuming the hardware identifier obtained from the scan is... The first in the device tree The compatibility identifier for each node is Then the exact matching function Defined as: , If no exact matching node is found, a general matching is performed. The general matching calculates a matching score based on the vendor, architecture, bus type, and functional category. : , The system selects the highest-scoring system that exceeds the threshold. The node serves as a general-purpose driver node: , Among them I vendor Indicates whether the manufacturer is compatible, I arch Indicates whether the architecture matches, I bus Indicates whether the bus type matches, I func Indicates whether the functional category matches; ω1 represents the weight coefficient of the vendor matching item, ω2 represents the weight coefficient of the architecture matching item, ω3 represents the weight coefficient of the bus matching item, and ω4 represents the weight coefficient of the functional category matching item; S min S represents the scoring threshold for general matching. n This represents the general matching score of the nth node, where n is the nth node. * This indicates the node number with the highest score; The system creates a hash index using the compatible field: , Where compatible represents the hardware compatibility identifier field, Hash represents the hash function, and Index represents the hash index value calculated from the compatibility identifier using the hash function; Once the hardware scan obtains a compatibility identifier, the system first locates candidate nodes using a hash index, and then performs an exact match or a general match.

5. The method according to claim 4, characterized in that, In step S3, let the input video frame be... , Frame attributes for: , in, and These represent the width and height of the resolution, respectively. Indicates frame format, Indicates frame rate, Represents a timestamp. Indicates the video source type; The video preprocessing unit converts the input frames into the luminance / chrominance color format YUV, the red-green-blue color format RGB, or the blue-green-red color format BGR. For video frames requiring noise reduction, the system employs wavelet transform for noise reduction; for YUV420 video frames, the system extracts the luminance component. Perform second-level wavelet decomposition: , Among them, Y t 1 represents the luminance component of the t-th frame of video; LL2 represents the low-frequency approximation coefficients of the second-order wavelet decomposition; LH2 represents the second-order horizontal high-frequency coefficients, HL2 represents the second-order vertical high-frequency coefficients, HH2 represents the second-order diagonal high-frequency coefficients; LH1 represents the first-order horizontal high-frequency coefficients, HL1 represents the first-order vertical high-frequency coefficients, and HH1 represents the first-order diagonal high-frequency coefficients. The noise intensity is estimated using the following formula: , in This represents the noise intensity estimate at time t, median indicates median operation, and HH1 represents the first-order diagonal high-frequency coefficient. soft threshold Defined as: , Among them, coefficient Defined as: , in, The basic threshold coefficient; Indicates the scene intensity factor; Indicates the current degree of video quality degradation; Indicates the strength of the task's real-time constraint; , , This is the adjustment coefficient; For any high-frequency coefficient Soft thresholding function for: , Where sign represents the sign function; After processing the high-frequency coefficients, the system performs wavelet reconstruction to obtain the denoised luminance components. Then, with color components , Synthesized output frames : , in U represents the luminance component after noise reduction and reconstruction. t V represents the blue chromaticity component of the t-th frame. t This represents the red chroma component of the t-th frame. Indicates the luminance component after noise reduction With chromaticity component U t V t The output video frames are recombined using the synthesis function Combine; Calculate the peak signal-to-noise ratio (PSNR): , , Among them, MSE t W represents the mean square error between the original brightness and the brightness after noise reduction in frame t. t H represents the width of the frame. t Y represents the height of the frame. t (x,y) represents the pixel value of the original luminance component at coordinates (x,y). This represents the pixel value of the luminance component at coordinates (x, y) after noise reduction, where x represents the horizontal coordinate of the pixel and y represents the vertical coordinate of the pixel; PSNR t MAX represents the peak signal-to-noise ratio of frame t. Y This indicates the maximum value of the luminance component.

6. The method according to claim 5, characterized in that, In step S4, the system establishes a mapping relationship database for hardware characteristics, task types, algorithms, and parameter ranges. The m-th record r in the mapping relationship database... m Represented as: , in, Indicates hardware conditions. Indicates the task type. This refers to a recommendation algorithm. Indicates the parameter range. This represents historical performance statistics; When selecting an algorithm, the system first filters candidate records based on task type, and then calculates the fit score based on hardware capability vectors and task constraints. , Here, Score represents the overall suitability of the candidate algorithm and hardware combination for the current task; a higher value indicates a higher degree of suitability. compute Indicates the degree of matching in computing power, Fit latency Indicates the degree of delay satisfaction, Fit accuracy Indicates the degree of precision achieved, Fit resource μ1 represents the weight of the degree of resource utilization suitability; μ2 represents the weight of the degree of computing power matching; μ3 represents the weight of the degree of latency satisfaction; and μ4 represents the weight of the degree of accuracy satisfaction. The confidence threshold of the object detection algorithm is expressed as: The input size is represented as The tracking algorithm search window is represented as The parameter selection target is: , Where Ω represents the candidate set of parameters, J(Θ) represents the comprehensive performance function under the current task and hardware conditions, and Θ * This represents the optimal combination of parameters that maximizes the overall performance function.

7. The method according to claim 6, characterized in that, In step S5, the adaptive decision-making unit models the task scheduling process as a Markov decision process (MDP): , in, For state space, For the action space, Let be the state transition probability. For the reward function, Discount factor; The state vector can be defined as a 12-dimensional continuous vector s t : , in, Indicates CPU utilization. Indicates GPU load. Indicates NPU utilization. This indicates the amount of remaining GPU memory. Indicates the amount of memory remaining. Indicates storage bandwidth. , , These represent the number of high, medium, and low priority tasks, respectively. This indicates the average latency of the current task. Indicates the length of the task queue. This indicates the timeout rate for unfinished tasks; The system introduces a scene adaptation coefficient matrix. Let the first The basic weights for each state dimension are: The current scenario's influencing factor is , No. The sensitivity of each dimension in the current scenario is: The dynamic weights are then: , Among them, w i 0 φ represents the basic weight of the i-th state dimension. s κ represents the influencing factor of the current scenario. i,s w represents the sensitivity of the i-th state dimension in the current scenario. i t This represents the dynamic weight of the i-th state dimension after adjustment in the current scenario at time t; The state fusion vector is defined as: , in, Represents the normalization function; in, This represents the state fusion value of the i-th state dimension at time t, which is equal to the dynamic weight w. i t With normalized state components Norm(s) i,t The product of ); s i,t This represents the original value of the i-th state dimension at time t, and Norm(·) represents the normalization function; Action space is defined as a set of discrete actions, which are divided into resource allocation actions, task migration actions, algorithm adjustment actions, and parameter adjustment actions. Resource allocation actions include adjusting the resource ratios of CPU, GPU, and NPU: , Where A resource Represents a set of resource allocation actions, CPU 20 This indicates that the CPU resource usage ratio will be adjusted to 20%, and the GPU... 20 This indicates that the GPU resource usage ratio will be adjusted to 20%, NPU 20 This indicates that the NPU resource usage ratio will be adjusted to 20%. Task migration operations include migrating high-priority or medium-priority tasks from one hardware unit to another; Algorithm adjustments include switching object detection models, object tracking algorithms, image segmentation algorithms, or video compression strategies; Parameter adjustment actions include adjusting the confidence threshold, input resolution, tracking search window, noise reduction threshold, inter-frame sampling interval, and batch size; Introducing a two-factor penalty coefficient based on scenario and task: , in, Indicates time The dynamic penalty coefficient; Basic penalty coefficient; As a scene penalty factor; Based on the urgency of the situation; As a task penalty factor; This represents the percentage of current task types or the percentage of critical tasks. Reward function R t Defined as: , in, Rewards are given based on task completion rate. Incentives for resource utilization As a reward for low latency, As a reward for video quality, This is a timeout penalty item. , , , Dynamic weights; Dynamic weights are determined by the task queue and resource status: , , , , in, This indicates the percentage of high-priority tasks. Indicates the percentage of real-time tasks. and These represent GPU resource utilization and NPU resource utilization, respectively. Indicates storage bandwidth usage. Indicates the average delay. Indicates task latency constraints. Indicates the current video quality. This indicates the video quality requirement; f1 represents the percentage of high-priority tasks ρ. high and the proportion of real-time tasks ρ real Calculate the task completion rate reward weight α t The mapping function, f2, represents the mapping function based on GPU utilization U. gpu NPU utilization rate U npu and storage bandwidth usage B store Calculate the reward weight β for resource utilization t The mapping function, f3, represents the average time delay. and task delay constraint L bound Calculate the low-latency reward weight γ t The mapping function, f4, represents the mapping function based on the current video quality Q. t And video quality requirements Q bound Calculate the video quality reward weight ζ t The mapping function; Task completion rate reward R comp Represented as: , in, Number of task types As a weight for task type, For the first Number of tasks completed by class For the first Total number of tasks of each type; Low latency reward R latency Represented as: , Among them, l t L represents the current task delay. bound Indicates task latency constraints; Video quality bonus R quality Represented as: , Where R quality Indicates a video quality reward, Q t Q indicates the current video quality. bound Indicates video quality requirements, Q min Indicates the minimum acceptable video quality; Resource utilization rate reward R util Represented as: , in, For the number of hardware units, For the first Utilization of each hardware unit Use the center of the interval to target the objective; The system constructs the initial state transition probability based on historical data. : , in, Represents the state in historical data Execute action After transitioning to state Number of times, Indicates the state in historical data Execute action Total number of times; Introducing environmental fluctuation coefficient : , in, Represents the environment vector at time t; These are the normalization coefficients; x is the maximum correction factor; t-1 Let ||| represent the environment vector at the previous moment; |||2 represents the L2 norm of the vector; Corrected state transition probabilities for: , in, This represents the compensation probability distribution under the current environment; For the sliding window length Recent probability Represented as: , in, Indicates recent Statistical frequency within a period For smoothing coefficients, Let the number of states be defined as follows: 。 8. The method according to claim 7, characterized in that, In step S6, let there be a total of Each evaluation object, There are 1 evaluation index, and the original index matrix X is: , in, Indicates the first The evaluation object is in the first The original values ​​of each indicator; m represents the number of evaluation objects, and n represents the number of evaluation indicators; For positive indicators, the standardization is as follows: , Among them, z ij This represents the standardized value of the i-th evaluation object on the j-th indicator; For negative indicators, the standardization is as follows: , For interval-type indicators, the definition is: , Where, x j target D represents the center of the target interval for the j-th indicator. j This represents the maximum permissible deviation of the j-th indicator; When calculating entropy weights, first calculate the first... The first indicator The proportion of each evaluation object : , No. Information entropy of each indicator e j for: , Coefficient of difference g j for: , Indicator weights for: , After obtaining the weights, construct the weighted normalization matrix: , Where v ij This represents the weighted standardized value of the i-th evaluation object on the j-th indicator, which is equal to the weight ω of the j-th indicator. j With standardized value z ij The product; Positive Ideal Solution and negative ideal solution They are respectively: , No. Distance from each evaluation object to the ideal solution Distance to the negative ideal solution for: , , Relative closeness for: , The system will As the effectiveness score of the current strategy, if Below the threshold , or continuous If the evaluation period decreases, optimization is triggered: , Where C i C represents the relative proximity of the i-th evaluated object and serves as the effectiveness score of the current strategy; min Indicates the performance score threshold; r represents the number of evaluation periods with consecutive declines; Trigger t This represents the optimization trigger flag at time t. A value of 1 triggers optimization, while a value of 0 does not trigger optimization. The trigger result enters the decision optimization unit.

9. The method according to claim 8, characterized in that, In step S7, the decision optimization unit determines the problem type based on the performance evaluation results, and the mapping relationship library is corrected as follows: , in, Indicates the first The mapping relationship at time Historical performance estimates Historical retention factor; If the mapping relationship is continuously below the threshold : , The system lowers the recommendation priority of the mapping relationship or narrows the parameter range; if another candidate mapping relationship is more efficient in similar tasks, the recommendation priority is increased. The parameter range correction is performed as follows: Let the algorithm parameters be... The current recommendation range is The current optimal parameters are If the current strategy scores high, then focus on... Narrowing the search scope: , , Where, θ min θ represents the lower limit of the recommended range for the parameter. max Δ represents the upper limit of the recommended range for the parameter. θ This indicates a narrowing step size for the parameter range. This indicates the lower limit of the corrected parameter range. Indicates the upper limit of the corrected parameter range; For long-term strategy prediction, a regression model is established based on historical performance data, with the feature vector z denoted as z. t for: , in, Let be the hardware capability vector at time t. Let be the state vector at time t. Let be the algorithm parameters at time t. Encoding the task at time t. The current video quality metric is represented by the actual performance score of the t-th sample. Establish a linear regression model with the variable as the dependent variable: , Where z t This represents the eigenvector at time t. a represents the performance score predicted by the regression model. reg Let λ represent the regression coefficient vector, b represent the bias term of the regression model, and λ represent the regression coefficient vector. r Represents the regularization coefficient; The model training objective is: , in, N is the regularization coefficient. s This indicates the number of historical samples used to train the regression model; the superscript T in the regression model indicates the transpose of the vector; During policy write-back, the decision optimization unit writes the updated mapping relationship, parameter range, and probability matrix into the policy cache and notifies the adaptive decision unit via the policy push service. The adaptive decision unit reads the latest policy in the next round of task scheduling. The system sets a minimum update interval or a policy stability window. , in, This is the time since the last strategy update. For the minimum update interval, Update t This represents the policy update flag at time t; a value of 1 indicates that a policy update is performed, and a value of 0 indicates that no update is performed. last τ represents the time of the last policy update. min This indicates the minimum update interval.

10. The method according to claim 9, characterized in that, In step S8, the closed-loop feedback relationship is defined by data type, and the hardware adaptation unit outputs computing power parameters to the adaptive decision unit; the video preprocessing unit outputs video quality indicators to the adaptive decision unit and the performance evaluation unit. The adaptive decision-making unit outputs scheduling actions, hardware load status, and task execution results to the performance evaluation unit; the performance evaluation unit outputs performance scores and optimization trigger signals to the decision optimization unit; the decision optimization unit writes back the mapping relationship correction results, parameter range correction results, and probability matrix update results to the adaptive decision-making unit. This data stream is represented as: , , , , , The symbol → indicates the direction of data flow; Caps t Representing computing power characteristics, Decision t This represents the scheduling decision generated by the adaptive decision unit, Quality. t Indicates video quality metrics, Action t Indicates a scheduling action, Load t Indicates load status, Result t Evaluation indicates the result of task execution. t Indicates the performance evaluation results, Optimization t Represents the policy optimization process. t+1 This indicates the scheduling strategy for the next round.