Adaptive neural network configuration method for video stream target detection
Through the adaptive neural network configuration method for video stream object detection, dynamic evaluation and real-time decision-making, the problems of fixed network configuration, insufficient resource perception and unstable configuration switching in the prior art are solved, and the adaptive configuration of the neural network and the efficient and stable operation of the system are realized.
Patent Information
- Application Number
- CN202510053261.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, the video stream object detection network is configured in a fixed manner, unable to adapt to dynamically changing scenario requirements and hardware resource conditions, lack resource perception capabilities, limited configuration switching mechanism, and simple optimization strategy.
An adaptive neural network configuration method for video stream object detection is proposed. Through dynamic evaluation, real-time decision-making and smooth switching, a collection of object detection neural networks is constructed and screened to realize the adaptive configuration of neural networks.
It realizes dynamic adaptation of neural network configuration, provides complete resource perception capabilities, establishes a smooth configuration switching mechanism, realizes multi-objective optimization based on deep reinforcement learning, and improves the adaptability and stability of the video stream object detection system.
Smart Images

Figure CN120107843A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of video stream target detection, and in particular relates to an adaptive neural network configuration method for video stream target detection. Background Art
[0002] With the rapid development of smart cities, smart security, and unmanned driving, the demand for video stream processing in edge computing environments is growing. Especially in high-computation tasks such as target detection, how to configure and manage neural networks has become a key challenge; currently, there are the following main problems in neural network configuration:
[0003] Fixed network configuration scheme: Existing neural network configuration methods usually adopt static deployment, that is, a specific network structure and parameter configuration are pre-selected for deployment; this fixed configuration method cannot adapt to the dynamically changing scene requirements and hardware resource conditions in actual applications, resulting in the inability to adjust the network structure according to the complexity of the video scene; it cannot respond to changes in system load to allocate resources, and it is difficult to balance the performance requirements in different scenarios;
[0004] Insufficient resource perception: Traditional neural network configuration methods lack the ability to perceive hardware resources and cannot accurately evaluate the resource overhead of different network configurations, making it difficult to achieve efficient configuration under the limited resources of edge devices.
[0005] Limited configuration switching mechanism: When the network configuration needs to be adjusted, the existing methods often require service interruption for switching. There is a lack of a smooth transition configuration switching mechanism, making it difficult to ensure service continuity and stability.
[0006] Simple configuration optimization strategy: Current configuration methods mostly use simple rule-based optimization, lack a systematic performance evaluation and feedback mechanism, and are unable to make optimization decisions with multi-objective trade-offs.
[0007] In view of the above problems, the present invention proposes an adaptive neural network configuration method for video stream target detection, which realizes the adaptive configuration of the neural network through technical means such as dynamic evaluation, real-time decision-making and smooth switching. Summary of the invention
[0008] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide an adaptive neural network configuration method for video stream target detection, so as to solve the problems in the prior art of video stream target detection such as fixed network configuration, insufficient resource perception, limited switching mechanism and simple optimization strategy.
[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0010] The present invention provides an adaptive neural network configuration method for video stream target detection, comprising the following steps:
[0011] 1) Build and select a set of target detection neural networks;
[0012] 2) Using the neural network in the selected target detection neural network set to perform video stream target detection; judging whether the current neural network configuration needs to be updated in a certain decision cycle, if so, returning to step 2); otherwise, entering step 3);
[0013] 3) Update the neural network configuration.
[0014] Furthermore, the step 1) specifically includes:
[0015] 11) Construct a set of target detection neural networks;
[0016] 12) Filter the target detection neural network collection.
[0017] Furthermore, the target detection neural network set in step 11) is a weight-sharing neural network set based on a hypernetwork, a discrete weight-independent neural network set, or a combination of the two; the weight-sharing neural network set based on a hypernetwork is formed by constructing a large-scale weight-sharing hypernetwork and extracting subnetworks of different configurations therefrom; the discrete weight-independent neural network set is composed of multiple independently trained target detection networks; specifically includes:
[0018] 111) Construct a set of weight-sharing neural networks based on a hypernetwork as follows:
[0019] 1111) Construct a detection backbone network with configurable parameters, wherein the configurable parameters include: the number of backbone network depth layers D, the ratio of the number of channels in each layer W, the convolution kernel size K, and the input resolution R; wherein D = {d 1 ,d 2 ,…d m},d i ∈N+,W={w 1 ,w 2 ,…w n},w i ∈R+,K={k 1 ,k 2 ,…k p},k i ∈N+,R={r},r∈N+,d i represents the depth of the i-th block, w i represents the channel number multiplier of the i-th layer, k i represents the convolution kernel size of the i-th layer, and r represents the input resolution;
[0020] 1112) construct a feature pyramid network including configurable parameters, wherein the configurable parameters include: the number of feature pyramid levels L and the multiple of the number of channels of each feature pyramid layer W′, where L={l 1 ,l 2 ,…l r},l i ∈{0,1},W′={m 1 ,m 2 ,…m s},m i ∈R+,l i Indicates whether to select the feature of the i-th layer (0 for not selecting, 1 for selecting), m i Indicates the channel number multiplier of the i-th layer;
[0021] 1113) After constructing the feature pyramid network, a detection head network is constructed, and a two-stage detection head structure based on region proposal (such as Faster R-CNN) or a single-stage detection head structure based on anchor-free (such as FCOS) is selected to achieve the final target detection and generate target category and location information;
[0022] 1114) combining the networks in step 1111), step 1112) and step 1113) to obtain a weight-sharing neural network based on a hypernetwork;
[0023] For any configuration parameter combination c, c = (d, w, k, r, l, m), d ∈ D, w ∈ W, k ∈ K, r ∈ R, l ∈ L, m ∈ W ′, the corresponding sub-network can be obtained from the super-network. The resulting neural network set is expressed as:
[0024] M 1 ={s(c)|c∈C}
[0025] Where C is the Cartesian product of the parameter space, C = D × W × K × R × L × W ′, s(c) represents the subnetwork with parameter configuration c;
[0026] 1115) train all sub-networks in step 1114), update the shared weights by randomly sampling sub-networks during training, and obtain the shared weights of the super-network through training; specifically, for the currently sampled sub-network s(c), update the shared weights by gradient descent, as follows:
[0027]
[0028] Among them, θ s(c) is the shared weight of the current sampling sub-network s(c), η is the learning rate, For the subnetwork in the dataset The loss function on is θ s(c) The gradient of
[0029] Get a set of optional network sets M 1 , and the shared weights of the sub-networks in the network set share , complete the construction of a set of weight-sharing neural networks based on a hypernetwork;
[0030] Optionally, a multi-scale feature distillation approach can be used. After randomly sampling the sub-networks, the output of the largest network in the search space or the output of a pre-trained teacher network is used to guide the training of the currently sampled sub-network.
[0031] 112) Construct a discrete set of weight-independent neural networks as follows:
[0032] 1121) Build a set of target detection networks with configurable parameters. The network types include: networks based on convolutional neural networks (CNN) (such as YOLOv5 series, EfficientDet series) and networks based on self-attention mechanisms (such as DETR series);
[0033] Different network variants can be designed for each network. The optional configuration parameters of the network variants include: input resolution R, number of backbone network depth layers D and the ratio of the number of channels in each layer W, R = {r}, r∈N+, D = {d 1 ,d 2 ,…d m},d i ∈N+,W={w 1 ,w 2 ,…w n},w i ∈R+;
[0034] By combining the above configuration parameters, a variety of network variants are generated, and the resulting neural network set is expressed as:
[0035] M 2 ={m(c)|c∈C}
[0036] Where C is the Cartesian product of the parameter space, C = R × D × W, m(c) represents a network variant instance with parameter configuration c;
[0037] 1122) Each network variant is trained separately to obtain independent weights of each network; specifically, for each network variant, the independent weights are updated by gradient descent as follows:
[0038]
[0039] Among them, θ m(c) is the network variant instance m(c), η is the learning rate, For network variant instances in the dataset The loss function on ;
[0040] Get a set of optional neural network sets M 2 , and the independent weight set corresponding to each instance in the set {Weight 1 ,Weight 2 ,…Weight m(c)},c∈C, completing the discrete weight-independent neural network set.
[0041] Furthermore, the step 12) specifically includes: screening a subset M′ that meets the requirements of a specific scenario on the target device through the target detection neural network set M established in step 11).
[0042] Furthermore, the step 12) specifically includes:
[0043] 121) Define performance evaluation indicators;
[0044] 1211) Accuracy evaluation index: Mean Average Precision (mAP) is used as the accuracy evaluation index of the target detection network; specifically, the pycocotools toolkit is used to calculate the mAP value (mAP@[IoU=0.50:0.95]) of the network under the condition of an IoU threshold of 0.5-0.95. The mAP value (mAP@[IoU=0.50:0.95]) is obtained by the following method:
[0045] Convert the detection results output by the network (including the target bounding box coordinates and confidence) into COCO evaluation format;
[0046] Initialize the estimator using the pycocotools.COCOeval class;
[0047] Evaluate the detection results on the test set and obtain the mAP@[IoU=0.50:0.95] value in the evaluation results;
[0048] The higher the calculated mAP@[IoU=0.50:0.95] value, the better the detection accuracy of the network;
[0049] 1212) Delay evaluation index: The inference delay is used as the index for network performance evaluation. The evaluation method is as follows:
[0050] The time cost of statistical network processing a single frame of image;
[0051] Record the starting timestamp t1 before network inference;
[0052] After obtaining the test results, record the end timestamp t2;
[0053] Calculate the single-frame inference delay, delay = t2-t1;
[0054] Continuously process N frames of images;
[0055] Calculate the sum of the single-frame inference latency of N frames and divide it by N to get the average single-frame inference latency;
[0056] Record the statistical characteristics of the delay, including maximum value, minimum value, and standard deviation;
[0057] Evaluate the real-time performance of the network through statistical analysis of network inference latency;
[0058] 122) Evaluate the performance of neural network ensembles;
[0059] The evaluation indicators defined in step 121) are systematically tested and analyzed. Each optional neural network configuration generates an evaluation record. Each evaluation record is described in the form of a triple, as follows:
[0060] Record=(Arch params ,accuracy,latency)
[0061] Among them, Arch params Represents the architecture parameter configuration of the neural network, accuracy represents the detection accuracy on the validation set, and latency represents the average inference latency on the target device;
[0062] 1221) Build a performance test environment;
[0063] A distributed evaluation system architecture is adopted to decouple the performance evaluation process into accuracy evaluation and latency evaluation, where accuracy evaluation is performed on the server and latency evaluation is performed on the target device.
[0064] The server is used as the main control node of the performance test to schedule and control the performance evaluation process; the target device is used as the execution node of the performance test to measure the average single-frame inference latency;
[0065] The server is used as a socket server to listen to the specified port; the target device is used as a socket client to connect to the server; after the two parties establish a TCP connection, the following interactions are performed:
[0066] (a) The server selects a specific architecture configuration parameter and performs accuracy evaluation locally;
[0067] (b) The server sends architecture configuration parameters to the target device;
[0068] (c) The target device loads the corresponding configuration and, depending on the type of configuration (weight-sharing supernetwork configuration or discrete neural network configuration), activates a subnetwork defined in step 111), or loads an independent network defined in step 112);
[0069] (d) The target device performs inference latency measurement;
[0070] (e) The target device returns the measurement results to the server;
[0071] Repeat steps (a) to (e) until a set of performance evaluation records is obtained;
[0072] According to the specific type of the target detection neural network set in step 11), it is divided into a hypernetwork-based neural network evaluation and a discrete neural network set evaluation;
[0073] 1222) HyperNetwork-based Neural Network Evaluation: Using Optuna-based automated hyperparameter optimization methods to search for optimal neural network configurations through intelligent sampling and multi-objective optimization, specifically including:
[0074] 12221) Sub-network sampling method: Sampling and evaluating the configurable parameters in the super-network. The specific steps are as follows:
[0075] Architecture parameter sampling: mapping the optional configuration parameters in the parameter space C in step 111) into a number of discrete adjustable parameters, and the value range of each adjustable parameter constitutes a search space; during the sampling process, the search space of each parameter is effectively explored through an optimization algorithm to generate a set of architecture parameter configurations;
[0076] Gradually refine the evaluation: As the sampling iteration proceeds, the amount of data used in each sampling evaluation is gradually increased;
[0077] 12222) Pareto optimal solution search: Find the optimal balance point between different evaluation indicators, as follows:
[0078] (f) Initialize the search space: determine all adjustable architecture parameters of the hypernetwork and their value ranges, including the number of network layers and channel width; define optimization objectives, including maximizing accuracy and minimizing inference latency;
[0079] (g) Generate initial population: Randomly select a set of architecture configurations from the search space as the set of initial solutions to ensure population diversity to cover a wider range of configuration space;
[0080] (h) Evaluate performance: For each architecture configuration, measure its accuracy and inference latency and record the evaluation results;
[0081] (i) Pareto sorting: Based on the evaluation results, the solutions in the population are sorted by non-dominated levels and each solution is assigned to a different Pareto level. Non-dominated solutions are preferentially retained, i.e., configurations that are not inferior to other solutions in any objective.
[0082] (j) Select the next generation of solutions: Select solutions with higher Pareto levels from the current population as the basis, and introduce a diversity mechanism to avoid falling into local optimality; on the basis of maintaining solution diversity, streamline the solution set;
[0083] (k) Generate new solutions: Apply optimization operations (such as mutation and crossover) to the current solution set to generate a new generation of architectural configurations to ensure the exploration of new solutions and the improvement of existing solutions;
[0084] (l) Iterative optimization: Repeat steps (h) to (k) to gradually optimize the population; in each round of iteration, dynamically update the Pareto frontier;
[0085] (m) Output configuration set: After the preset number of iterations is reached or the stopping condition is met, the final optimized configuration set is generated. The set contains the configurations that achieve the best trade-off between accuracy and inference latency, and records the architectural parameters and corresponding performance data of each configuration;
[0086] 1223) Discrete neural network set evaluation: perform latency and accuracy evaluation on each network configuration in the discrete neural network set, and record the architecture parameters and corresponding performance data of each configuration;
[0087] 123) Filtering a subset of network configurations;
[0088] Based on the performance evaluation results of step 122), according to the specific requirements of the actual application scenario, the screening criteria are set to filter the neural network set, and a subset of network configurations that meet the deployment requirements is selected. The screening process is as follows:
[0089] 1231) screening base constraints;
[0090] Set hard requirements for neural network performance, including:
[0091] Minimum accuracy threshold min ): The network detection accuracy must be higher than the minimum accuracy threshold;
[0092] Maximum allowed delay max ): Network inference latency must be lower than the maximum allowed latency;
[0093] For any network configuration (α, A, T), it must satisfy the following conditions at the same time:
[0094] A≥Accuracy min
[0095] T≤Latency max ;
[0096] For unsatisfactory network configurations, remove the network configuration subset;
[0097] 1232) Filter redundant configuration;
[0098] Grid clustering is performed on the accuracy-delay plane as follows:
[0099] Divide the accuracy dimension and delay dimension into grids of size δA and δT respectively, where δA is the grid unit interval of accuracy (e.g., 5%), and δT is the grid unit interval of delay (e.g., 0.05s);
[0100] For multiple configurations that fall within the same grid cell, only the one with the best performance is retained;
[0101] For configurations i and j in a grid cell, if |Ai-Aj|≤δA and |Ti-Tj|≤δT are satisfied, the configuration with higher accuracy is selected and retained;
[0102] 1233) output configuration set;
[0103] Output the final subset of neural network configurations that pass the screening. Each configuration contains a single neural network configuration architecture parameter, accuracy, and latency information. The configuration subset will serve as a candidate set for subsequent neural network deployment.
[0104] Furthermore, the step 2) specifically includes:
[0105] 21) Initial neural network configuration;
[0106] 22) receiving video stream;
[0107] 23) Perform target detection on the received video stream, determine whether the current neural network configuration needs to be updated with a certain decision cycle, and perform corresponding processing.
[0108] Furthermore, the step 21) specifically includes:
[0109] Select the initial neural network configuration m from the subset M′ in step 1) init , give priority to selecting the neural network configuration with a balance between accuracy and latency as the initial configuration of the current neural network architecture parameters (for example, sort the optional configurations by latency or accuracy, and take the middle configuration as the initial configuration); for different neural network configuration types, perform different initialization steps, as follows:
[0110] 211) If it is a subnetwork configuration of a weight-sharing supernetwork, then load the supernetwork shared weight Weight share , and set the current active sub-network according to the current neural network architecture parameter configuration. After activating the current active sub-network, update the statistics of the batch normalization layer in the current active sub-network (when switching sub-networks in a weight-sharing super-network, the statistics of the batch normalization layer need to be recalculated because the number of activation channels of each layer may have changed, otherwise it will cause a huge drop in accuracy);
[0111] If the batch normalization parameter in the subnetwork weight is (μ x ,σ x ), where μ x is the mean of the input features, σ x is the standard deviation of the input features, then the calibration dataset To update:
[0112]
[0113] The statistics of the batch normalization layer are pre-calibrated and stored as a separate weight file for fast sub-network switching;
[0114] 212) If it is a discrete weight-independent neural network configuration, directly load the corresponding independent weight Weight discrete .
[0115] Furthermore, the step 22) specifically includes:
[0116] Send a video stream start request containing stream_id to the server, read the video frame through OpenCV, encode the image into JPEG format and convert it into base64 string; construct a frame data packet containing its meta information (including stream_id, frame_id and timestamps) and image data, and send it to the server through Socket; when the video stream processing is terminated, send a stream stop request and release related resources.
[0117] Furthermore, the step 23) specifically includes:
[0118] 231) Create a fixed-size frame processing queue for caching frames to be processed. The queue adopts a FIFO (first-in-first-out) strategy. When receiving video frame data, decode the base64 string into an image array, record the frame reception timestamp, and add the frame to be processed and its meta information (including stream_id, frame_id and timestamps) to the processing queue; when the queue reaches the maximum length, the entry of a new frame will cause the earliest frame to be discarded to ensure the real-time performance of the system;
[0119] 232) looping to obtain frames to be processed from the queue, for each frame to be processed, calling the currently active target detection neural network to perform calculations to obtain a detection result, the detection result including a detection box, a category, and a confidence score;
[0120] Detection box: Each detected object corresponds to a rectangular box, which can be represented as b i =(x min ,y min ,x max ,y max ), where (x min ,y min ),(x max ,y max ) are the coordinates of the upper left corner and lower right corner of the rectangular box respectively;
[0121] Category: The category to which the object detected in each detection frame belongs;
[0122] Confidence score: Each detection box corresponds to a confidence score c i ∈[0,1], indicating the probability of predicting that the object in the detection box is of a specific category;
[0123] Perform visual rendering on the original image (such as drawing detection boxes, category labels, etc.) and record the processing completion timestamp;
[0124] 233) re-encode the processed frame into base64 format, construct a data packet containing the detection result and timestamp information and output it;
[0125] 234) Update the performance statistics of the current service and provide query APIs to the outside world, including:
[0126] Record the detection accuracy based on the mAP@[IoU=0.50:0.95] value of the current network;
[0127] The processing delay is calculated by the difference between the processing completion timestamp and the processing start timestamp;
[0128] Count the current queue length as a system load indicator;
[0129] Calculate the average confidence of the test results, where the average confidence is the average of all confidence scores in the test results;
[0130] Calculate the average object size of the detection results. The average object size is the average length and width of all detection boxes in the detection results.
[0131] Calculate the brightness and contrast of the current image;
[0132] 235) Determine whether the current neural network configuration needs to be updated, and perform corresponding processing, specifically including:
[0133] 2351)Configure update status detection:
[0134] Detect the current configuration update status. If it is detected that the configuration update status needs to be updated, read the target configuration parameters and verify the validity of the configuration update, including the update timestamp and the integrity of the configuration parameters.
[0135] 2352) Processing according to the test results:
[0136] When it is detected that the configuration update status is that an update is required, the current timestamp and update information are recorded, and the process returns to step 21) to update the neural network configuration. After the configuration update is completed, the process returns to step 232) to continue processing the video stream.
[0137] When it is detected that the configuration update state does not need to be updated, the current neural network configuration is maintained, and the process returns to step 232 to continue processing the video stream;
[0138] 2353)Configure switching status record:
[0139] Record the configuration-related status information during the execution of step 2352), including the configuration check timestamp, configuration update status, and configuration switch execution status; at the same time, keep tracking the configuration status, record the currently active network configuration information, and maintain the configuration update history record.
[0140] Furthermore, the step 3) specifically includes:
[0141] 31) Timing state feature update;
[0142] 311) Use a time window structure based on a double-ended queue to cache the system operation status: set the max_window maximum window size to limit the number of historical state records; set the observation window size observation_window to determine the length of the state sequence used for each decision; set the decision time interval decision_interval to control the execution frequency of the neural network configuration decision;
[0143] 312) Obtain the latest performance statistics, including: detection accuracy accuracy, processing delay latency, current queue length queue_length, average confidence of detection results, average target size average_size of detection results, and brightness and contrast of the current image; add the obtained performance statistics as the current system state to the state cache queue;
[0144] 313) The collected performance statistics are normalized and each indicator is mapped to a unified numerical range; among them, the accuracy is normalized relative to the maximum accuracy max_accuracy, the delay keeps the original value to reflect the actual processing time, the queue length is normalized relative to the maximum system cache queue length queue_max_length, the confidence is normalized, and the image features are normalized relative to their respective theoretical maximum values;
[0145] 32) Neural network configuration decision;
[0146] 321) Define the decision network structure based on LSTM-DQN;
[0147] 3211) defining state input features, including all state indicators normalized in step 313);
[0148] 3212) Construct a time series feature extraction layer, use the LSTM network structure to process the continuous state sequence, and extract the time series features and dependencies of the state sequence through the multi-layer network structure;
[0149] 3213) Construct a decision output layer and use a multi-layer fully connected network to map the time series features to the action value space to evaluate the expected benefits of different configuration actions;
[0150] 3214) Introducing the dropout mechanism and gradient clipping strategy to improve the stability and generalization ability of network training;
[0151] 322) Based on the neural network configuration subset M′, a discrete action space is defined:
[0152] 3221) Establish a mapping relationship from configuration to action, map each candidate neural network configuration to a discrete action, and form an action set A = {a 1 ,a 2 ,…,a n}, where a i ∈M′, n is the size of the neural network configuration subset |M′|;
[0153] 3222) constructing an action-to-configuration mapping c=map(a), ensuring that the action a selected from the action space A corresponds to the neural network configuration parameter c;
[0154] 323) define state representation;
[0155] Define the state s at time t t is the state sequence of observation_window length, s t ={x t-w+1 ,x t-w+2 ,...,x t}, where w = observation_window, x t-i represents the system state at time ti, and the system state x at each time t Contains all normalized state indicators in step 313);
[0156] 324) Processing state feature sequence;
[0157] 3241) Based on the state cache in step 31), obtain the most recent observation_window state samples to form the state s in step 323) t , forming a state feature sequence;
[0158] 3242) Preprocess the state feature sequence. When the number of state samples is less than observation_window, the sequence is filled by copying the earliest state sample to ensure the integrity of the state feature sequence;
[0159] 3243) Performing tensor conversion on the processed state feature sequence as input of the decision network;
[0160] 325) Action selection based on ε-greedy strategy;
[0161] 3251) Perform random exploration with probability ε, randomly select an action with medium probability from the action space A to explore new possibilities in the environment;
[0162] 3252) Select the optimal action with probability 1-ε, input the current state sequence into the decision network, obtain the expected value estimates of all actions, and select the action with the largest expected value;
[0163] 3253) Set the initial exploration probability ε init , minimum exploration probability ε min and the decay rate ε decay ,With the increase of the number of decisions, the probability of random exploration is gradually reduced, and a smooth transition from exploration to utilization is achieved;
[0164] 33) converting the action selected in step 325) into an actual neural network configuration and performing switching;
[0165] 331) Action mapping conversion: For the currently selected action a t , get the corresponding configuration parameter c t =map(a t );
[0166] 332) configuration update action execution;
[0167] 3321) Build configuration update status information, including target configuration parameters ct and switch timestamp;
[0168] 3322) using the information obtained in step 3321) to update the configuration update status in step 235) and perform network configuration switching;
[0169] 3323) Record the current configuration update action, including the initiation time, target configuration and switching state identification;
[0170] 333) Switching status monitoring;
[0171] 3331) receiving the configuration switching progress information returned in step 235);
[0172] 3332) Tracking the intermediate states during the switching process, including configuration loading state and neural network switching state;
[0173] 3333) Detect abnormal conditions during the switching process, record error information when an abnormality occurs, and perform alarm and abnormality processing;
[0174] 334) Switching completion confirmation;
[0175] 3341) receiving a switching completion notification;
[0176] 3342) Verify the effectiveness of the new configuration;
[0177] 3343) Update the current configuration record of the system;
[0178] 34) Decision-making network optimization mechanism;
[0179] 341) Reward function design;
[0180] 3411) Based on the state cache in step 31), obtain the state sequence {x t-decision_interval+1 ,x t-decision_interval+2 ,...,x t}, where decision_interval is the decision time interval;
[0181] 3412) Define the queue pressure coefficient p queue :
[0182] p queue =min(1.0,queue_length / queue_max_length);
[0183] 3413) Design a dynamic weighted reward calculation method:
[0184] Performance score perf , score perf =wacc accuracy+w conf ·confidence, where w acc and w conf is the weight coefficient;
[0185] Delay score latency , score latency =exp(-latency);
[0186] Adaptively adjust weights based on queue pressure:
[0187] w accuracy =1.0-p queue
[0188] w latency =p queu e
[0189] Among them, w accuracy is the weight of the accuracy index, w latency is the weight of the delay indicator;
[0190] 3414) Calculate the comprehensive reward value r t ,as follows:
[0191] r t =w accuracy ×score perf +w latency ×score latency +r penalty
[0192] Among them, r penalty As an additional reward adjustment, a penalty is imposed when the queue length exceeds a threshold;
[0193] 342) Experience replay pool maintenance and update;
[0194] 3421) Maintain a fixed-capacity experience replay cache D to store decision transition quads (s t ,a t ,r t ,s t+1 );
[0195] 3422) When executing action s t After that, calculate the reward r t , store the conversion records in the experience replay cache;
[0196] 3423) Randomly sample batch data from the experience replay buffer D for training;
[0197] 343) Time series value estimation:
[0198] 3431) Define the discount factor γ to balance immediate rewards and long-term benefits;
[0199] 3432) Calculate the weighted sum of the current immediate reward and the future expected return as follows:
[0200] Q target =r t +γ×max{Q(s t+1 ,a)|a∈A}
[0201] Among them, Q target Represents the current instant reward r t The weighted sum of the expected future returns is used to guide the target value of strategy optimization; A is the action space, max{Q(s t+1 ,a)|a∈A} represents the next state s t+1 The maximum expected return among all possible actions;
[0202] 3433) Calculate the value of the decision network for the current state and action combination as follows:
[0203] Q current =Q(s t ,a t )
[0204] Among them, Q current Represents the decision network for the current state s t and action a t The combined value estimate reflects the current policy's evaluation of the state-action pair;
[0205] 3434) Define the timing difference error δt as follows:
[0206] δt=Q target -Q current ;
[0207] 344) Gradient optimization strategy;
[0208] 3441) The loss function is calculated based on the time difference error δt as follows:
[0209] L(θ)=E[(Q target -Q current ) 2 ]
[0210] Among them, θ is the parameter of the decision network, and the loss function L(θ) represents the deviation between the predicted value of the decision network and the target value;
[0211] 3442) The gradient descent method is used to update the decision network parameters as follows:
[0212]
[0213] Among them, α is the learning rate.
[0214] Beneficial effects of the present invention:
[0215] (1) The present invention realizes dynamic self-adaptation of neural network configuration: by establishing a decision-making mechanism based on time series state, the neural network configuration can be automatically adjusted according to factors such as the complexity of the real-time video scene and the system load condition, effectively solving the problem of fixed configuration and inability to dynamically adjust in the prior art;
[0216] (2) The present invention provides complete resource perception capabilities: based on the collection and analysis of multi-dimensional indicators of the system operation status, including detection accuracy, processing delay, queue length, etc., the system can accurately evaluate the resource overhead of different configurations and provide a reliable basis for neural network configuration decisions;
[0217] (3) The present invention establishes a smooth configuration switching mechanism: through the configuration switching state monitoring and exception handling mechanism, the stability of the neural network configuration switching process is ensured, avoiding the service interruption problem caused by configuration switching in the traditional method;
[0218] (4) The present invention realizes multi-objective optimization based on deep reinforcement learning: by adopting a dynamic weight reward mechanism, the system can adaptively balance between accuracy and real-time performance to meet the performance requirements in different scenarios;
[0219] (5) The present invention has strong versatility and scalability: the neural network configuration method proposed in the present invention supports both hypernetwork-based weight-sharing neural networks and discrete weight-independent neural networks, and has good versatility; and its state feature definition, reward calculation and other mechanisms maintain strong scalability.
[0220] The present invention can effectively improve the adaptability and stability of the video stream target detection system in the edge computing environment and has important practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0221] Figure 1 is a flow chart of the method of the present invention;
[0222] Figure 2 Build and select flowcharts for neural network collections;
[0223] Figure 3 Flowchart for video stream processing and configuration switching:
[0224] Figure 4 Schematic diagram for configuring an adaptive neural network. DETAILED DESCRIPTION
[0225] In order to facilitate the understanding of those skilled in the art, the present invention is further described below in conjunction with embodiments and drawings. The contents mentioned in the implementation modes are not intended to limit the present invention.
[0226] Reference Figures 1 to 4 As shown, an adaptive neural network configuration method for video stream target detection of the present invention comprises the following steps:
[0227] 1) Build and select a set of target detection neural networks; specifically include:
[0228] 11) Construct a set of target detection neural networks;
[0229] 12) Filter the target detection neural network collection.
[0230] The target detection neural network set in step 11) is a weight-sharing neural network set based on a hypernetwork, a discrete weight-independent neural network set, or a combination of the two; the weight-sharing neural network set based on a hypernetwork is formed by constructing a large-scale weight-sharing hypernetwork and extracting subnetworks of different configurations therefrom; the discrete weight-independent neural network set is composed of multiple independently trained target detection networks; specifically includes:
[0231] 111) Construct a set of weight-sharing neural networks based on a hypernetwork as follows:
[0232] 1111) Construct a detection backbone network with configurable parameters, wherein the configurable parameters include: the number of backbone network depth layers D, the ratio of the number of channels in each layer W, the convolution kernel size K, and the input resolution R; wherein D = {d 1 ,d 2 ,…d m},d i ∈N+,W={w 1 ,w 2 ,…w n},w i ∈R+,K={k 1 ,k 2 ,…k p},k i ∈N+,R={r},r∈N+,d i represents the depth of the i-th block, w i represents the channel number multiplier of the i-th layer, k i represents the convolution kernel size of the i-th layer, and r represents the input resolution;
[0233] 1112) construct a feature pyramid network including configurable parameters, wherein the configurable parameters include: the number of feature pyramid levels L and the multiple of the number of channels of each feature pyramid layer W′, where L={l 1,l 2 ,…l r},l i ∈{0,1},W′={m 1 ,m 2 ,…m s},m i ∈R+,l i Indicates whether to select the feature of the i-th layer (0 for not selecting, 1 for selecting), m i Indicates the channel number multiplier of the i-th layer;
[0234] 1113) After constructing the feature pyramid network, a detection head network is constructed, and a two-stage detection head structure based on region proposal (such as Faster R-CNN) or a single-stage detection head structure based on anchor-free (such as FCOS) is selected to achieve the final target detection and generate target category and location information;
[0235] 1114) combining the networks in step 1111), step 1112) and step 1113) to obtain a weight-sharing neural network based on a hypernetwork;
[0236] For any configuration parameter combination c, c = (d, w, k, r, l, m), d ∈ D, w ∈ W, k ∈ K, r ∈ R, l ∈ L, m ∈ W ′, the corresponding sub-network can be obtained from the super-network. The resulting neural network set is expressed as:
[0237] M 1 ={s(c)|c∈C}
[0238] Where C is the Cartesian product of the parameter space, C = D × W × K × R × L × W ′, s(c) represents the subnetwork with parameter configuration c;
[0239] 1115) train all sub-networks in step 1114), update the shared weights by randomly sampling sub-networks during training, and obtain the shared weights of the super-network through training; specifically, for the currently sampled sub-network s(c), update the shared weights by gradient descent, as follows:
[0240]
[0241] Among them, θ s(c) is the shared weight of the current sampling sub-network s(c), η is the learning rate, is the loss function of the sub-network on the dataset D, is θ s(c) The gradient of
[0242] Get a set of optional network sets M 1 , and the shared weights of the sub-networks in the network setshare , complete the construction of a set of weight-sharing neural networks based on a hypernetwork;
[0243] Optionally, a multi-scale feature distillation approach can be used. After randomly sampling the sub-networks, the output of the largest network in the search space or the output of a pre-trained teacher network is used to guide the training of the currently sampled sub-network.
[0244] 112) Construct a discrete set of weight-independent neural networks as follows:
[0245] 1121) Build a set of target detection networks with configurable parameters. The network types include: networks based on convolutional neural networks (CNN) (such as YOLOv5 series, EfficientDet series) and networks based on self-attention mechanisms (such as DETR series);
[0246] Different network variants can be designed for each network. The optional configuration parameters of the network variants include: input resolution R, number of backbone network depth layers D and the ratio of the number of channels in each layer W, R = {r}, r∈N+, D = {d 1 ,d 2 ,…d m},d i ∈N+,W={w 1 ,w 2 ,…w n},w i ∈R+;
[0247] By combining the above configuration parameters, a variety of network variants are generated, and the resulting neural network set is expressed as:
[0248] M 2 ={m(c)|c∈C}
[0249] Where C is the Cartesian product of the parameter space, C = R × D × W, m(c) represents a network variant instance with parameter configuration c;
[0250] 1122) Each network variant is trained separately to obtain independent weights of each network; specifically, for each network variant, the independent weights are updated by gradient descent as follows:
[0251]
[0252] Among them, θ m(c) is the network variant instance m(c), η is the learning rate, For network variant instances in the dataset The loss function on ;
[0253] Get a set of optional neural network sets M 2, and the independent weight set corresponding to each instance in the set {Weight 1 ,Weight 2 ,…Weight m(c)},c∈C, completing the discrete weight-independent neural network set.
[0254] Wherein, the step 12) specifically includes: through the target detection neural network set M established in step 11), screening a subset M′ that meets the requirements of a specific scenario on the target device.
[0255] Wherein, the step 12) specifically includes:
[0256] 121) Define performance evaluation indicators;
[0257] 1211) Accuracy evaluation index: Mean Average Precision (mAP) is used as the accuracy evaluation index of the target detection network; specifically, the pycocotools toolkit is used to calculate the mAP value (mAP@[IoU=0.50:0.95]) of the network under the condition of an IoU threshold of 0.5-0.95. The mAP value (mAP@[IoU=0.50:0.95]) is obtained by the following method:
[0258] Convert the detection results output by the network (including the target bounding box coordinates and confidence) into COCO evaluation format;
[0259] Initialize the estimator using the pycocotools.COCOeval class;
[0260] Evaluate the detection results on the test set and obtain the mAP@[IoU=0.50:0.95] value in the evaluation results;
[0261] The higher the calculated mAP@[IoU=0.50:0.95] value, the better the detection accuracy of the network;
[0262] 1212) Delay evaluation index: The inference delay is used as the index for network performance evaluation. The evaluation method is as follows:
[0263] The time cost of statistical network processing a single frame of image;
[0264] Record the starting timestamp t1 before network inference;
[0265] After obtaining the test results, record the end timestamp t2;
[0266] Calculate the single-frame inference delay, delay = t2-t1;
[0267] Continuously process N frames of images;
[0268] Calculate the sum of the single-frame inference latency of N frames and divide it by N to get the average single-frame inference latency;
[0269] Record the statistical characteristics of the delay, including maximum value, minimum value, and standard deviation;
[0270] Evaluate the real-time performance of the network through statistical analysis of network inference latency;
[0271] 122) Evaluate the performance of neural network ensembles;
[0272] The evaluation indicators defined in step 121) are systematically tested and analyzed. Each optional neural network configuration generates an evaluation record. Each evaluation record is described in the form of a triple, as follows:
[0273] Record=(Arch params ,accuracy,latency)
[0274] Among them, Arch params Represents the architecture parameter configuration of the neural network, accuracy represents the detection accuracy on the validation set, and latency represents the average inference latency on the target device;
[0275] 1221) Build a performance test environment;
[0276] A distributed evaluation system architecture is adopted to decouple the performance evaluation process into accuracy evaluation and latency evaluation, where accuracy evaluation is performed on the server and latency evaluation is performed on the target device.
[0277] The server is used as the main control node of the performance test to schedule and control the performance evaluation process; the target device is used as the execution node of the performance test to measure the average single-frame inference latency;
[0278] The server is used as a socket server to listen to the specified port; the target device is used as a socket client to connect to the server; after the two parties establish a TCP connection, the following interactions are performed:
[0279] (a) The server selects a specific architecture configuration parameter and performs accuracy evaluation locally;
[0280] (b) The server sends architecture configuration parameters to the target device;
[0281] (c) The target device loads the corresponding configuration and, depending on the type of configuration (weight-sharing supernetwork configuration or discrete neural network configuration), activates a subnetwork defined in step 111), or loads an independent network defined in step 112);
[0282] (d) The target device performs inference latency measurement;
[0283] (e) The target device returns the measurement results to the server;
[0284] Repeat steps (a) to (e) until a set of performance evaluation records is obtained;
[0285] According to the specific type of the target detection neural network set in step 11), it is divided into a hypernetwork-based neural network evaluation and a discrete neural network set evaluation;
[0286] 1222) HyperNetwork-based Neural Network Evaluation: Using Optuna-based automated hyperparameter optimization methods to search for optimal neural network configurations through intelligent sampling and multi-objective optimization, specifically including:
[0287] 12221) Sub-network sampling method: Sampling and evaluating the configurable parameters in the super-network. The specific steps are as follows:
[0288] Architecture parameter sampling: mapping the optional configuration parameters in the parameter space C in step 111) into a number of discrete adjustable parameters, and the value range of each adjustable parameter constitutes a search space; during the sampling process, the search space of each parameter is effectively explored through an optimization algorithm to generate a set of architecture parameter configurations;
[0289] Gradually refine the evaluation: As the sampling iteration proceeds, the amount of data used in each sampling evaluation is gradually increased;
[0290] 12222) Pareto optimal solution search: Find the optimal balance point between different evaluation indicators, as follows:
[0291] (f) Initialize the search space: determine all adjustable architecture parameters of the hypernetwork and their value ranges, including the number of network layers and channel width; define optimization objectives, including maximizing accuracy and minimizing inference latency;
[0292] (g) Generate initial population: Randomly select a set of architecture configurations from the search space as the set of initial solutions to ensure population diversity to cover a wider range of configuration space;
[0293] (h) Evaluate performance: For each architecture configuration, measure its accuracy and inference latency and record the evaluation results;
[0294] (i) Pareto sorting: Based on the evaluation results, the solutions in the population are sorted by non-dominated levels and each solution is assigned to a different Pareto level. Non-dominated solutions are preferentially retained, i.e., configurations that are not inferior to other solutions in any objective.
[0295] (j) Select the next generation of solutions: Select solutions with higher Pareto levels from the current population as the basis, and introduce a diversity mechanism to avoid falling into local optimality; on the basis of maintaining solution diversity, streamline the solution set;
[0296] (k) Generate new solutions: Apply optimization operations (such as mutation and crossover) to the current solution set to generate a new generation of architectural configurations to ensure the exploration of new solutions and the improvement of existing solutions;
[0297] (l) Iterative optimization: Repeat steps (h) to (k) to gradually optimize the population; in each round of iteration, dynamically update the Pareto frontier;
[0298] (m) Output configuration set: After the preset number of iterations is reached or the stopping condition is met, the final optimized configuration set is generated. The set contains the configurations that achieve the best trade-off between accuracy and inference latency, and records the architectural parameters and corresponding performance data of each configuration;
[0299] 1223) Discrete neural network set evaluation: perform latency and accuracy evaluation on each network configuration in the discrete neural network set, and record the architecture parameters and corresponding performance data of each configuration;
[0300] 123) Filtering a subset of network configurations;
[0301] Based on the performance evaluation results of step 122), according to the specific requirements of the actual application scenario, the screening criteria are set to filter the neural network set, and a subset of network configurations that meet the deployment requirements is selected. The screening process is as follows:
[0302] 1231) screening base constraints;
[0303] Set hard requirements for neural network performance, including:
[0304] Minimum accuracy threshold min ): The network detection accuracy must be higher than the minimum accuracy threshold;
[0305] Maximum allowed delay max ): Network inference latency must be lower than the maximum allowed latency;
[0306] For any network configuration (α, A, T), it must satisfy the following conditions at the same time:
[0307] A≥Accuracy min
[0308] T≤Latency max ;
[0309] For unsatisfactory network configurations, remove the network configuration subset;
[0310] 1232) Filter redundant configuration;
[0311] Grid clustering is performed on the accuracy-delay plane as follows:
[0312] Divide the accuracy dimension and delay dimension into grids of size δA and δT respectively, where δA is the grid unit interval of accuracy (e.g., 5%), and δT is the grid unit interval of delay (e.g., 0.05s);
[0313] For multiple configurations that fall within the same grid cell, only the one with the best performance is retained;
[0314] For configurations i and j in a grid cell, if |Ai-Aj|≤δA and |Ti-Tj|≤δT are satisfied, the configuration with higher accuracy is selected and retained;
[0315] 1233) output configuration set;
[0316] Output the final subset of neural network configurations that pass the screening. Each configuration contains a single neural network configuration architecture parameter, accuracy, and latency information. The configuration subset will serve as a candidate set for subsequent neural network deployment.
[0317] 2) Using the neural network in the selected target detection neural network set to perform video stream target detection; judging whether the current neural network configuration needs to be updated in a certain decision cycle, if so, returning to step 2); otherwise, entering step 3);
[0318] Wherein, the step 2) specifically includes:
[0319] 21) Initial neural network configuration;
[0320] 22) receiving video stream;
[0321] 23) Perform target detection on the received video stream, determine whether the current neural network configuration needs to be updated with a certain decision cycle, and perform corresponding processing.
[0322] Specifically, the step 21) specifically includes:
[0323] Select the initial neural network configuration m from the subset M′ in step 1) init , give priority to selecting the neural network configuration with a balance between accuracy and latency as the initial configuration of the current neural network architecture parameters (for example, sort the optional configurations by latency or accuracy, and take the middle configuration as the initial configuration); for different neural network configuration types, perform different initialization steps, as follows:
[0324] 211) If it is a subnetwork configuration of a weight-sharing supernetwork, then load the supernetwork shared weight Weight share , and set the current active sub-network according to the current neural network architecture parameter configuration. After activating the current active sub-network, update the statistics of the batch normalization layer in the current active sub-network (when switching sub-networks in a weight-sharing super-network, the statistics of the batch normalization layer need to be recalculated because the number of activation channels of each layer may have changed, otherwise it will cause a huge drop in accuracy);
[0325] If the batch normalization parameter in the subnetwork weight is (μ x ,σ x ), where μ x is the mean of the input features, σ x is the standard deviation of the input features, then the calibration dataset To update:
[0326]
[0327] The statistics of the batch normalization layer are pre-calibrated and stored as a separate weight file for fast sub-network switching;
[0328] 212) If it is a discrete weight-independent neural network configuration, directly load the corresponding independent weight Weight discrete .
[0329] Specifically, the step 22) specifically includes:
[0330] Send a video stream start request containing stream_id to the server, read the video frame through OpenCV, encode the image into JPEG format and convert it into base64 string; construct a frame data packet containing its meta information (including stream_id, frame_id and timestamps) and image data, and send it to the server through Socket; when the video stream processing is terminated, send a stream stop request and release related resources.
[0331] Specifically, the step 23) specifically includes:
[0332] 231) Create a fixed-size frame processing queue for caching frames to be processed. The queue adopts a FIFO (first-in-first-out) strategy. When receiving video frame data, decode the base64 string into an image array, record the frame reception timestamp, and add the frame to be processed and its meta information (including stream_id, frame_id and timestamps) to the processing queue; when the queue reaches the maximum length, the entry of a new frame will cause the earliest frame to be discarded to ensure the real-time performance of the system;
[0333] 232) looping to obtain frames to be processed from the queue, for each frame to be processed, calling the currently active target detection neural network to perform calculations to obtain a detection result, the detection result including a detection box, a category, and a confidence score;
[0334] Detection box: Each detected object corresponds to a rectangular box, which can be represented as b i =(x min ,y min ,x max ,y max ), where (x min ,y min ),(x max ,y max ) are the coordinates of the upper left corner and lower right corner of the rectangular box respectively;
[0335] Category: The category to which the object detected in each detection frame belongs;
[0336] Confidence score: Each detection box corresponds to a confidence score c i ∈[0,1], indicating the probability of predicting that the object in the detection box is of a specific category;
[0337] Perform visual rendering on the original image (such as drawing detection boxes, category labels, etc.) and record the processing completion timestamp;
[0338] 233) re-encode the processed frame into base64 format, construct a data packet containing the detection result and timestamp information and output it;
[0339] 234) Update the performance statistics of the current service and provide query APIs to the outside world, including:
[0340] Record the detection accuracy based on the mAP@[IoU=0.50:0.95] value of the current network;
[0341] The processing delay is calculated by the difference between the processing completion timestamp and the processing start timestamp;
[0342] Count the current queue length as a system load indicator;
[0343] Calculate the average confidence of the test results, where the average confidence is the average of all confidence scores in the test results;
[0344] Calculate the average object size of the detection results. The average object size is the average length and width of all detection boxes in the detection results.
[0345] Calculate the brightness and contrast of the current image;
[0346] 235) Determine whether the current neural network configuration needs to be updated, and perform corresponding processing, specifically including:
[0347] 2351)Configure update status detection:
[0348] Detect the current configuration update status. If it is detected that the configuration update status needs to be updated, read the target configuration parameters and verify the validity of the configuration update, including the update timestamp and the integrity of the configuration parameters.
[0349] 2352) Processing according to the test results:
[0350] When it is detected that the configuration update status is that an update is required, the current timestamp and update information are recorded, and the process returns to step 21) to update the neural network configuration. After the configuration update is completed, the process returns to step 232) to continue processing the video stream.
[0351] When it is detected that the configuration update state does not need to be updated, the current neural network configuration is maintained, and the process returns to step 232 to continue processing the video stream;
[0352] 2353)Configure switching status record:
[0353] Record the configuration-related status information during the execution of step 2352), including the configuration check timestamp, configuration update status, and configuration switch execution status; at the same time, keep tracking the configuration status, record the currently active network configuration information, and maintain the configuration update history record.
[0354] 3) Update the neural network configuration, including:
[0355] 31) Timing state feature update;
[0356] 311) Use a time window structure based on a double-ended queue to cache the system operation status: set the max_window maximum window size to limit the number of historical state records; set the observation window size observation_window to determine the length of the state sequence used for each decision; set the decision time interval decision_interval to control the execution frequency of the neural network configuration decision;
[0357] 312) Obtain the latest performance statistics, including: detection accuracy accuracy, processing delay latency, current queue length queue_length, average confidence of detection results, average target size average_size of detection results, and brightness and contrast of the current image; add the obtained performance statistics as the current system state to the state cache queue;
[0358] 313) The collected performance statistics are normalized and each indicator is mapped to a unified numerical range; among them, the accuracy is normalized relative to the maximum accuracy max_accuracy, the delay keeps the original value to reflect the actual processing time, the queue length is normalized relative to the maximum system cache queue length queue_max_length, the confidence is normalized, and the image features are normalized relative to their respective theoretical maximum values;
[0359] 32) Neural network configuration decision;
[0360] 321) Define the decision network structure based on LSTM-DQN;
[0361] 3211) defining state input features, including all state indicators normalized in step 313);
[0362] 3212) Construct a time series feature extraction layer, use the LSTM network structure to process the continuous state sequence, and extract the time series features and dependencies of the state sequence through the multi-layer network structure;
[0363] 3213) Construct a decision output layer and use a multi-layer fully connected network to map the time series features to the action value space to evaluate the expected benefits of different configuration actions;
[0364] 3214) Introducing the dropout mechanism and gradient clipping strategy to improve the stability and generalization ability of network training;
[0365] 322) Based on the neural network configuration subset M′, a discrete action space is defined:
[0366] 3221) Establish a mapping relationship from configuration to action, map each candidate neural network configuration to a discrete action, and form an action set A = {a 1 ,a 2 ,…,a n}, where a i ∈M′, n is the size of the neural network configuration subset |M′|;
[0367] 3222) constructing an action-to-configuration mapping c=map(a), ensuring that the action a selected from the action space A corresponds to the neural network configuration parameter c;
[0368] 323) define state representation;
[0369] Define the state s at time t t is the state sequence of observation_window length, s t ={x t-w+1 ,x t-w+2 ,...,x t}, where w = observation_window, x t-i represents the system state at time ti, and the system state x at each time t Contains all normalized state indicators in step 313);
[0370] 324) Processing state feature sequence;
[0371] 3241) Based on the state cache in step 31), obtain the most recent observation_window state samples to form the state s in step 323) t , forming a state feature sequence;
[0372] 3242) Preprocess the state feature sequence. When the number of state samples is less than observation_window, the sequence is filled by copying the earliest state sample to ensure the integrity of the state feature sequence;
[0373] 3243) Performing tensor conversion on the processed state feature sequence as input of the decision network;
[0374] 325) Action selection based on ε-greedy strategy;
[0375] 3251) Perform random exploration with probability ε, randomly select an action with medium probability from the action space A to explore new possibilities in the environment;
[0376] 3252) Select the optimal action with probability 1-ε, input the current state sequence into the decision network, obtain the expected value estimates of all actions, and select the action with the largest expected value;
[0377] 3253) Set the initial exploration probability ε init , minimum exploration probability ε min and the decay rate ε decay ,With the increase of the number of decisions, the probability of random exploration is gradually reduced, and a smooth transition from exploration to utilization is achieved;
[0378] 33) converting the action selected in step 325) into an actual neural network configuration and performing switching;
[0379] 331) Action mapping conversion: For the currently selected action a t , get the corresponding configuration parameter c t =map(a t );
[0380] 332) configuration update action execution;
[0381] 3321) Build configuration update status information, including target configuration parameters ct and switch timestamp;
[0382] 3322) using the information obtained in step 3321) to update the configuration update status in step 235) and perform network configuration switching;
[0383] 3323) Record the current configuration update action, including the initiation time, target configuration and switching state identification;
[0384] 333) Switching status monitoring;
[0385] 3331) receiving the configuration switching progress information returned in step 235);
[0386] 3332) Tracking the intermediate states during the switching process, including configuration loading state and neural network switching state;
[0387] 3333) Detect abnormal conditions during the switching process, record error information when an abnormality occurs, and perform alarm and abnormality processing;
[0388] 334) Switching completion confirmation;
[0389] 3341) receiving a switching completion notification;
[0390] 3342) Verify the effectiveness of the new configuration;
[0391] 3343) Update the current configuration record of the system;
[0392] 34) Decision-making network optimization mechanism;
[0393] 341) Reward function design;
[0394] 3411) Based on the state cache in step 31), obtain the state sequence {x t-decision_interval+1 ,x t-decision_interval+2 ,...,x t}, where decision_interval is the decision time interval;
[0395] 3412) Define the queue pressure coefficient p queue :
[0396] p queue =min(1.0,queue_length / queue_max_length);
[0397] 3413) Design a dynamic weighted reward calculation method:
[0398] Performance score perf , score perf =wacc acuracy+w conf ·confidence, where w acc and w conf is the weight coefficient;
[0399] Delay score latency , score latency =exp(-latency);
[0400] Adaptively adjust weights based on queue pressure:
[0401] w accuracy =1.0-p queue
[0402] w latency =p queue
[0403] Among them, w accuracy is the weight of the accuracy index, w latency is the weight of the delay indicator;
[0404] 3414) Calculate the comprehensive reward value r t ,as follows:
[0405] r t =w accuracy ×score perf +w latency ×score latency +r penalty
[0406] Among them, r penalty As an additional reward adjustment, a penalty is imposed when the queue length exceeds a threshold;
[0407] 342) Experience replay pool maintenance and update;
[0408] 3421) Maintain a fixed-capacity experience replay cache D to store decision transition quads (s t ,a t ,r t ,s t+1 );
[0409] 3422) When executing action s t After that, calculate the reward r t , store the conversion records in the experience replay cache;
[0410] 3423) Randomly sample batch data from the experience replay buffer D for training;
[0411] 343) Time series value estimation:
[0412] 3431) Define the discount factor γ to balance immediate rewards and long-term benefits;
[0413] 3432) Calculate the weighted sum of the current immediate reward and the future expected return as follows:
[0414] Q target =r t +γ×max{Q(s t+1 ,a)|a∈A}
[0415] Among them, Q target Represents the current instant reward r t The weighted sum of the expected future returns is used to guide the target value of strategy optimization; A is the action space, max{Q(s t+1 ,a)|a∈A} represents the next state s t+1 The maximum expected return among all possible actions;
[0416] 3433) Calculate the value of the decision network for the current state and action combination as follows:
[0417] Q current =Q(s t ,a t )
[0418] Among them, Q current Represents the decision network for the current state s t and action a t The combined value estimate reflects the current policy's evaluation of the state-action pair;
[0419] 3434) Define the timing difference error δt as follows:
[0420] δt=Q target -Q current ;
[0421] 344) Gradient optimization strategy;
[0422] 3441) The loss function is calculated based on the time difference error δt as follows:
[0423] L(θ)=E[(Q target -Q current ) 2 ]
[0424] Among them, θ is the parameter of the decision network, and the loss function L(θ) represents the deviation between the predicted value of the decision network and the target value;
[0425] 3442) The gradient descent method is used to update the decision network parameters as follows:
[0426]
[0427] Among them, α is the learning rate.
[0428] The present invention has many specific application paths. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements can be made without departing from the principle of the present invention. These improvements should also be regarded as the protection scope of the present invention.
Claims
1. An adaptive neural network configuration method for video stream target detection, characterized in that: Here are the steps: 1) Build and select a set of target detection neural networks; 2) Using the neural network in the selected target detection neural network set to perform video stream target detection; judging whether the current neural network configuration needs to be updated in a certain decision cycle, if so, returning to step 2); otherwise, entering step 3); 3) Update the neural network configuration.
2. The method for configuring an adaptive neural network for video stream target detection according to claim 1, characterized in that: The step 1) specifically includes: 11) Construct a set of target detection neural networks; 12) Filter the target detection neural network collection.
3. The method for configuring an adaptive neural network for video stream target detection according to claim 2, characterized in that: The target detection neural network set in step 11) is a weight-sharing neural network set based on a hypernetwork, a discrete weight-independent neural network set, or a combination of the two; the weight-sharing neural network set based on a hypernetwork is formed by constructing a large-scale weight-sharing hypernetwork and extracting subnetworks with different configurations therefrom to form a network set; The discrete weight-independent neural network set is composed of multiple independently trained target detection networks; specifically, it includes: 111) Construct a set of weight-sharing neural networks based on a hypernetwork as follows: 1111) Construct a detection backbone network with configurable parameters, wherein the configurable parameters include: the number of backbone network depth layers D, the ratio of the number of channels in each layer W, the convolution kernel size K, and the input resolution R; where D = {d1, d2, ... d m },d i ∈N+,W={w1,w2,…w n },w i ∈R+,K={k1,k2,…k p },k i ∈N+,R={r},r∈N+,d i represents the depth of the i-th block, w i represents the channel number multiplier of the i-th layer, k i represents the convolution kernel size of the i-th layer, and r represents the input resolution; 1112) construct a feature pyramid network containing configurable parameters, wherein the configurable parameters include: the number of feature pyramid levels L and the multiple W′ of the number of channels of each feature pyramid layer, where L = {l1, l2, ... l r },l i ∈{0,1},W′={m1,m2,…m s },m i ∈R+,l i Indicates whether to select the features of the i-th layer, m i Indicates the channel number multiplier of the i-th layer; 1113) After constructing the feature pyramid network, construct a detection head network, select a two-stage detection head structure based on region proposal, or a single-stage detection head structure based on no anchor point, to achieve the final target detection and generate target category and location information; 1114) combining the networks in step 1111), step 1112) and step 1113) to obtain a weight-sharing neural network based on a super network; For any configuration parameter combination c, c = (d, w, k, r, l, m), d ∈ D, w ∈ W, k ∈ K, r ∈ R, l ∈ L, m ∈ W ′ , the corresponding sub-network can be obtained from the super-network, and the obtained neural network set is expressed as: M1={s(c)|c∈C} Where C is the Cartesian product of the parameter space, C = D × W × K × R × L × W ′ , s(c) represents the sub-network with parameter configuration c; 1115) train all sub-networks in step 1114), update the shared weights by randomly sampling sub-networks during training, and obtain the shared weights of the super-network through training; specifically, for the currently sampled sub-network s(c), update the shared weights by gradient descent, as follows: Among them, θ s(c) is the shared weight of the current sampling sub-network s(c), η is the learning rate, For the subnetwork in the dataset The loss function on is θ s(c) The gradient of Get a set of optional network sets M1 and the shared weights Weight of the sub-networks in the network set share , complete the construction of a set of weight-sharing neural networks based on a hypernetwork; 112) Construct a discrete set of weight-independent neural networks as follows: 1121) constructing a target detection network set with configurable parameters, the network types include: a network based on a convolutional neural network and a network based on a self-attention mechanism; Different network variants can be designed for each network. The optional configuration parameters of the network variants include: input resolution R, number of backbone network depth layers D and the ratio of the number of channels in each layer W, R = {r}, r∈N+, D = {d1, d2, …d m },d i ∈N+,W={w1,w2,…w n },w i ∈R+; By combining the above configuration parameters, a variety of network variants are generated, and the resulting neural network set is expressed as: M2={m(c)|c∈C} Where C is the Cartesian product of the parameter space, C = R × D × W, m(c) represents a network variant instance with parameter configuration c; 1122) Each network variant is trained separately to obtain independent weights of each network; specifically, for each network variant, the independent weights are updated by gradient descent as follows: Among them, θ m(c) is the network variant instance m(c), η is the learning rate, For network variant instances in the dataset The loss function on ; Get a set of optional neural network sets M2, and the independent weight set {Weight1,Weight2,…Weight m(c) },c∈C, completing the discrete weight-independent neural network set.
4. The method for configuring an adaptive neural network for video stream target detection according to claim 3, characterized in that: The step 12) specifically includes: using the target detection neural network set M established in step 11), screening a subset M that meets the requirements of a specific scenario on the target device ′ .
5. The method for configuring an adaptive neural network for video stream target detection according to claim 3, characterized in that: The step 12) specifically includes: 121) Define performance evaluation indicators; 1211) Accuracy evaluation index: The average precision is used as the accuracy evaluation index of the target detection network; specifically, the pycocotools toolkit is used to calculate the mAP value of the network under the condition of IoU threshold of 0.5-0.95, mAP@[IoU=0.50:0.95], and the mAP value is obtained by the following method: Convert the detection results output by the network into COCO evaluation format; Initialize the estimator using the pycocotools.COCOeval class; Evaluate the detection results on the test set and obtain the mAP@[IoU=0.50:0.95] value in the evaluation results; The higher the calculated mAP@[IoU=0.50:0.95] value, the better the detection accuracy of the network; 1212) Delay evaluation index: The inference delay is used as the index for network performance evaluation. The evaluation method is as follows: The time cost of statistical network processing a single frame of image; Record the starting timestamp t1 before network inference; After obtaining the test results, record the end timestamp t2; Calculate the single-frame inference delay, delay = t2-t1; Continuously process N frames of images; Calculate the sum of the single-frame inference latency of N frames and divide it by N to get the average single-frame inference latency; Record the statistical characteristics of the delay, including maximum value, minimum value, and standard deviation; Evaluate the real-time performance of the network through statistical analysis of network inference latency; 122) Evaluate the performance of neural network ensembles; The evaluation indicators defined in step 121) are systematically tested and analyzed. Each optional neural network configuration generates an evaluation record. Each evaluation record is described in the form of a triple, as follows: Record=(Arch params ,accuracy,latency) Among them, Arch params Represents the architecture parameter configuration of the neural network, accuracy represents the detection accuracy on the validation set, and latency represents the average inference latency on the target device; 1221) Build a performance test environment; A distributed evaluation system architecture is adopted to decouple the performance evaluation process into accuracy evaluation and latency evaluation, where accuracy evaluation is performed on the server and latency evaluation is performed on the target device. The server is used as the main control node of the performance test to schedule and control the performance evaluation process; the target device is used as the execution node of the performance test to measure the average single-frame inference latency; The server is used as a socket server to listen to the specified port; the target device is used as a socket client to connect to the server; after the two parties establish a TCP connection, the following interactions are performed: (a) The server selects a specific architecture configuration parameter and performs accuracy evaluation locally; (b) The server sends architecture configuration parameters to the target device; (c) The target device loads the corresponding configuration and, depending on the type of configuration, activates a subnetwork defined in step 111) or loads an independent network defined in step 112); (d) The target device performs inference latency measurement; (e) The target device returns the measurement results to the server; Repeat steps (a) to (e) until a set of performance evaluation records is obtained; According to the specific type of the target detection neural network set in step 11), it is divided into a hypernetwork-based neural network evaluation and a discrete neural network set evaluation; 1222) HyperNetwork-based Neural Network Evaluation: Using Optuna-based automated hyperparameter optimization methods to search for optimal neural network configurations through intelligent sampling and multi-objective optimization, specifically including: 12221) Sub-network sampling method: Sampling and evaluating the configurable parameters in the super-network. The specific steps are as follows: Architecture parameter sampling: mapping the optional configuration parameters in the parameter space C in step 111) into a number of discrete adjustable parameters, and the value range of each adjustable parameter constitutes a search space; during the sampling process, the search space of each parameter is effectively explored through an optimization algorithm to generate a set of architecture parameter configurations; Gradually refine the evaluation: As the sampling iteration proceeds, the amount of data used in each sampling evaluation is gradually increased; 12222) Pareto optimal solution search: Find the optimal balance point between different evaluation indicators, as follows: (f) Initialize the search space: determine all adjustable architecture parameters of the hypernetwork and their value ranges, including the number of network layers and channel width; define optimization objectives, including maximizing accuracy and minimizing inference latency; (g) Generate initial population: Randomly select a set of architecture configurations from the search space as the set of initial solutions to ensure population diversity to cover a wider range of configuration space; (h) Evaluate performance: For each architecture configuration, measure its accuracy and inference latency and record the evaluation results; (i) Pareto sorting: Based on the evaluation results, the solutions in the population are sorted by non-dominated levels and each solution is assigned to a different Pareto level. Non-dominated solutions are preferentially retained, i.e., configurations that are not inferior to other solutions in any objective. (j) Select the next generation of solutions: Select solutions with higher Pareto levels from the current population as the basis, and introduce a diversity mechanism to avoid falling into local optimality; on the basis of maintaining solution diversity, streamline the solution set; (k) Generate new solutions: Apply optimization operations to the current solution set to generate a new generation of architectural configurations to ensure the exploratory nature of new solutions and the improvement of existing solutions; (l) Iterative optimization: Repeat steps (h) to (k) to gradually optimize the population; in each round of iteration, dynamically update the Pareto frontier; (m) Output configuration set: After the preset number of iterations is reached or the stopping condition is met, the final optimized configuration set is generated. The set contains the configurations that achieve the best trade-off between accuracy and inference latency, and records the architectural parameters and corresponding performance data of each configuration; 1223) Discrete neural network set evaluation: perform latency and accuracy evaluation on each network configuration in the discrete neural network set, and record the architecture parameters and corresponding performance data of each configuration; 123) Filtering a subset of network configurations; Based on the performance evaluation results of step 122), according to the specific requirements of the actual application scenario, the screening criteria are set to filter the neural network set, and a subset of network configurations that meet the deployment requirements is selected. The screening process is as follows: 1231) screening base constraints; Set hard requirements for neural network performance, including: Minimum accuracy threshold Accuracy min : The network detection accuracy must be higher than the minimum accuracy threshold; Maximum allowed delay max : The network inference delay must be lower than the maximum allowed delay; For any network configuration (α, A, T), it must satisfy the following conditions at the same time: A≥Accuracy min T≤Latency max ; For unsatisfactory network configurations, remove the network configuration subset; 1232) Filter redundant configuration; Grid clustering is performed on the accuracy-delay plane as follows: The accuracy dimension and delay dimension are divided into grids of size δA and δT respectively, where δA is the grid unit interval of accuracy and δT is the grid unit interval of delay; For multiple configurations that fall within the same grid cell, only the one with the best performance is retained; For configurations i and j in a grid cell, if |Ai-Aj|≤δA and |Ti-Tj|≤δT are satisfied, the configuration with higher accuracy is selected and retained; 1233) output configuration set; Output the final subset of neural network configurations that pass the screening. Each configuration contains a single neural network configuration architecture parameter, accuracy, and latency information. The configuration subset will serve as a candidate set for subsequent neural network deployment.
6. The method for configuring an adaptive neural network for video stream target detection according to claim 5, characterized in that: The step 2) specifically includes: 21) Initial neural network configuration; 22) receiving video stream; 23) Perform target detection on the received video stream, determine whether the current neural network configuration needs to be updated with a certain decision cycle, and perform corresponding processing.
7. The method for configuring an adaptive neural network for video stream target detection according to claim 6, characterized in that: The step 21) specifically includes: Select the initial neural network configuration m from the subset M′ in step 1) init , prioritize the neural network configuration that balances accuracy and latency as the initial configuration of the current neural network architecture parameters; for different neural network configuration types, perform different initialization steps as follows: 211) If it is a subnetwork configuration of a weight-sharing supernetwork, then load the supernetwork shared weight Weight share , and set the current active sub-network according to the current neural network architecture parameter configuration. After activating the current active sub-network, update the statistics of the batch normalization layer in the current active sub-network; If the batch normalization parameter in the subnetwork weight is (μ x ,σ x ), where μ x is the mean of the input features, σ x is the standard deviation of the input features, then the calibration dataset To update: The statistics of the batch normalization layer are pre-calibrated and stored as a separate weight file for fast sub-network switching; 212) If it is a discrete weight-independent neural network configuration, directly load the corresponding independent weight Weight discrete .
8. The method for configuring an adaptive neural network for video stream target detection according to claim 7, characterized in that: The step 22) specifically includes: Send a video stream start request containing stream_id to the server, read the video frame through OpenCV, encode the image into JPEG format and convert it into base64 string; construct a frame data packet containing its meta information and image data, and send it to the server through Socket; when the video stream processing is terminated, send a stream stop request and release related resources.
9. The method for configuring an adaptive neural network for video stream target detection according to claim 8, characterized in that: The step 23) specifically includes: 231) Create a fixed-size frame processing queue for caching frames to be processed. The queue adopts a FIFO strategy. When receiving video frame data, decode the base64 string into an image array, record the frame receiving timestamp, and add the frame to be processed and its meta information to the processing queue; when the queue reaches the maximum length, the entry of a new frame will cause the earliest frame to be discarded to ensure the real-time performance of the system; 232) looping to obtain frames to be processed from the queue, for each frame to be processed, calling the currently active target detection neural network to perform calculations to obtain a detection result, the detection result including a detection box, a category, and a confidence score; Detection box: Each detected object corresponds to a rectangular box, which can be represented as b i =(x min ,y min ,x max ,y max ), where (x min ,y min ),(x max ,y max ) are the coordinates of the upper left corner and lower right corner of the rectangular box respectively; Category: The category to which the object detected in each detection frame belongs; Confidence score: Each detection box corresponds to a confidence score c i ∈[0,1], indicating the probability of predicting that the object in the detection box is of a specific category; Perform visual rendering on the original image and record the processing completion timestamp; 233) re-encode the processed frame into base64 format, construct a data packet containing the detection result and timestamp information and output it; 234) Update the performance statistics of the current service and provide query APIs to the outside world, including: Record the detection accuracy based on the mAP@[IoU=0.50:0.95] value of the current network; The processing delay is calculated by the difference between the processing completion timestamp and the processing start timestamp; Count the current queue length as a system load indicator; Calculate the average confidence of the test results, where the average confidence is the average of all confidence scores in the test results; Calculate the average object size of the detection results. The average object size is the average length and width of all detection boxes in the detection results. Calculate the brightness and contrast of the current image; 235) Determine whether the current neural network configuration needs to be updated, and perform corresponding processing, specifically including: 2351)Configuration update status detection: Detect the current configuration update status. If it is detected that the configuration update status needs to be updated, read the target configuration parameters and verify the validity of the configuration update, including the update timestamp and the integrity of the configuration parameters. 2352) Processing according to the test results: When it is detected that the configuration update status is that an update is required, the current timestamp and update information are recorded, and the process returns to step 21) to update the neural network configuration. After the configuration update is completed, the process returns to step 232) to continue processing the video stream. When it is detected that the configuration update state does not need to be updated, the current neural network configuration is maintained, and the process returns to step 232 to continue processing the video stream; 2353)Configure switching status record: Record the configuration-related status information during the execution of step 2352), including the configuration check timestamp, configuration update status, and configuration switch execution status; at the same time, keep tracking the configuration status, record the currently active network configuration information, and maintain the configuration update history record.
10. The method for configuring an adaptive neural network for video stream target detection according to claim 9, characterized in that: The step 3) specifically includes: 31) Update of timing status characteristics; 311) Use a time window structure based on a double-ended queue to cache the system operation status: set the max_window maximum window size to limit the number of historical state records; set the observation window size observation_window to determine the length of the state sequence used for each decision; set the decision time interval decision_interval to control the execution frequency of the neural network configuration decision; 312) Obtain the latest performance statistics, including: detection accuracy accuracy, processing delay latency, current queue length queue_length, average confidence of detection results, average target size average_size of detection results, and brightness and contrast of the current image; add the obtained performance statistics as the current system state to the state cache queue; 313) The collected performance statistics are normalized and each indicator is mapped to a unified numerical range; among them, the accuracy is normalized relative to the maximum accuracy max_accuracy, the delay keeps the original value to reflect the actual processing time, the queue length is normalized relative to the maximum system cache queue length queue_max_length, the confidence is normalized, and the image features are normalized relative to their respective theoretical maximum values; 32) Neural network configuration decision; 321) Define the decision network structure based on LSTM-DQN; 3211) defining state input features, including all state indicators normalized in step 313); 3212) Construct a time series feature extraction layer, use the LSTM network structure to process the continuous state sequence, and extract the time series features and dependencies of the state sequence through the multi-layer network structure; 3213) Construct a decision output layer and use a multi-layer fully connected network to map the time series features to the action value space to evaluate the expected benefits of different configuration actions; 3214) Introducing the dropout mechanism and gradient clipping strategy to improve the stability and generalization ability of network training; 322) Based on the neural network configuration subset M ′ , define the discrete action space: 3221) Establish a mapping relationship from configuration to action, map each candidate neural network configuration to a discrete action, and form an action set A = {a1, a2, ..., a n }, where a i ∈M′, n is the size of the neural network configuration subset |M′|; 3222) constructing an action-to-configuration mapping c=map(a), ensuring that the action a selected from the action space A corresponds to the neural network configuration parameter c; 323) define state representation; Define the state s at time t t is the state sequence of observation_window length, s t ={x t-w+1 ,x t-w+2 ,...,x t }, where w = observation_window, x t-i represents the system state at time ti, and the system state x at each time t Contains all normalized state indicators in step 313); 324) Processing state feature sequence; 3241) Based on the state cache in step 31), obtain the most recent observation_window state samples to form the state s in step 323) t , forming a state feature sequence; 3242) Preprocess the state feature sequence. When the number of state samples is less than observation_window, the sequence is filled by copying the earliest state sample to ensure the integrity of the state feature sequence; 3243) Performing tensor conversion on the processed state feature sequence as input of the decision network; 325) Action selection based on ε-greedy strategy; 3251) Perform random exploration with probability ε, randomly select an action with medium probability from the action space A to explore new possibilities in the environment; 3252) Select the optimal action with probability 1-ε, input the current state sequence into the decision network, obtain the expected value estimates of all actions, and select the action with the largest expected value; 3253) Set the initial exploration probability ε init , minimum exploration probability ε min and the decay rate ε decay ,With the increase of the number of decisions, the probability of random exploration is gradually reduced, and a smooth transition from exploration to utilization is achieved; 33) converting the action selected in step 325) into an actual neural network configuration and performing switching; 331) Action mapping conversion: For the currently selected action a t , get the corresponding configuration parameter c t =map(a t ); 332) configuration update action execution; 3321) Build configuration update status information, including target configuration parameters c t and switch timestamp; 3322) using the information obtained in step 3321) to update the configuration update status in step 235) and perform network configuration switching; 3323) Record the current configuration update action, including the initiation time, target configuration and switching state identification; 333) Switching status monitoring; 3331) receiving the configuration switching progress information returned in step 235); 3332) Tracking the intermediate states during the switching process, including configuration loading state and neural network switching state; 3333) Detect abnormal conditions during the switching process, record error information when an abnormality occurs, and perform alarm and abnormality processing; 334) Switching completion confirmation; 3341) receiving a switching completion notification; 3342) Verify the effectiveness of the new configuration; 3343) Update the current configuration record of the system; 34) Decision-making network optimization mechanism; 341) Reward function design; 3411) Based on the state cache in step 31), obtain the state sequence {x t-decision_interval+1 ,x t-decision_interval+2 ,...,x t }, where decision_interval is the decision time interval; 3412) Define the queue pressure coefficient p queue : p queue =min(1.0,queue_length / queue_max_length); 3413) Design a dynamic weighted reward calculation method: Performance score perf , score perf =w acc accuracy+w conf ·confidence, where w acc and w conf is the weight coefficient; Delay score latency , score latency =exp(-latency); Adaptively adjust weights based on queue pressure: w accuracy =1.0-p queue w latency =p queue Among them, w accuracy is the weight of the accuracy index, w latency is the weight of the delay indicator; 3414) Calculate the comprehensive reward value r t ,as follows: r t =w accuracy ×score perf +w latency ×score latency +r penalty Among them, r penalty As an additional reward adjustment, a penalty is imposed when the queue length exceeds a threshold; 342) Experience replay pool maintenance and update; 3421) Maintain a fixed-capacity experience replay cache D to store decision transition quads (s t ,a t ,r t ,s t+1 ); 3422) When executing action s t After that, calculate the reward r t , store the conversion records in the experience replay cache; 3423) Randomly sample batch data from the experience replay buffer D for training; 343) Time series value estimation: 3431) Define the discount factor γ to balance immediate rewards and long-term benefits; 3432) Calculate the weighted sum of the current immediate reward and the future expected return as follows: Q target =r t +γ×max{Q(s t+1 ,a)|a∈A} Among them, Q tar get Represents the current instant reward r t The weighted sum of the expected future returns is used to guide the target value of strategy optimization; A is the action space, max{Q(s t+1 ,a)|a∈A} represents the next state s t+1 The maximum expected return among all possible actions; 3433) Calculate the value of the decision network for the current state and action combination as follows: Q current =Q(s t ,a t ) Among them, Q current Represents the decision network for the current state s t and action a t The combined value estimate reflects the current policy's evaluation of the state-action pair; 3434) Define the timing difference error δt as follows: δt=Q target -Q current ; 344) Gradient optimization strategy; 3441) The loss function is calculated based on the time difference error δt as follows: L(θ)=E[(Q target -Q current ) 2 ] Among them, θ is the parameter of the decision network, and the loss function L(θ) represents the deviation between the predicted value of the decision network and the target value; 3442) The gradient descent method is used to update the decision network parameters as follows: Among them, α is the learning rate.