Intelligent collaborative management method and system for network security operation and maintenance work

By adopting multi-objective optimization model, deep reinforcement learning and Pareto optimal solution algorithm on the network security operation and maintenance management platform, an intelligent collaborative management method is solved, and the multi-objective optimization imbalance and insufficient feature representation capabilities of work order distribution and task scheduling in the existing technology is solved, and efficient and intelligent operation and maintenance management is achieved.

CN120218539AActive Publication Date: 2025-06-27BEIJING YUHONG XINAN TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510349281.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-27
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing network security operation and maintenance management platform has problems such as multi-objective optimization imbalance, insufficient representation capabilities of operation and maintenance personnel and work order feature, and inability to dynamically optimize static scheduling strategies in terms of work order distribution and task scheduling.

Method used

A multi-objective optimization model is used to combine deep reinforcement learning and Pareto optimal solution algorithm to build an intelligent collaborative management method. By obtaining characteristic information of work tickets and operation and maintenance personnel, using multi-layer perceptrons and dual Q networks for state coding and strategy training, dynamically adjust the work ticket distribution plan and monitor the task status in real time to generate a collaborative scheduling strategy.

Benefits of technology

It has achieved multi-objective dynamic optimization of work order distribution, in-depth representation of operation and maintenance characteristics and real-time coordination of task scheduling, which has significantly improved the efficiency and intelligence level of network security operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218539A_ABST
    Figure CN120218539A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent collaborative management method and system for network security operation and maintenance work, and relates to the technical field of network security, and the method comprises the steps: obtaining operation and maintenance work order information containing a work order priority and a processing time limit, and converting the operation and maintenance work order information into a work order feature vector; inputting the skill score, the historical completion rate and the workload of the operation and maintenance personnel into a multi-objective optimization model, and calculating a work order matching degree score and an emergency degree score; performing state coding on the work order feature vector, the personnel feature and the constraint condition by adopting a multi-layer perceptron, extracting spatial correlation by utilizing a double Q network, and calculating a target Q value; and generating a work order distribution scheme based on a Pareto optimal solution algorithm, monitoring a processing state in real time through a task collaborative scheduling model, and generating a collaborative scheduling strategy. The intelligent level of work order distribution is improved, collaborative management of operation and maintenance tasks is realized, and the operation and maintenance efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to network security technologies, and particularly to an intelligent collaborative management method and system for network security operation and maintenance work. Background Art

[0002] With the increasing complexity of network security threats, network security operation and maintenance work faces challenges such as a sharp increase in the number of work orders, higher requirements for processing timeliness, and complex task allocation. Network security operation and maintenance work requires intelligent scheduling and collaborative management based on multi-dimensional factors such as work order priority, processing time limit, and skills of operation and maintenance personnel. Currently, mainstream operation and maintenance management platforms have begun to use artificial intelligence technologies for work order distribution and task collaboration, and achieve reasonable scheduling of operation and maintenance resources through multi-objective optimization algorithms.

[0003] However, the existing technologies still have the following problems: First, traditional work order distribution methods often only consider single-objective optimization and cannot effectively balance multiple objectives such as work order matching degree and urgency, resulting in unreasonable resource allocation; second, the existing task scheduling models lack the ability to deeply represent the characteristics of operation and maintenance personnel and work orders, and it is difficult to accurately depict the complex relationships between multi-dimensional characteristics; third, the current collaborative scheduling strategies generally adopt static rules and cannot be dynamically optimized and adjusted according to the real-time work order processing status, affecting operation and maintenance efficiency.

[0004] The technical problem to be solved by the present invention is: how to construct an intelligent collaborative management method for network security operation and maintenance based on multi-objective optimization and deep reinforcement learning to achieve multi-objective dynamic optimization of work order distribution, deep representation of operation and maintenance characteristics, and real-time collaboration of task scheduling. Summary of the Invention

[0005] Embodiments of the present invention provide an intelligent collaborative management method and system for network security operation and maintenance work, which can solve the problems in the existing technologies.

[0006] In the first aspect of the embodiments of the present invention, an intelligent collaborative management method for network security operation and maintenance work is provided, including: Obtain network security operation and maintenance work order information, where the network security operation and maintenance work order information includes work order priority and work order processing time limit, convert the network security operation and maintenance work order information into a work order feature vector, and input the work order feature vector into a multi-objective optimization model trained based on historical work order processing data; The multi-objective optimization model calculates a work order matching degree score based on the skill score of operation and maintenance personnel, historical work order completion rate, and current workload, and calculates a work order urgency score based on the work order priority and work order processing time limit; input the work order matching degree score and the work order urgency score into the Pareto optimal solution algorithm to generate a multi-objective optimized work order distribution plan; Use a multi-layer perceptron to perform state encoding on the work order feature vector, operation and maintenance personnel features, and constraint conditions. Utilize a dual Q-network structure to extract the spatial correlation of the state encoding and calculate the target Q-value. Train the parameters of the dual Q-network based on a multi-objective reward function. Combine the Pareto dominance relationship and the reference point method to select the optimal solution, and continuously optimize the work order distribution plan through a dynamic adjustment mechanism; Determine the target operation and maintenance personnel according to the work order distribution plan, and issue work order processing tasks to the target operation and maintenance personnel; Use the Pareto optimal solution algorithm to construct a task collaborative scheduling model, and the task collaborative scheduling model monitors the work order processing status of the target operation and maintenance personnel in real time; When the work order processing status is abnormal, the task collaborative scheduling model generates a task collaborative scheduling strategy based on the work order processing status and the work order processing task.

[0007] The multi-objective optimization model calculates the work order matching degree score based on the skill score, historical work order completion rate, and current workload of the operation and maintenance personnel, and calculates the work order urgency score based on the work order priority and work order processing time limit; Input the work order matching degree score and the work order urgency score into the Pareto optimal solution algorithm to generate a multi-objective optimized work order distribution plan, including: Obtain the skill score, historical work order completion rate, and current workload of the operation and maintenance personnel, and construct a skill score vector for the skill score according to the skill dimension; Calculate the historical work order completion rate based on the time decay weight, and the time decay weight decays as the statistical time interval increases; Calculate the current workload according to the ratio of the remaining workload of the task to the deadline; Calculate the priority score according to the work order priority and the urgency coefficient, and calculate the time limit score based on the difference between the work order processing time limit and the current time; Input the skill score vector, historical work order completion rate, and current workload into the feature fusion layer, and perform weighted fusion through the feature weight matrix to obtain the fusion feature; Calculate the work order matching degree score based on the fusion feature, and the work order matching degree score is obtained through the inner product operation of the fusion feature and the work order feature vector; Obtain the work order urgency score by weighting the priority score and the time limit score through the weight coefficient, and the weight coefficient is optimized and determined based on historical distribution data; Construct a multi-objective optimization function, which includes a work order matching degree objective function and a work order urgency objective function. The work order matching degree objective function is constructed based on the work order matching degree score, and the work order urgency objective function is constructed based on the work order urgency score; Input the multi-objective optimization function into the Pareto optimal solution algorithm, and screen the non-dominated solution set based on the Pareto dominance relationship. The solutions in the non-dominated solution set all satisfy the processing capacity constraint; Use an adaptive weight adjustment mechanism to dynamically optimize the feature weight matrix and the weight coefficient, and implement feedback compensation control based on the performance deviation to generate the optimal work order distribution plan.

[0008] Input the multi-objective optimization function into the Pareto optimal solution algorithm, screen the non-dominated solution set based on the Pareto dominance relationship, and the solutions in the non-dominated solution set all satisfy the processing capacity constraint; adopt an adaptive weight adjustment mechanism to dynamically optimize the feature weight matrix and weight coefficient, and implement feedback compensation control based on the performance deviation to generate the optimal work order distribution plan, including: The multi-objective optimization function includes a first objective function calculated based on the work order matching degree score and a second objective function calculated based on the work order urgency score, and impose the constraints of the upper limit of the processing capacity of the operation and maintenance personnel and the unique allocation of work orders on the multi-objective optimization function; Input the multi-objective optimization function into the Pareto optimal solution algorithm, and the Pareto optimal solution algorithm performs multi-dimensional space mapping on the first objective function and the second objective function, and determines and screens the non-dominated solution set that meets the constraints based on the Pareto dominance relationship. Each solution in the non-dominated solution set corresponds to a candidate work order distribution plan; Construct an adaptive weight adjustment mechanism. The adaptive weight adjustment mechanism performs performance evaluation on the candidate work order distribution plans in the non-dominated solution set to obtain the performance deviation, and implements feedback compensation using a proportional-integral controller based on the performance deviation, dynamically adjusts the feature fusion weight matrix and the objective function weight coefficient, and selects the optimal work order distribution plan from the non-dominated solution set.

[0009] Use the double Q-network structure to extract the spatial correlation of the state encoding and calculate the target Q value, train the parameters of the double Q-network based on the multi-objective reward function, select the optimal plan by combining the Pareto dominance relationship and the reference point method, and continuously optimize the work order distribution plan through the dynamic adjustment mechanism, including: Input the work order feature vector, the operation and maintenance personnel features, and the system constraint condition matrix into the multi-layer perceptron, and obtain the state feature vector through state encoding by the multi-layer perceptron; Construct a double Q-network structure. The double Q-network structure includes a current network and a target network. Input the state feature vector into the current network, extract the spatial correlation of the state features through the convolutional neural network, input the extracted spatial correlation into the target network and calculate the target Q value; Construct a multi-objective reward function based on the target Q value. The multi-objective reward function includes a matching degree reward, an urgency reward, and a constraint satisfaction reward. The matching degree reward is obtained by calculating the cosine similarity between the work order skill requirements and the operation and maintenance personnel skill scores. The urgency reward is calculated based on the remaining processing time of the work order. The constraint satisfaction reward is calculated based on the degree of satisfaction of the constraint conditions; Adopt an experience replay mechanism to store state transition samples, select training samples based on priority sampling, input the training samples into the double Q-network structure for training, and update the network parameters of the double Q-network structure by minimizing the temporal difference error; Substitute the network parameters of the updated double Q - network structure into the ε - greedy strategy, explore in the action space to generate a set of candidate work order distribution schemes, and use the Pareto dominance relationship to screen the set of candidate work order distribution schemes to obtain the non - dominated solution set; construct a reference - point method model based on the non - dominated solution set. The reference - point method model calculates the utility function value based on the ideal - point coordinates and target weights, and selects the work order distribution scheme with the largest utility function value as the optimal work order distribution scheme; Monitor the performance of the optimal work order distribution scheme. When the change rate of the performance index exceeds the preset change threshold, trigger the dynamic adjustment mechanism, update the target weights based on the gradient - descent method, and feedback the updated target weights to the multi - objective reward function to continuously optimize the optimal work order distribution scheme.

[0010] Using the Pareto dominance relationship to screen the set of candidate work order distribution schemes to obtain the non - dominated solution set; constructing a reference - point method model based on the non - dominated solution set, and the reference - point method model calculates the utility function value based on the ideal - point coordinates and target weights, and selecting the work order distribution scheme with the largest utility function value as the optimal work order distribution scheme includes: Calculate the dominance degree and the dominated degree of each scheme in the set of candidate work order distribution schemes based on the Pareto dominance relationship, where the dominance degree represents the number of other schemes dominated by the corresponding scheme, and the dominated degree represents the number of schemes that dominate the corresponding scheme; Stratify and screen the set of candidate work order distribution schemes according to the dominance degree and the dominated degree, and divide the schemes with a dominated degree of zero into multiple layers of non - dominated solution sets; Select the first - layer non - dominated solution set in the multiple - layer non - dominated solution sets to construct a reference - point method model, and determine the ideal - point coordinates based on the function values corresponding to the multi - objective reward functions of each scheme in the multiple - layer non - dominated solution sets; Calculate the normalized objective values of each scheme in the multiple - layer non - dominated solution sets relative to the ideal - point coordinates, and calculate the utility function value in combination with the preset target weights, and select the work order distribution scheme with the largest utility function value as the optimal work order distribution scheme.

[0011] Use the Pareto - optimal solution algorithm to construct a task collaborative scheduling model. The task collaborative scheduling model monitors the work order processing status of the target operation and maintenance personnel in real - time; when the work order processing status is abnormal, the task collaborative scheduling model generates a task collaborative scheduling strategy according to the work order processing status and the work order processing tasks, including: Use the Pareto - optimal solution algorithm to construct a task collaborative scheduling model. The task collaborative scheduling model obtains the work order processing status of the target operation and maintenance personnel. The work order processing status includes the work order processing duration, the work order processing quality, and the resource utilization rate. Convert the work order processing status into a state feature matrix, and the state feature matrix generates a historical state encoding through a bidirectional gated recurrent unit network; Obtain time-series state feature data based on historical state encoding and perform multi-scale decomposition. Dynamically adjust the learning rate according to the state change rate, calculate dynamic statistics by combining exponential moving average and median statistics, and perform weighted fusion on the dynamic statistics through reliability evaluation and impose a smoothing constraint to obtain the final statistical result; Calculate the dynamic mean and dynamic standard deviation of the state feature matrix using the final statistical result, and standardize the state feature matrix through the dynamic mean and dynamic standard deviation to obtain a standardized state vector; construct a multi-dimensional reward and punishment function, and the multi-dimensional reward and punishment function calculates the reward and punishment values based on the deviation of the work order processing duration, the difference in processing quality, and the resource utilization rate distance. When the standardized state vector exceeds the reward and punishment value, it is determined that the work order processing state is abnormal; When it is detected that the work order processing state is abnormal, obtain the real-time work order processing progress and the information of deployable operation and maintenance resources, and generate candidate scheduling strategies based on the real-time work order processing progress and the information of deployable operation and maintenance resources; input the candidate scheduling strategies into the task collaborative scheduling model, calculate the evaluation scores of each strategy using the multi-dimensional reward and punishment function, and select the non-dominated solutions as the task collaborative scheduling strategies through the Pareto optimal solution algorithm.

[0012] Obtain time-series state feature data based on historical state encoding and perform multi-scale decomposition. Dynamically adjust the learning rate according to the state change rate, calculate dynamic statistics by combining exponential moving average and median statistics, and perform weighted fusion on the dynamic statistics through reliability evaluation and impose a smoothing constraint to obtain the final statistical result including: Extract time-series state feature data based on historical state encoding, and perform wavelet transform on the time-series state feature data to obtain multi-scale decomposition coefficients; calculate the difference between the time-series state feature data at adjacent times to obtain the instantaneous change rate, and perform exponential smoothing on the instantaneous change rate to obtain the cumulative change rate; Generate an adaptive learning rate through mapping by the Sigmoid function based on the cumulative change rate, and the value range of the adaptive learning rate is limited by the preset upper bound and lower bound of the learning rate; perform exponential moving average operation on the time-series state feature data using the adaptive learning rate to obtain the dynamic mean, and calculate the dynamic standard deviation of the time-series state feature data based on the dynamic mean; Calculate the median of the time-series state feature data within a sliding time window of a preset length to obtain the robust mean, and calculate the median of the absolute deviation of the time-series state feature data relative to the robust mean to obtain the robust standard deviation; calculate the relative deviation between the dynamic mean and the robust mean, and convert the relative deviation through an exponential function to obtain the reliability evaluation score; Calculate the fusion weights based on the reliability evaluation scores, and use the fusion weights to perform weighted fusion on the dynamic mean and the robust mean to obtain the final mean, and perform weighted fusion on the dynamic standard deviation and the robust standard deviation to obtain the final standard deviation; impose a temporal smoothing constraint on the final mean to limit the change amplitude at adjacent times, and impose a value range constraint on the final standard deviation to ensure its stability; use the final mean and the final standard deviation as the final statistical results.

[0013] In the second aspect of the embodiments of the present invention, An intelligent collaborative management system for network security operation and maintenance work is provided, including: A first unit, configured to obtain network security operation and maintenance work order information, where the network security operation and maintenance work order information includes work order priority and work order processing time limit, convert the network security operation and maintenance work order information into a work order feature vector, and input the work order feature vector into a multi-objective optimization model trained based on historical work order processing data; A second unit, configured to calculate the work order matching degree score by the multi-objective optimization model according to the skill scores of operation and maintenance personnel, the historical work order completion rate, and the current workload, and calculate the work order urgency score based on the work order priority and the work order processing time limit; input the work order matching degree score and the work order urgency score into the Pareto optimal solution algorithm to generate a work order distribution plan for multi-objective optimization; A third unit, configured to perform state encoding on the work order feature vector, the operation and maintenance personnel features, and the constraint conditions by using a multi-layer perceptron, extract the spatial correlation of the state encoding by using a double Q network structure and calculate the target Q value, train the parameters of the double Q network based on a multi-objective reward function, select the optimal solution by combining the Pareto dominance relationship and the reference point method, and continuously optimize the work order distribution plan through a dynamic adjustment mechanism; A fourth unit, configured to determine the target operation and maintenance personnel according to the work order distribution plan, and send the work order processing task to the target operation and maintenance personnel; construct a task collaborative scheduling model by using the Pareto optimal solution algorithm, and the task collaborative scheduling model monitors the work order processing status of the target operation and maintenance personnel in real time; when the work order processing status is abnormal, the task collaborative scheduling model generates a task collaborative scheduling strategy according to the work order processing status and the work order processing task.

[0014] In the third aspect of the embodiments of the present invention, An electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0015] In the fourth aspect of the embodiments of the present invention, Provided is a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.

[0016] The beneficial effects of the present application are as follows: 1. By constructing a multi-objective optimization model, the present invention simultaneously considers two optimization objectives of work order matching degree and urgency, fuses and models multi-dimensional features such as the skill scores, historical completion rates, and workloads of operation and maintenance personnel, and uses the Pareto optimal solution algorithm for multi-objective trade-off, effectively solving the problem that it is difficult for traditional single-objective optimization methods to balance multi-dimensional resource allocation, and significantly improving the rationality of work order distribution and resource utilization efficiency.

[0017] 2. The present invention adopts a deep learning architecture combining a multi-layer perceptron and a double Q network to perform deep state encoding and spatial correlation extraction on work order feature vectors, operation and maintenance personnel features, and constraint conditions, and trains network parameters through a multi-objective reward function, realizing accurate modeling and feature representation of complex operation and maintenance scenarios, greatly improving the accuracy and robustness of work order distribution decisions, and the work order distribution accuracy rate is increased by more than 30% compared with traditional methods.

[0018] 3. The present invention innovatively applies the Pareto optimal solution algorithm to the construction of a task collaborative scheduling model, realizes real-time monitoring of work order processing status and rapid response to abnormal status, dynamically generates task collaborative scheduling strategies, and solves the problem of poor adaptability of traditional static scheduling schemes. Through continuous optimization of the dynamic adjustment mechanism, the overall operation and maintenance efficiency of the system is increased by 40%, and the abnormal handling response time is shortened by 50%, significantly enhancing the intelligent collaborative management ability of network security operation and maintenance work. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic flowchart of the intelligent collaborative management method for network security operation and maintenance work in an embodiment of the present invention; Figure 2 It is a heat map of processing time under different work order quantities and urgency levels in an embodiment of the present invention; Figure 3 It is a scatter plot comparison table diagram of the matching degree of the work order distribution scheme in an embodiment of the present invention; Figure 4 It is a performance improvement table diagram of the present technical solution relative to other algorithms in an embodiment of the present invention; Figure 5 It is a schematic diagram of a multi-scale time series feature adaptive statistical analysis system in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0022] Figure 1 It is a schematic flowchart of the intelligent collaborative management method for network security operation and maintenance work in the embodiments of the present invention. As Figure 1 shown, the method includes: Obtain network security operation and maintenance work order information, where the network security operation and maintenance work order information includes work order priority and work order processing time limit. Convert the network security operation and maintenance work order information into a work order feature vector, and input the work order feature vector into a multi-objective optimization model trained based on historical work order processing data; The multi-objective optimization model calculates the work order matching degree score based on the skill score of the operation and maintenance personnel, the historical work order completion rate, and the current workload, and calculates the work order urgency score based on the work order priority and work order processing time limit; Input the work order matching degree score and the work order urgency score into the Pareto optimal solution algorithm to generate a multi-objective optimized work order distribution plan; Use a multi-layer perceptron to perform state encoding on the work order feature vector, operation and maintenance personnel features, and constraint conditions, use a dual Q-network structure to extract the spatial correlation of the state encoding and calculate the target Q value, train the parameters of the dual Q-network based on a multi-objective reward function, select the optimal plan by combining the Pareto dominance relationship and the reference point method, and continuously optimize the work order distribution plan through a dynamic adjustment mechanism; Determine the target operation and maintenance personnel according to the work order distribution plan, and issue the work order processing task to the target operation and maintenance personnel; Use the Pareto optimal solution algorithm to construct a task collaborative scheduling model, and the task collaborative scheduling model monitors the work order processing status of the target operation and maintenance personnel in real time; When the work order processing status is abnormal, the task collaborative scheduling model generates a task collaborative scheduling strategy based on the work order processing status and the work order processing task.

[0023] In an alternative embodiment, the multi-objective optimization model calculates the work order matching degree score based on the skill score of the operation and maintenance personnel, the historical work order completion rate, and the current workload, and calculates the work order urgency score based on the work order priority and work order processing time limit; Input the work order matching degree score and the work order urgency score into the Pareto optimal solution algorithm to generate a multi-objective optimized work order distribution plan, including: Obtain the skill scores, historical work order completion rates, and current workloads of operation and maintenance personnel. Construct a skill score vector for the skill scores according to skill dimensions. Calculate the historical work order completion rate based on time-decaying weights, where the time-decaying weights decay as the statistical time interval increases. Calculate the current workload according to the ratio of the remaining workload of the task to the deadline. Calculate the priority score according to the work order priority and the urgency coefficient, and calculate the time limit score based on the difference between the work order processing time limit and the current time. Input the skill score vector, historical work order completion rate, and current workload into the feature fusion layer, and perform weighted fusion through the feature weight matrix to obtain the fused features. Calculate the work order matching degree score based on the fused features, where the work order matching degree score is obtained through the inner product operation of the fused features and the work order feature vector. Obtain the work order urgency score by weighting the priority score and the time limit score through the weight coefficient, and the weight coefficient is optimized and determined based on historical distribution data. Construct a multi-objective optimization function, which includes a work order matching degree objective function and a work order urgency objective function. The work order matching degree objective function is constructed based on the work order matching degree score, and the work order urgency objective function is constructed based on the work order urgency score. Input the multi-objective optimization function into the Pareto optimal solution algorithm, and screen the non-dominated solution set based on the Pareto dominance relationship. The solutions in the non-dominated solution set all satisfy the processing capacity constraint. Adopt an adaptive weight adjustment mechanism to dynamically optimize the feature weight matrix and the weight coefficient, and implement feedback compensation control based on the performance deviation to generate the optimal work order distribution plan.

[0024] Obtain the basic feature information of operation and maintenance personnel. For the skill scores, construct a skill score vector according to different skill dimensions such as network security, system maintenance, and vulnerability repair. For the calculation of the historical work order completion rate, adopt a weighted statistical method based on time decay, set a benchmark time window, and the weight of the historical completion data farther away from the current time is smaller. Determine the time-decaying weight through an exponential decay function. For the evaluation of the current workload, count the remaining workload of all unfinished work orders of the operation and maintenance personnel, and calculate the standardized workload value in combination with the deadline of each work order.

[0025] The calculation of the work order urgency score includes two dimensions. The priority score is based on the priority level of the work order, and a corresponding urgency coefficient is set in combination with the business scenario for weighted calculation. The time limit score is obtained by normalizing the time difference between the work order processing time limit and the current time. These two scores are weighted and fused through a dynamic weight coefficient, and the initial value of the weight coefficient is determined by analyzing the time limit compliance rate and priority satisfaction degree in the historical work order distribution data.

[0026] The calculation of the work order matching degree score adopts a feature fusion method. The skill score vector, historical work order completion rate, and current workload are input into the feature fusion layer, which contains a trainable feature weight matrix. Different features are weighted according to their importance through this matrix. The inner product of the fused feature vector and the work order feature vector is calculated to obtain the final matching degree score.

[0027] In the multi-objective optimization stage, two objective functions of work order matching degree and work order urgency are constructed. The work order matching degree objective function aims to maximize the matching degree score, and the work order urgency objective function aims to minimize the urgency loss. These two objective functions form a multi-objective optimization problem, and at the same time, the upper limit constraint of the operation and maintenance personnel's processing capacity is imposed.

[0028] The Pareto optimal solution algorithm is used to solve the multi-objective optimization problem. The algorithm first generates an initial solution set, determines the Pareto dominance relationship for each solution in the solution set, and filters out the non-dominated solution set. To further optimize the quality of the solution, an adaptive weight adjustment mechanism is introduced. This mechanism obtains the performance deviation by monitoring the work order distribution effect, and dynamically adjusts the feature weight matrix and the objective function weight coefficient according to the deviation value to achieve feedback compensation control. Finally, the solution with the optimal comprehensive performance is selected from the optimized non-dominated solution set as the final work order distribution plan.

[0029] In an alternative implementation, the multi-objective optimization function is input into the Pareto optimal solution algorithm, and the non-dominated solution set is filtered based on the Pareto dominance relationship. The solutions in the non-dominated solution set all satisfy the processing capacity constraint; the adaptive weight adjustment mechanism is used to dynamically optimize the feature weight matrix and the weight coefficient, and feedback compensation control is implemented based on the performance deviation to generate the optimal work order distribution plan, including: The multi-objective optimization function includes a first objective function based on the calculation of the work order matching degree score and a second objective function based on the calculation of the work order urgency score. Constraints on the upper limit of the operation and maintenance personnel's processing capacity and the unique allocation of work orders are imposed on the multi-objective optimization function; The multi-objective optimization function is input into the Pareto optimal solution algorithm. The Pareto optimal solution algorithm performs multi-dimensional space mapping on the first objective function and the second objective function, and filters out the non-dominated solution set that satisfies the constraint conditions based on the determination of the Pareto dominance relationship. Each solution in the non-dominated solution set corresponds to a candidate work order distribution plan; An adaptive weight adjustment mechanism is constructed. The adaptive weight adjustment mechanism performs performance evaluation on the candidate work order distribution plans in the non-dominated solution set to obtain the performance deviation, and implements feedback compensation using a proportional-integral controller based on the performance deviation, dynamically adjusting the feature fusion weight matrix and the objective function weight coefficient, and selecting the optimal work order distribution plan from the non-dominated solution set.

[0030] Construct a multi-objective optimization function, which includes two sub-objective functions: the first objective function is calculated based on the work order matching score, and the second objective function is calculated based on the work order urgency score. The first objective function is expressed as the sum of the matching scores of all work orders and the assigned operation and maintenance personnel, and the goal is to maximize this sum; the second objective function represents the weighted sum of the urgency scores of all work orders and the expected start time of processing, and the goal is to minimize this weighted sum.

[0031] There are 3 work orders and 2 operation and maintenance personnel in the system. The matching score matrix between work orders and operation and maintenance personnel is: the matching scores of work order 1 with personnel A / B are 85 / 70 respectively, the matching scores of work order 2 with personnel A / B are 65 / 90 respectively, and the matching scores of work order 3 with personnel A / B are 75 / 80 respectively. If the assignment plan is: work order 1 is assigned to personnel A, work order 2 is assigned to personnel B, and work order 3 is assigned to personnel A, then the value of the first objective function is 85 + 90 + 75 = 250.

[0032] Suppose the urgency scores of work orders 1, 2, and 3 are 90, 85, and 95 respectively. Personnel A has no current work order, and personnel B has a work order expected to be completed in 1 hour. Under the above assignment plan, the processing start times of work orders 1, 2, and 3 are 0, 1, and 2 hours respectively. Then the value of the second objective function is 90×0 + 85×1 + 95×2 = 275.

[0033] When constructing the multi-objective optimization function, two types of constraint conditions are imposed: the upper limit constraint of the processing capacity of operation and maintenance personnel and the unique assignment constraint of work orders. If the upper limits of the processing capacities of personnel A and B are 10 and 8 work units respectively, and the workloads of work orders 1, 2, and 3 are 3, 4, and 5 work units respectively, then the above plan meets the constraint conditions.

[0034] Input the constructed multi-objective optimization function into the Pareto optimal solution algorithm. This algorithm first randomly generates an initial population, and each individual represents a work order distribution plan. Calculate the values of the two objective functions for each individual in the population, and perform non-dominated sorting based on the Pareto dominance relationship. The Pareto dominance relationship is defined as: if plan A is better than plan B in at least one objective function and not worse than plan B in another objective function, then plan A dominates plan B.

[0035] For example, the objective function values of three plans are: plan 1(250, -275), plan 2(240, -260), plan 3(260, -290). After comparison, plan 1 and plan 2 belong to the first level (non-dominating each other), and plan 3 belongs to the second level (dominated by plan 2).

[0036] Within the same non-dominated level, sort based on the crowding distance. Generate a new population through selection, crossover, and mutation operations, and iterate and optimize until convergence. Finally, the algorithm outputs all individuals in the first non-dominated level, forming a non-dominated solution set.

[0037] To select the final solution, an adaptive weight adjustment mechanism is constructed. This mechanism first performs a performance evaluation on the candidate solutions, calculating the deviations between indicators such as the timely processing rate of work orders and the workload balance degree and the target values. For example, if the target timely processing rate is 95% and a certain solution is predicted to be 92%, then the deviation is -3%.

[0038] Based on the performance deviation, a proportional-integral controller is used to implement feedback compensation, dynamically adjusting the feature fusion weight matrix and the target function weight coefficients. If the timely processing rate is lower than the target, the weight of urgency is increased; if the workload distribution is uneven, the allocation strategy is adjusted. For example, for a deviation of -3%, the proportional coefficient is 0.1, the integral coefficient is 0.05, and the historical cumulative deviation is -10%, then the weight of urgency should be increased by 0.1×3% + 0.05×10% = 0.8%.

[0039] Based on the adjusted weights, the candidate solutions are re-evaluated, and the solution with the optimal comprehensive performance is selected. If the weights of the first / second target functions are 0.55 / 0.45, then the comprehensive scores of solutions 1, 2, and 3 are 13.75, 15, and 13 respectively, and the solution 2 with the highest score is selected as the final work order distribution solution.

[0040] Figure 2 The following is the heat map of the processing time under different work order quantities and urgency levels in the embodiments of the present invention: The left figure shows the performance data of the present technical solution: when the system scale ranges from "very low" to "very high" and the work order quantity ranges from 100 to 500, the performance indicators show a gradient upward trend. The specific data shows that the performance value is 8.3 at the lowest configuration (100 work orders, very low scale), and it increases to 13.4 as the work order quantity increases to 500; when the system scale is increased to "very high", the performance value ranges from 12.4 (100 work orders) to 23.8 (500 work orders). The data shows a relatively gentle growth trend, indicating that the present solution has good scalability.

[0041] The right figure shows the performance data of the traditional method: under the same conditions, the performance indicators are generally higher than those of the present technical solution but fluctuate more. When the system scale is "very low", the performance value increases from 15.1 (100 work orders) to 22.3 (500 work orders); when the system scale reaches "very high", the performance value increases sharply from 22.6 (100 work orders) to 54.8 (500 work orders). The data shows that the traditional method has a significant increase in performance consumption under high load conditions and has poor scalability.

[0042] By comparison, although the present technical solution is lower than the traditional method in terms of absolute performance value, it has better performance stability and scalability, especially with obvious advantages in high load scenarios.

[0043] In an alternative implementation, a dual Q-network structure is used to extract the spatial correlation of the state encoding and calculate the target Q-value. The parameters of the dual Q-network are trained based on a multi-objective reward function. The optimal solution is selected by combining the Pareto dominance relationship and the reference point method. The continuous optimization of the work order distribution scheme through a dynamic adjustment mechanism includes: Input the work order feature vector, the characteristics of the operation and maintenance personnel, and the system constraint condition matrix into a multi-layer perceptron, and perform state encoding through the multi-layer perceptron to obtain a state feature vector; Construct a dual Q-network structure, which includes a current network and a target network. Input the state feature vector into the current network, extract the spatial correlation of the state features through a convolutional neural network, and input the extracted spatial correlation into the target network to calculate the target Q-value; Construct a multi-objective reward function based on the target Q-value. The multi-objective reward function includes a matching degree reward, an urgency reward, and a constraint satisfaction reward. The matching degree reward is obtained by calculating the cosine similarity between the skill requirements of the work order and the skill scores of the operation and maintenance personnel. The urgency reward is calculated based on the remaining processing time of the work order. The constraint satisfaction reward is calculated based on the degree of satisfaction of the constraint conditions; Adopt an experience replay mechanism to store state transition samples, select training samples based on priority sampling, input the training samples into the dual Q-network structure for training, and update the network parameters of the dual Q-network structure by minimizing the temporal difference error; Substitute the network parameters of the updated dual Q-network structure into the ε-greedy strategy, explore in the action space to generate a set of candidate work order distribution schemes, and use the Pareto dominance relationship to screen the set of candidate work order distribution schemes to obtain a non-dominated solution set. Construct a reference point method model based on the non-dominated solution set. The reference point method model calculates the utility function value based on the ideal point coordinates and the target weights, and selects the work order distribution scheme with the largest utility function value as the optimal work order distribution scheme; Monitor the performance of the optimal work order distribution scheme. When the change rate of the performance index exceeds the preset change threshold, trigger the dynamic adjustment mechanism, update the target weights based on the gradient descent method, and feedback the updated target weights to the multi-objective reward function to continuously optimize the optimal work order distribution scheme.

[0044] The system obtains the work order feature vector, the characteristics of the operation and maintenance personnel, and the system constraint condition matrix as input data. The work order feature vector includes features such as work order type, skill requirements, urgency, and estimated processing time. The characteristics of the operation and maintenance personnel include skill scores, historical completion rates, current workloads, etc. The system constraint condition matrix includes processing capacity constraints, time window constraints, etc. These input data are passed into a multi-layer perceptron for state encoding to obtain a state feature vector. The multi-layer perceptron consists of an input layer, multiple hidden layers, and an output layer, and uses the ReLU activation function to enhance the nonlinear expression ability of the network.

[0045] The dimension of the work order feature vector is 10, the dimension of the operation and maintenance personnel feature is 8, the dimension of the system constraint condition matrix is 6×4, and the structure of the multi-layer perceptron is [10 + 8 + 24, 32, 64], that is, the number of neurons in the input layer is 42 (10 + 8 + 24), the number of neurons in the first hidden layer is 32, the number of neurons in the second hidden layer is 64, and the dimension of the output layer, that is, the state feature vector, is 64.

[0046] Construct a double Q-network structure, including a current network and a target network. The two networks have the same network structure but different parameter update frequencies. The introduction of the double Q-network can effectively solve the overestimation problem in Q-learning and improve the learning stability. Input the state feature vector into the current network, and extract the spatial correlation of the state features through a convolutional neural network. The convolutional neural network contains multiple convolutional layers, pooling layers, and fully connected layers, and can capture the local correlation between features.

[0047] The specific structure of the convolutional neural network is as follows: the first convolutional layer uses 32 3×3 convolutional kernels with a stride of 1; the second convolutional layer uses 64 3×3 convolutional kernels with a stride of 1; a max-pooling layer with a pooling kernel size of 2×2 is connected after each convolutional layer; then there are two fully connected layers with the number of neurons being 128 and 64 respectively. The number of neurons in the final output layer is equal to the size of the action space, that is, the number of optional work order distribution schemes. The spatial correlation features extracted by the convolutional neural network are input into the target network to calculate the target Q-values of each possible action (work order distribution scheme).

[0048] Construct a multi-objective reward function based on the target Q-values, including three parts: matching degree reward, urgency reward, and constraint satisfaction reward. The matching degree reward is obtained by calculating the cosine similarity between the work order skill requirements and the operation and maintenance personnel's skill scores, and measures the rationality of work order allocation. For example, if the work order requirement skill vector is [0.8, 0.3, 0.5, 0.2] and the operation and maintenance personnel's skill score vector is [0.7, 0.4, 0.6, 0.3], then the matching degree reward is the cosine similarity of the two vectors, and the calculation result is approximately 0.96, indicating a high matching degree.

[0049] The urgency reward is calculated based on the remaining processing time of the work order. The shorter the remaining time, the higher the reward, which prompts the system to process urgent work orders first. For example, set the basic urgency reward to 10. When the remaining processing time of the work order is less than 2 hours, the urgency reward is 10; when the remaining time is between 2 and 24 hours, the urgency reward is 10 multiplied by (26 minus the remaining hours) divided by 24; when the remaining time is greater than 24 hours, the urgency reward is 1.

[0050] The constraint satisfaction reward is calculated based on the degree of constraint satisfaction. The reward is highest when the constraints are fully satisfied, and a penalty is given when the constraints are violated. For example, the satisfaction degree of the operation and maintenance personnel's processing capacity constraint can be calculated by the ratio of their current workload to the upper limit of processing capacity: if the load ratio is less than 80%, the reward is 5; if the ratio is between 80% and 100%, the reward is 5 multiplied by (100% minus the load ratio) divided by 20%; if it exceeds 100%, the penalty is -10.

[0051] The experience replay mechanism is adopted to store state transition samples, including the current state, the executed action, the obtained reward, the next state, and the flag of whether it is terminated. The capacity of the experience replay buffer is set to 10,000, and samples with a batch size of 64 are sampled from it each time for training. The priority sampling strategy is introduced, and different weights are assigned according to the temporal difference error of the samples. The higher the temporal difference error of the sample, the higher the sampling probability, which accelerates the network convergence.

[0052] The training samples are input into the double Q-network structure for training, and the network parameters are updated by minimizing the temporal difference error. The temporal difference error is the difference between the actual Q value and the target Q value. The gradient is calculated through the backpropagation algorithm and the Adam optimizer is used to update the current network parameters. The target network parameters are copied from the current network every 100 training steps to maintain learning stability. The learning rate is set to 0.001, and the discount factor is set to 0.95.

[0053] After the training is completed, the network parameters of the updated double Q-network structure are substituted into the ε-greedy strategy to explore in the action space to generate a set of candidate work order distribution schemes. The ε-greedy strategy selects the action with the highest Q value with a probability of 1 - ε and randomly selects an action with a probability of ε. The ε value is initially set to 0.9 and gradually decays to 0.1 as the training progresses to balance exploration and exploitation.

[0054] The Pareto dominance relationship is used to screen the set of candidate work order distribution schemes to obtain the non-dominated solution set. The definition of the Pareto dominance relationship: If scheme A is superior to scheme B in at least one objective (such as matching degree, urgency), and is not inferior to scheme B in other objectives, then scheme A dominates scheme B. The schemes in the non-dominated solution set are not dominated by any other scheme in the set.

[0055] Based on the non-dominated solution set, a reference point method model is constructed. By setting the ideal point coordinates and target weights, the utility function value is calculated, and the work order distribution scheme with the largest utility function value is selected as the optimal scheme. The ideal point coordinates represent the ideal optimal values of each objective, and the target weights reflect the decision maker's preferences for different objectives. The utility function value is calculated as the weighted Chebyshev distance between the scheme and the ideal point. The smaller the distance, the larger the utility function value.

[0056] Let the ideal point coordinates be [1.0, 1.0, 1.0] (representing the ideal values of matching degree, urgency, and constraint satisfaction respectively), and the target weights be [0.4, 0.4, 0.2]; the target values of Plan A are [0.9, 0.8, 0.95], and the target values of Plan B are [0.85, 0.9, 0.9]. Calculate the weighted Chebyshev distance: For Plan A, it is max(0.4×0.1, 0.4×0.2, 0.2×0.05) = 0.08; for Plan B, it is max(0.4×0.15, 0.4×0.1, 0.2×0.1) = 0.06. The distance of Plan B is smaller and the utility function value is larger, so Plan B is selected as the optimal plan.

[0057] Monitor the performance of the optimal work order distribution plan and collect actual operation data such as work order completion rate, processing time limit, and workload balance degree of operation and maintenance personnel. When the change rate of these performance indicators exceeds the preset change threshold (such as 5%), trigger the dynamic adjustment mechanism. The dynamic adjustment mechanism updates the target weights based on the gradient descent method and adjusts the weight values in the direction of performance improvement. For example, if the work order processing time limit decreases, increase the weight of the urgency reward; if the workload distribution is uneven, increase the weight of the constraint satisfaction reward.

[0058] Feed back the updated target weights to the multi-objective reward function, retrain the double Q network, and generate a new optimal work order distribution plan. Through this closed-loop feedback mechanism, the system can continuously optimize itself, adapt to environmental changes, and continuously improve the work order distribution efficiency.

[0059] Figure 3 For the scatter plot comparison table of the matching degree of the work order distribution plan in the embodiment of the present invention: This chart shows the comparison data of the matching degree scores of different work order distribution algorithms under various technical complexities. It can be seen from the table that this technical solution maintains a high and stable matching degree score at each technical complexity level (1.0 - 6.0), ranging from 0.84 to 0.86, with an average value of 0.85 and a standard deviation of only 0.01, showing extremely high stability. In contrast, for the traditional greedy algorithm, as the technical complexity increases, the matching degree drops sharply from 0.75 to 0.64, with an average value of 0.70 and a standard deviation of 0.04; the heuristic allocation algorithm performs the worst, dropping from 0.72 to 0.61, with an average value of only 0.67 and a standard deviation of 0.04; the linear programming algorithm drops from 0.78 to 0.65, with an average value of 0.72 and a standard deviation of 0.05; the benchmark DQN method drops from 0.80 to 0.67, with an average value of 0.74 and a standard deviation of 0.05. The data clearly shows that this technical solution not only has the highest matching degree score at all complexity levels, but also its performance is hardly affected as the technical complexity increases, which proves that this method has excellent adaptability and robustness in complex work order distribution scenarios. Especially in high-complexity work orders (5.0 - 6.0), the advantages of this technical solution are more significant.

[0060] Traditional work order distribution methods usually adopt rule matching or simple priority sorting, and it is difficult to take into account multiple objectives simultaneously; some improved methods introduce machine learning techniques, but mostly based on supervised learning, relying on a large amount of labeled data and having limited adaptability; there are also methods that use single Q-learning or policy gradient algorithms, but they face problems such as unstable convergence and difficulty in dealing with multiple objectives. Existing implementation means lack effective extraction of the correlation of the state space, cannot balance multi-objective conflicts, and lack a dynamic adjustment mechanism.

[0061] This application extracts the correlation of the state space through a dual Q-network structure, reducing the redundancy of the state space representation; introduces a multi-objective reward function to balance the relationship between matching degree, urgency, and constraint satisfaction; combines the Pareto dominance relationship and the reference point method to select the optimal solution based on ensuring the diversity of solutions; and introduces a dynamic adjustment mechanism to optimize decision-making parameters according to the actual operation effect. It enables the system to more accurately capture the core characteristics of the work order distribution problem, balance multi-objective requirements, and have an adaptive learning ability, significantly improving the intelligent level of work order distribution and the overall efficiency of the system.

[0062] In an optional implementation manner, the Pareto dominance relationship is used to screen the set of candidate work order distribution solutions to obtain the non-dominated solution set; a reference point method model is constructed based on the non-dominated solution set. The reference point method model calculates the utility function value based on the ideal point coordinates and objective weights, and selects the work order distribution solution with the largest utility function value as the optimal work order distribution solution, including: Calculate the domination degree and the dominated degree of each solution in the candidate work order distribution solution set based on the Pareto domination relationship, where the domination degree represents the number of other solutions dominated by the corresponding solution, and the dominated degree represents the number of solutions that dominate the corresponding solution; Perform hierarchical screening on the candidate work order distribution solution set according to the domination degree and the dominated degree, and divide the solutions with a dominated degree of zero into multiple layers of leaderless solution sets; Select the first-layer leaderless solution set in the multiple layers of leaderless solution sets to construct a reference point method model, and determine the ideal point coordinates based on the function values corresponding to the multi-objective reward functions of each solution in the multiple layers of leaderless solution sets; Calculate the normalized objective values of each solution in the multiple layers of leaderless solution sets relative to the ideal point coordinates, calculate the utility function values in combination with the preset objective weights, and select the work order distribution solution with the largest utility function value as the optimal work order distribution solution.

[0063] Calculate the domination degree and the dominated degree of each solution in the candidate work order distribution solution set based on the Pareto domination relationship. The Pareto domination relationship means that for two work order distribution solutions A and B, if A is not inferior to B in all optimization objectives and is superior to B in at least one optimization objective, then A is said to dominate B. The domination degree represents the number of other solutions dominated by the corresponding solution, and the dominated degree represents the number of solutions that dominate the corresponding solution.

[0064] The system first obtains the candidate work order distribution solution set from the double Q network and the ε-greedy strategy. Suppose there are N candidate solutions in total. Each solution has M optimization objectives, which are derived from the multi-objective reward function, including matching degree reward, urgency reward, and constraint satisfaction reward. The system constructs an N×N domination relationship matrix, and the matrix element (i,j) represents whether solution i dominates solution j, where 1 means domination and 0 means non-domination.

[0065] Suppose there are 5 candidate solutions, and each solution has 3 optimization objectives (matching degree, urgency, and constraint satisfaction degree), and the objective values are as follows: Solution 1: [0.85, 0.76, 0.92]; Solution 2: [0.78, 0.82, 0.88]; Solution 3: [0.82, 0.79, 0.86]; Solution 4: [0.81, 0.75, 0.91]; Solution 5: [0.87, 0.81, 0.85].

[0066] Compare Solution 1 and Solution 2: Solution 1 is superior to Solution 2 in terms of matching degree (0.85>0.78) and constraint satisfaction degree (0.92>0.88), but inferior to Solution 2 in terms of urgency (0.76<0.82). Therefore, neither of them dominates the other. After comprehensive comparison, the domination relationship matrix is obtained. By summing the rows, the domination degree of each solution is obtained; by summing the columns, the dominated degree of each solution is obtained.

[0067] The calculated domination degrees and dominated degrees of each plan may be as follows: Plan 1: domination degree = 2 (dominating Plan 3 and Plan 4), dominated degree = 0; Plan 2: domination degree = 1 (dominating Plan 3), dominated degree = 1 (dominated by Plan 5); Plan 3: domination degree = 0, dominated degree = 3 (dominated by Plan 1, Plan 2, and Plan 5); Plan 4: domination degree = 0, dominated degree = 2 (dominated by Plan 1 and Plan 5); Plan 5: domination degree = 3 (dominating Plan 2, Plan 3, and Plan 4), dominated degree = 0.

[0068] Based on the domination degrees and dominated degrees, the set of candidate work order distribution plans is hierarchically screened. The plans with a dominated degree of zero are classified into the first-layer non-dominated solution set, that is, the Pareto front. After removing the plans in the first-layer non-dominated solution set from the candidate plan set, the dominated degrees of the remaining plans are recalculated, and the new plans with a dominated degree of zero are classified into the second-layer non-dominated solution set, and so on until all plans are classified into different-layer non-dominated solution sets.

[0069] The first-layer non-dominated solution set contains Plan 1 and Plan 5 (both with a dominated degree of 0); after removing these two plans, the dominated degree of Plan 2 becomes 0 and is classified into the second-layer non-dominated solution set; finally, Plan 3 and Plan 4 are classified into the third-layer non-dominated solution set.

[0070] Select the first-layer non-dominated solution set in the multi-layer non-dominated solution sets to construct a reference point method model. The reference point method is a commonly used solution selection method in multi-objective optimization. It calculates the utility function value of each solution by setting a reference point (ideal point) and a weight vector, and selects the solution with the maximum utility as the final plan.

[0071] Determine the ideal point coordinates based on the function values corresponding to the multi-objective reward functions of each plan in the multi-layer non-dominated solution sets. The ideal point represents the ideal optimal value of each objective, usually taking the optimal value of each objective among all plans. For maximization objectives, take the maximum value of each plan for that objective; for minimization objectives, take the minimum value of each plan for that objective.

[0072] The ideal point coordinates are [0.87, 0.82, 0.92], representing the optimal values of the matching degree, urgency, and constraint satisfaction degree respectively (the matching degree and constraint satisfaction degree take the maximum values, and the urgency is also assumed to take the maximum value). These optimal values come from Plan 5, Plan 2, and Plan 1 respectively.

[0073] Calculate the normalized objective values of each plan in the multi-layer non-dominated solution sets relative to the ideal point coordinates. The normalized objective value represents the deviation degree between the plan objective value and the ideal point, usually obtained by dividing the plan objective value by the corresponding value of the ideal point. After normalization, the closer the objective value is to 1, the closer it is to the ideal point.

[0074] In the first layer of non-dominated solution sets (Scenario 1 and Scenario 5), the normalized objective values of the scenarios are: Scenario 1: [0.85 / 0.87 = 0.977, 0.76 / 0.82 = 0.927, 0.92 / 0.92 = 1.000]; Scenario 5: [0.87 / 0.87 = 1.000, 0.81 / 0.82 = 0.988, 0.85 / 0.92 = 0.924].

[0075] Calculate the utility function value by combining the preset objective weights. The objective weights reflect the degree of importance that decision-makers attach to different objectives. The larger the weight value, the more important the corresponding objective. The utility function value can be calculated by the Chebyshev weighted distance, that is, taking the maximum value among the weighted deviations of each objective, and then taking the opposite number as the utility value. In this way, the smaller the deviation, the larger the utility value.

[0076] Assume that the preset objective weights are [0.4, 0.4, 0.2], corresponding to the weights of matching degree, urgency, and constraint satisfaction respectively. Calculate the utility function values of each scenario in the first layer of non-dominated solution sets: The maximum weighted deviation of Scenario 1 is 0.4×(1 - 0.977) = 0.009 or 0.4×(1 - 0.927) = 0.029 or 0.2×(1 - 1.000) = 0.000. Taking the maximum value of 0.029, the utility function value is -0.029; The maximum weighted deviation of Scenario 5 is 0.4×(1 - 1.000) = 0.000 or 0.4×(1 - 0.988) = 0.005 or 0.2×(1 - 0.924) = 0.015. Taking the maximum value of 0.015, the utility function value is -0.015.

[0077] Since the utility function value of Scenario 5 (-0.015) is greater than that of Scenario 1 (-0.029), the system selects Scenario 5 as the optimal work order distribution plan. This means that considering the decision-maker's preferences for each objective, Scenario 5 performs best in comprehensively balancing all objectives.

[0078] Adopt a method that combines the Pareto dominance relationship and the reference point method to select the optimal work order distribution plan, which effectively solves the problem of conflicts between different objectives in multi-objective optimization. The non-dominated solution set is screened through the Pareto dominance relationship to ensure that the selected plan is not dominated by other plans; The plan that best meets the decision-maker's preferences is selected from the non-dominated solution set through the reference point method, realizing a reasonable balance among multiple objectives. This method overcomes the limitations of traditional single-objective optimization or simple weighting methods, can find the best balance point among multiple objectives such as work order matching degree, processing urgency, and system constraint satisfaction, and significantly improves the intelligence level and operation efficiency of the work order distribution system.

[0079] Figure 4 For the performance improvement table graph of this technical solution of this embodiment of the present invention compared with other algorithms: The table shows the relative improvement percentage of this technical solution in eight core evaluation dimensions compared with four existing algorithms (traditional greedy algorithm, heuristic allocation algorithm, linear programming algorithm and benchmark DQN method). From the data, it can be seen that this solution has achieved significant improvements in all dimensions, especially in standard deviation improvement (75.0%-80.0%) and performance degradation rate improvement (91.8%-92.8%). The comprehensive improvement relative to the heuristic allocation algorithm is the largest, reaching 46.6%, followed by the traditional greedy algorithm (43.6%), the linear programming algorithm (42.4%) and the benchmark DQN method (40.3%). In terms of high-complexity stability, it is improved by 37.7% relative to the heuristic allocation algorithm, while the average matching degree is improved by 26.9%. The uniformity of matching degree distribution is improved by 82.8% relative to the linear programming algorithm, and the resource utilization is improved by 22.3% relative to the heuristic allocation algorithm. In low-complexity scenarios, the advantage over the heuristic allocation algorithm is 19.4%, and the computational efficiency is improved by 18.7%. These data fully demonstrate the excellent performance of this technical solution, especially its significant advantages in stability, consistency and scalability, and its ability to effectively cope with work order distribution scenarios of various complexities.

[0080] In an optional implementation, a Pareto optimal solution algorithm is used to construct a task collaborative scheduling model, which monitors the work order processing status of the target operation and maintenance personnel in real time; when the work order processing status is abnormal, the task collaborative scheduling model generates a task collaborative scheduling strategy based on the work order processing status and the work order processing task, including: The Pareto optimal solution algorithm is used to build a task collaborative scheduling model. The task collaborative scheduling model obtains the work order processing status of the target operation and maintenance personnel. The work order processing status includes the work order processing time, work order processing quality and resource utilization. The work order processing status is converted into a state feature matrix. The state feature matrix generates a historical state code through a bidirectional gated recurrent unit network. Based on historical state coding, the time series state feature data is obtained and multi-scale decomposition is performed. The learning rate is dynamically adjusted according to the state change rate. The dynamic statistics are calculated by combining the exponential moving average and median statistics. The dynamic statistics are weighted and fused through reliability evaluation and smoothing constraints are imposed to obtain the final statistical results. The dynamic mean and dynamic standard deviation of the state feature matrix are calculated using the final statistical results. The state feature matrix is ​​standardized by the dynamic mean and dynamic standard deviation to obtain a standardized state vector. A multidimensional reward and punishment function is constructed. The multidimensional reward and punishment function calculates the reward and punishment value based on the deviation of the work order processing time, the difference in processing quality, and the distance of resource utilization. When the standardized state vector exceeds the reward and punishment value, it is determined that the work order processing state is abnormal. When an abnormal work order processing status is detected, obtain the real-time work order processing progress and information on deployable operation and maintenance resources, and generate candidate scheduling strategies based on the real-time work order processing progress and information on deployable operation and maintenance resources; input the candidate scheduling strategies into the task collaborative scheduling model, calculate the evaluation scores of each strategy using a multi-dimensional reward and punishment function, and select the non-dominated solutions as the task collaborative scheduling strategies through the Pareto optimal solution algorithm.

[0081] Use the Pareto optimal solution algorithm to construct a task collaborative scheduling model. The core function of the task collaborative scheduling model is to obtain the work order processing status of target operation and maintenance personnel and generate collaborative scheduling strategies in case of anomalies. The work order processing status includes three key indicators: work order processing duration, work order processing quality, and resource utilization rate.

[0082] The work order processing duration refers to the actual time taken by operation and maintenance personnel to complete a work order, which can be calculated by recording the time from the start to the completion of the work order. For example, if the expected processing duration of a network configuration work order is 2 hours and the actual processing duration is 2.5 hours, the duration deviation is 0.5 hours. The work order processing quality is calculated through multi-dimensional evaluation indicators, including the fault resolution rate, user satisfaction, recurrence rate, etc. The resource utilization rate measures the utilization of system resources by operation and maintenance personnel during the process of processing work orders, including CPU usage, memory occupancy, network bandwidth usage, etc.

[0083] The system converts these work order processing status data into a status feature matrix. The rows of the status feature matrix represent status records at different time points, and the columns represent different status indicators. For example, for a certain operation and maintenance personnel, the work order processing data for the past 30 days may be collected, with 10 status indicators recorded every day, forming a 30×10 status feature matrix.

[0084] The status feature matrix is processed through a Bidirectional Gated Recurrent Unit (BiGRU) network to generate historical status encodings. The BiGRU network structure includes a forward GRU and a backward GRU, which can capture the forward and backward dependencies of the sequence simultaneously. The hidden layer dimension of the BiGRU network is set to 128, and it consists of 2 layers of BiGRU structures. The input layer receives each row in the status feature matrix as the input for one time step, and the output layer fuses the bidirectional hidden states to output the historical status encodings. The training of the BiGRU network uses the stochastic gradient descent method with a batch size of 64, the initial learning rate is set to 0.001, and the Adam optimizer is used.

[0085] Obtain the time-series state feature data based on historical state encoding and perform multi-scale decomposition on it. The multi-scale decomposition decomposes the time-series features into three parts: trend term, periodic term, and random term, in order to analyze the state changes more precisely. The trend term is extracted by the moving average method with a window size set to 7; the periodic term extracts the main frequency components through Fourier transform; the random term is the residual obtained by subtracting the trend term and the periodic term from the original features.

[0086] Dynamically adjust the learning rate according to the state change rate, which is obtained by calculating the average value of the absolute values of the state differences at adjacent time points. When the state change rate is large, increase the learning rate so that the model can respond quickly to the changes; when the state change rate is small, decrease the learning rate to improve the model stability. Specifically, define the base learning rate as 0.01. When the change rate exceeds the preset threshold (such as 0.05), the learning rate is adjusted to the base learning rate multiplied by the ratio of the change rate to the threshold; when the change rate is less than the threshold, the learning rate is reduced to half of the base learning rate.

[0087] Combine exponential moving average and median statistics to calculate the dynamic statistic. The exponential moving average assigns higher weights to recent data, and the calculation formula is: the current exponential moving average value is equal to the smoothing coefficient multiplied by the current observed value plus (1 minus the smoothing coefficient) multiplied by the exponential moving average value at the previous moment. The smoothing coefficient is set to 0.2, and the smaller smoothing coefficient makes the historical data have a greater impact on the current statistic, enhancing the statistical stability. The median statistic is insensitive to outliers, and calculates the median of the recent n observed values as the current statistic, with the n value set to 15.

[0088] Perform weighted fusion on the dynamic statistic through reliability evaluation and impose a smoothing constraint to obtain the final statistical result. The reliability evaluation assigns weights to each statistical method based on factors such as the sample size, data distribution characteristics, and the proportion of outliers. For example, when the sample size is sufficient and the distribution is normal, the weights of the exponential moving average and the median are set to 0.6 and 0.4 respectively; when there are more outliers, increase the median weight to 0.7 and decrease the exponential moving average weight to 0.3. The smoothing constraint ensures that the statistical results at adjacent time points do not change too much. When the change in the statistical result exceeds the preset threshold (such as 20%), an limiter is applied to limit the change within the threshold range.

[0089] Calculate the dynamic mean and dynamic standard deviation of the state feature matrix using the final statistical result. The dynamic mean is calculated based on the weighted average of the recent m time points, and the weights decay with time; the dynamic standard deviation is calculated based on the square root of the sum of the squared deviations of the recent m time points from the dynamic mean, with the m value set to 30. Standardize the state feature matrix through the dynamic mean and dynamic standard deviation to obtain the standardized state vector. The standardization process subtracts the dynamic mean from the original feature value and then divides by the dynamic standard deviation, making the ranges and distributions of each feature dimension more consistent for subsequent analysis.

[0090] Construct a multi-dimensional reward and punishment function, which calculates the reward and punishment values based on the deviation of work order processing duration, the difference in processing quality, and the distance of resource utilization rate. For the work order processing duration, define the threshold of the reward and punishment function as plus or minus 1.5 standard deviations; for the processing quality, the threshold is plus or minus 2 standard deviations; for the resource utilization rate, the threshold is plus or minus 1 standard deviation. When a certain dimension value in the standardized status vector exceeds the corresponding reward and punishment value, it is determined that there is an abnormality in the work order processing status.

[0091] The standardized value of the work order processing duration of a certain operation and maintenance personnel is 2.3, exceeding the threshold of 1.5, indicating that the processing time is abnormally extended; the standardized value of the processing quality is -2.5, lower than the threshold of -2, indicating that the quality has abnormally declined; the standardized value of the resource utilization rate is 0.8, within the threshold range of plus or minus 1, indicating that the resource usage is normal. Comprehensive judgment shows that there is an abnormality in the work order processing status of this operation and maintenance personnel.

[0092] When an abnormality in the work order processing status is detected, the system obtains the real-time work order processing progress and information on deployable operation and maintenance resources. The real-time work order processing progress includes the percentage of completed tasks, the completion of the critical path, the current processing efficiency, etc.; the information on deployable operation and maintenance resources includes the list of idle operation and maintenance personnel, the skill scores of each person, the current workload, etc.

[0093] Generate candidate scheduling strategies based on the real-time work order processing progress and information on deployable operation and maintenance resources. The candidate strategies may include: resource reallocation strategy, task splitting strategy, priority adjustment strategy, collaborative processing strategy, etc. For example, for the situation where the processing time is abnormally extended, the following candidate strategies can be generated: Strategy 1: Assign an alternative operation and maintenance personnel with a skill score of 85 points and a current load of 30% to assist in processing; Strategy 2: Split the remaining work order tasks into two parts and assign them to the original operation and maintenance personnel and an alternative person with a skill score of 78 points and a current load of 25% respectively; Strategy 3: Increase the priority of the work order so that the original operation and maintenance personnel can focus on processing this work order and postpone the processing of other non-urgent work orders.

[0094] Input the candidate scheduling strategies into the task collaborative scheduling model, and use the multi-dimensional reward and punishment function to calculate the evaluation scores of each strategy. The evaluation dimensions include the expected completion time, resource consumption, impact on processing quality, etc. For example, for the above three candidate strategies, the evaluation results may be: Strategy 1: The expected completion time is reduced by 30%, the resource consumption is increased by 40%, and the processing quality is improved by 15%; Strategy 2: The expected completion time is reduced by 45%, the resource consumption is increased by 35%, and the processing quality is improved by 5%; Strategy 3: The expected completion time is reduced by 20%, the resource consumption remains unchanged, and the processing quality is improved by 10%.

[0095] The non-dominated solutions are selected as the task collaborative scheduling strategy through the Pareto optimal solution algorithm. The Pareto optimal solution algorithm compares each candidate strategy based on the dominance relationship. If strategy A is not inferior to strategy B in all evaluation dimensions and is superior to strategy B in at least one dimension, then strategy A dominates strategy B. The strategies in the non-dominated solution set are not dominated by any other strategy.

[0096] Strategy 1 and strategy 2 do not dominate each other (strategy 1 is better in processing quality, and strategy 2 is better in completion time), and strategy 3 is dominated by strategy 1 (strategy 1 is superior to strategy 3 in all dimensions). Therefore, the non-dominated solution set contains strategy 1 and strategy 2. The system can select one of the strategies according to the current business focus, or provide the two strategies for the decision maker to choose.

[0097] Through the task collaborative scheduling model constructed by adopting the Pareto optimal solution algorithm, the system can monitor the work order processing status of the operation and maintenance personnel in real time, detect abnormal situations in time, generate a multi-dimensional balanced collaborative scheduling strategy, improve the work order processing efficiency and quality, optimize the resource utilization rate, and finally achieve intelligent task collaborative scheduling.

[0098] In an alternative implementation, the time-series state feature data is obtained based on the historical state encoding and multi-scale decomposition is performed. The learning rate is dynamically adjusted according to the state change rate. The dynamic statistic is calculated by combining the exponential moving average and the median statistic. The final statistic result is obtained by weighted fusion of the dynamic statistic through reliability evaluation and applying a smoothing constraint, including: The time-series state feature data is extracted based on the historical state encoding, and the wavelet transform is performed on the time-series state feature data to obtain the multi-scale decomposition coefficients; the difference between the time-series state feature data at adjacent moments is calculated to obtain the instantaneous change rate, and the exponential smoothing is performed on the instantaneous change rate to obtain the cumulative change rate; An adaptive learning rate is generated by mapping through the Sigmoid function based on the cumulative change rate. The value range of the adaptive learning rate is limited by the preset upper bound and lower bound of the learning rate; the exponential moving average operation is performed on the time-series state feature data using the adaptive learning rate to obtain the dynamic mean, and the dynamic standard deviation of the time-series state feature data is calculated based on the dynamic mean; The median of the time-series state feature data is calculated within a sliding time window with a preset length to obtain the robust mean, and the median of the absolute deviation of the time-series state feature data relative to the robust mean is calculated to obtain the robust standard deviation; the relative deviation between the dynamic mean and the robust mean is calculated, and the relative deviation is converted through an exponential function to obtain the reliability evaluation score; Calculate the fusion weights based on the reliability evaluation scores, and use the fusion weights to perform weighted fusion on the dynamic mean and the robust mean to obtain the final mean, and perform weighted fusion on the dynamic standard deviation and the robust standard deviation to obtain the final standard deviation; impose a temporal smoothing constraint on the final mean to limit the change amplitude at adjacent times, and impose a value range constraint on the final standard deviation to ensure its stability; use the final mean and the final standard deviation as the final statistical results.

[0099] Extract the temporal state feature data based on the historical state encoding. The historical state encoding is generated by a bidirectional gated recurrent unit network and contains the temporal information of the work order processing status. Extract the temporal state feature data from the historical state encoding, specifically including the time series of key indicators such as the work order processing duration, processing quality, and resource utilization rate. For example, for the work order processing data of a certain operation and maintenance personnel in the past 30 days, extract the average processing duration, processing quality score, and average resource utilization rate per day to form three time series with a length of 30.

[0100] Perform wavelet transform on the extracted temporal state feature data to obtain the multi-scale decomposition coefficients. The wavelet transform is a time-frequency analysis tool that can decompose the temporal signal into components of different frequency scales. In the present invention, the db4 wavelet basis function is used to perform 4-level decomposition on the temporal features to obtain 4 detail coefficients and 1 approximation coefficient. Taking the work order processing duration sequence as an example, the detail coefficients D1, D2, D3, D4 and the approximation coefficient A4 reflecting the changes at different time scales can be obtained through wavelet transform. Among them, D1 reflects the short-term rapid changes, A4 reflects the long-term trend, and the other detail coefficients reflect the medium-term changes. This multi-scale decomposition helps to distinguish the noise, periodic fluctuations, and long-term trends in the temporal data.

[0101] Calculate the difference between the temporal state feature data at adjacent times to obtain the instantaneous change rate. Specifically, subtract the feature value at time t-1 from the feature value at time t to obtain the instantaneous change rate at time t. For example, if the work order processing durations of a certain operation and maintenance personnel on the 15th day and the 16th day are 2.5 hours and 3.2 hours respectively, then the instantaneous change rate on the 16th day is 0.7 hours. Perform exponential smoothing on the instantaneous change rate to obtain the cumulative change rate. The exponential smoothing calculates the weighted average by assigning decreasing weights to the historical instantaneous change rates. The smoothing parameter is set to 0.3, indicating that the weight of the current instantaneous change rate is 0.3 and the weight of the previous cumulative change rate is 0.7. This smoothing process can reduce the influence of random fluctuations and better reflect the change trend.

[0102] Generate an adaptive learning rate by mapping through the Sigmoid function based on the cumulative change rate. The Sigmoid function maps the input to a value between 0 and 1 and has the characteristic of smooth transition. In the specific implementation, subtract a preset threshold (such as 0.05) from the cumulative change rate and then divide it by a scale parameter (such as 0.02), and then map it to a value between 0 and 1 through the Sigmoid function. The value range of the adaptive learning rate is limited by a preset upper bound and a lower bound of the learning rate. The upper bound of the learning rate is set to 0.2, and the lower bound is set to 0.01. The final adaptive learning rate is equal to the lower bound plus the output value of the Sigmoid function multiplied by the difference between the upper bound and the lower bound. For example, if the cumulative change rate is 0.08 and the output of the calculated Sigmoid function is 0.7, then the adaptive learning rate is 0.01+(0.2 - 0.01)×0.7 = 0.143.

[0103] Perform an exponential moving average operation on the time-series state feature data using the adaptive learning rate to obtain a dynamic mean. The exponential moving average is a method of averaging that assigns higher weights to recent data, and its calculation formula is: the current exponential moving average value is equal to the adaptive learning rate multiplied by the current observation value plus (1 minus the adaptive learning rate) multiplied by the exponential moving average value at the previous moment. The initial value is set to the first observation value. Calculate the dynamic standard deviation of the time-series state feature data based on the dynamic mean. The calculation method is: first calculate the squared difference between each observation value and the dynamic mean at the corresponding moment, then apply the same exponential moving average algorithm to these squared differences, and finally take the square root. The dynamic standard deviation can reflect the fluctuation of the data in real time.

[0104] Calculate the median of the time-series state feature data within a sliding time window of a preset length to obtain a robust mean. The length of the sliding window is set to 15, indicating that each calculation includes the data at the current and the previous 14 moments. The median is the value at the middle position after sorting these data by size. For example, for the data [2.1, 2.3, 1.9, 2.5, 2.7, 2.2, 2.0, 3.5, 2.4, 2.3, 2.2, 2.1, 2.6, 2.5, 2.3] within the window, after sorting it is [1.9, 2.0, 2.1, 2.1, 2.2, 2.2, 2.3, 2.3, 2.3, 2.4, 2.5, 2.5, 2.6, 2.7, 3.5], and the median is 2.3, which is the robust mean. The median is not sensitive to outliers and can provide a more stable mean estimate.

[0105] Calculate the median of the absolute deviations of the time-series state feature data from the robust mean to obtain the robust standard deviation. The specific steps are as follows: First, calculate the absolute difference between each data point in the window and the robust mean, then take the median of these absolute differences, and multiply it by the constant 1.4826 (based on the assumption of a normal distribution). Taking the above data as an example, the absolute difference sequence is [0.2, 0.0, 0.4, 0.2, 0.4, 0.1, 0.3, 1.2, 0.1, 0.0, 0.1, 0.2, 0.3, 0.2, 0.0], the median is 0.2, and the robust standard deviation is 0.2 × 1.4826 = 0.2965. The robust standard deviation is also known as MAD (Median Absolute Deviation), which is a scale estimation method that is insensitive to outliers.

[0106] Calculate the relative deviation between the dynamic mean and the robust mean by dividing the absolute value of their difference by the robust standard deviation. For example, if the dynamic mean is 2.4, the robust mean is 2.3, and the robust standard deviation is 0.2965, then the relative deviation is |2.4 - 2.3| / 0.2965 = 0.337. Convert the relative deviation through an exponential function to obtain the reliability assessment score, which is calculated as the exponential function of the negative relative deviation, that is, e to the power of the negative relative deviation. In the above example, the reliability assessment score is e to the power of negative 0.337, approximately 0.714. The reliability assessment score ranges from 0 to 1, and the larger the value, the closer the dynamic mean is to the robust mean, and the more reliable the dynamic statistical result is.

[0107] Calculate the fusion weights based on the reliability assessment score. Specifically, the fusion weight of the dynamic mean is equal to the reliability assessment score, and the fusion weight of the robust mean is equal to 1 minus the reliability assessment score. In the above example, the fusion weight of the dynamic mean is 0.714, and the fusion weight of the robust mean is 0.286. Use the fusion weights to perform weighted fusion of the dynamic mean and the robust mean to obtain the final mean. The calculation method is: multiply the dynamic mean by its fusion weight and add the robust mean multiplied by its fusion weight. In the above example, the final mean is 2.4 × 0.714 + 2.3 × 0.286 = 2.371. Similarly, perform the same weighted fusion on the dynamic standard deviation and the robust standard deviation to obtain the final standard deviation.

[0108] Apply a time-series smoothing constraint to the final mean to limit the change amplitude between adjacent time moments. Specifically, if the change amplitude of the final mean relative to the previous moment exceeds a preset threshold (such as 50% of the final standard deviation of the previous moment), then limit the change amplitude within the threshold range. For example, if the final mean of the previous moment is 2.2 and the final standard deviation is 0.3, then the change amplitude of the final mean at the current moment should not exceed 0.3 × 50% = 0.15, that is, it should be within the range of [2.05, 2.35]. If the calculated final mean is 2.371, which exceeds the upper limit of 2.35, then adjust it to 2.35.

[0109] Impose a range constraint on the final standard deviation to ensure its stability. Specifically, set the allowable range of the final standard deviation to [σ_min, σ_max], where σ_min is 80% of the minimum standard deviation of historical data and σ_max is 120% of the maximum standard deviation of historical data. If the calculated final standard deviation exceeds this range, adjust it to the boundary value of the range. This constraint prevents extreme values from appearing in the standard deviation estimation and maintains the stability of statistical results.

[0110] Use the final mean and the final standard deviation as the final statistical results for subsequent state anomaly detection and collaborative scheduling strategy generation. The final statistical results have high stability and robustness to outliers, and can accurately reflect the true distribution characteristics of the work order processing status.

[0111] The time series feature processing methods in the prior art mainly have the following problems: First, the exponential moving average with a fixed learning rate is used, which cannot adapt to the dynamic characteristics of the data change rate, resulting in being too lagged during the rapid change period or being too sensitive during the stable period; second, there is no effective processing mechanism for outliers, and a single statistical method is easily affected by abnormal data, reducing the analysis accuracy; third, there is a lack of multi-scale analysis ability and it is unable to distinguish different frequency change patterns; fourth, the statistical results lack time series smoothness, and the statistical values at adjacent moments may jump, affecting the stability of subsequent decisions.

[0112] Figure 5 Schematic diagram of the multi-scale time series feature adaptive statistical analysis system for the embodiments of the present invention: The figure shows the interface of a multi-scale time-series feature adaptive statistical analysis system, which is divided into three main modules. The "Time-series State Features and Multi-scale Decomposition" module at the top shows the historical trend of the work order processing duration and its wavelet decomposition results. In the chart, the blue line is the original data, showing that the processing duration fluctuates between 1.5 and 3.5 hours. The red line is the final mean of about 2.5 hours, and the light blue area represents the confidence interval. The wavelet decomposition below shows the multi-scale signal decomposition from D1 (high frequency) to A4 (trend), and each scale captures different frequency change patterns. The "Change Rate and Adaptive Learning Rate" module in the lower left shows the relationship between the cumulative change rate (current value 0.073) and the adaptive learning rate (current value 0.143). The learning rate is dynamically adjusted between 0.01 and 0.2 through the Sigmoid function (output value 0.702). The "Statistic Calculation" module in the lower right presents various statistical indicators: dynamic mean 2.40 (confidence level 85.2%), dynamic standard deviation 0.352 (confidence level 82.5%), robust mean 2.30 (confidence level 93.7%), robust standard deviation 0.297 (confidence level 91.2%), and the final mean 2.37 and final standard deviation 0.335 after fusion. The system uses a sliding window of 15 time points and a MAD coefficient of 1.4826 for calculation, achieving robust statistical analysis of time-series data.

[0113] Introducing wavelet transform for multi-scale decomposition can capture feature change patterns from different time scales. Secondly, the learning rate is dynamically adjusted based on the state change rate, enabling the statistical method to adaptively respond to data changes. Combining the two complementary methods of exponential moving average and median statistics, the former is efficient in processing normal data, and the latter is robust to outliers. Intelligent weighted fusion is achieved through reliability evaluation, and a smoothing constraint is imposed to ensure the stability of the statistical results.

[0114] The starting point is to improve the accuracy, stability, and self-adaptability of time-series state feature analysis, and finally achieve precise monitoring and anomaly detection of the work order processing status. The effects are as follows: the resistance to outliers is increased by about 40%, the volatility of the statistical results is reduced by about 35%, the response speed to state changes is increased by about 25%, and the time-series consistency of the statistical results is increased by about 30%. It provides a more reliable data basis for the subsequent generation of collaborative scheduling strategies, significantly improving the efficiency and accuracy of task collaborative scheduling.

[0115] In the second aspect of the embodiments of the present invention, An intelligent collaborative management system for network security operation and maintenance work is provided, including: The first unit is used to obtain network security operation and maintenance work order information. The network security operation and maintenance work order information includes work order priority and work order processing time limit, convert the network security operation and maintenance work order information into a work order feature vector, and input the work order feature vector into a multi-objective optimization model trained based on historical work order processing data; The second unit is used for the multi-objective optimization model to calculate the work order matching degree score according to the skill score of the operation and maintenance personnel, the historical work order completion rate, and the current workload, and calculate the work order urgency score based on the work order priority and work order processing time limit; input the work order matching degree score and the work order urgency score into the Pareto optimal solution algorithm to generate a work order distribution plan for multi-objective optimization; The third unit is used to perform state encoding on the work order feature vector, operation and maintenance personnel features, and constraint conditions by using a multi-layer perceptron, extract the spatial correlation of the state encoding by using a double Q-network structure and calculate the target Q value, train the parameters of the double Q-network based on a multi-objective reward function, select the optimal solution by combining the Pareto dominance relationship and the reference point method, and continuously optimize the work order distribution plan through a dynamic adjustment mechanism; The fourth unit is used to determine the target operation and maintenance personnel according to the work order distribution plan and issue the work order processing task to the target operation and maintenance personnel; construct a task collaborative scheduling model by using the Pareto optimal solution algorithm, and the task collaborative scheduling model monitors the work order processing status of the target operation and maintenance personnel in real time; when the work order processing status is abnormal, the task collaborative scheduling model generates a task collaborative scheduling strategy according to the work order processing status and the work order processing task.

[0116] In the third aspect of the embodiments of the present invention, There is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0117] In the fourth aspect of the embodiments of the present invention, There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0118] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent collaborative management method for network security operation and maintenance, characterized in that: include: Obtain network security operation and maintenance work order information, which includes work order priority and work order processing time limit, convert the network security operation and maintenance work order information into a work order feature vector, and input the work order feature vector into a multi-objective optimization model trained based on historical work order processing data; The multi-objective optimization model calculates the work order matching score based on the skill score of the operation and maintenance personnel, the historical work order completion rate and the current workload, and calculates the work order urgency score based on the work order priority and the work order processing time limit; Input the work order matching score and the work order urgency score into the Pareto optimal solution algorithm to generate a multi-objective optimized work order distribution plan; A multi-layer perceptron is used to state-code the work order feature vector, operation and maintenance personnel characteristics, and constraints. The dual Q network structure is used to extract the spatial correlation of the state code and calculate the target Q value. The parameters of the dual Q network are trained based on the multi-objective reward function. The optimal solution is selected by combining the Pareto dominance relationship and the reference point method. The work order distribution solution is continuously optimized through a dynamic adjustment mechanism. The target operation and maintenance personnel are determined according to the work order distribution plan, and the work order processing tasks are issued to the target operation and maintenance personnel; the Pareto optimal solution algorithm is used to build a task collaborative scheduling model, and the task collaborative scheduling model monitors the work order processing status of the target operation and maintenance personnel in real time; when the work order processing status is abnormal, the task collaborative scheduling model generates a task collaborative scheduling strategy based on the work order processing status and work order processing tasks.

2. The method according to claim 1, characterized in that The multi-objective optimization model calculates the work order matching score based on the skill score of the operation and maintenance personnel, the historical work order completion rate and the current workload, and calculates the work order urgency score based on the work order priority and the work order processing time limit; Input the work order matching score and the work order urgency score into the Pareto optimal solution algorithm to generate a multi-objective optimized work order distribution solution including: Obtain the skill scores, historical work order completion rates, and current workloads of the operation and maintenance personnel, and construct a skill score vector based on the skill dimensions; calculate the historical work order completion rates based on the time decay weights, where the time decay weights decay as the statistical time interval increases; calculate the current workload based on the ratio of the remaining workload of the task to the deadline; The priority score is calculated based on the work order priority and urgency coefficient, and the time limit score is calculated based on the difference between the work order processing time limit and the current time. The skill score vector, historical work order completion rate, and current workload are input into the feature fusion layer, and the fusion feature is obtained by weighted fusion through the feature weight matrix. The work order matching score is calculated based on the fusion features. The work order matching score is obtained by the inner product operation of the fusion features and the work order feature vector. The priority score and the time limit score are weighted by the weight coefficient to obtain the work order urgency score. The weight coefficient is optimized and determined based on historical distribution data. A multi-objective optimization function is constructed, which includes a work order matching objective function and a work order urgency objective function. The work order matching objective function is constructed based on the work order matching score, and the work order urgency objective function is constructed based on the work order urgency score. The multi-objective optimization function is input into the Pareto optimal solution algorithm, and the non-dominated solution set is screened based on the Pareto dominance relationship. The solutions in the non-dominated solution set all meet the processing capacity constraint. The adaptive weight adjustment mechanism is used to dynamically optimize the feature weight matrix and weight coefficient, and feedback compensation control is implemented based on performance deviation to generate the optimal work order distribution plan.

3. The method according to claim 2, characterized in that The multi-objective optimization function is input into the Pareto optimal solution algorithm, and the non-dominated solution set is screened based on the Pareto dominance relationship. The solutions in the non-dominated solution set all meet the processing capacity constraint; Adopting the adaptive weight adjustment mechanism to dynamically optimize the feature weight matrix and weight coefficient, implementing feedback compensation control based on performance deviation, and generating the optimal work order distribution plan includes: The multi-objective optimization function includes a first objective function calculated based on the work order matching score and a second objective function calculated based on the work order urgency score, and the multi-objective optimization function is subject to the constraints of the upper limit of the operation and maintenance personnel's processing capacity and the unique allocation of work orders; The multi-objective optimization function is input into the Pareto optimal solution algorithm. The Pareto optimal solution algorithm performs multi-dimensional space mapping on the first objective function and the second objective function, and selects a non-dominated solution set that meets the constraint conditions based on the Pareto dominance relationship. Each solution in the non-dominated solution set corresponds to a candidate work order distribution plan. An adaptive weight adjustment mechanism is constructed. The adaptive weight adjustment mechanism performs performance evaluation on the candidate work order distribution schemes in the non-dominated solution set to obtain the performance deviation. Based on the performance deviation, a proportional-integral controller is used to implement feedback compensation, dynamically adjust the feature fusion weight matrix and the objective function weight coefficient, and select the optimal work order distribution scheme from the non-dominated solution set.

4. The method according to claim 1, characterized in that: The dual Q network structure is used to extract the spatial correlation of state encoding and calculate the target Q value. The parameters of the dual Q network are trained based on the multi-objective reward function. The optimal solution is selected by combining the Pareto dominance relationship and the reference point method. The work order distribution solution is continuously optimized through a dynamic adjustment mechanism, including: The work order feature vector, operation and maintenance personnel features, and system constraint matrix are input into the multi-layer perceptron, and the state feature vector is obtained by state encoding through the multi-layer perceptron; Constructing a dual Q network structure, which includes a current network and a target network, inputting the state feature vector into the current network, extracting the spatial correlation of the state feature through a convolutional neural network, inputting the extracted spatial correlation into the target network and calculating the target Q value; A multi-objective reward function is constructed based on the target Q value. The multi-objective reward function includes matching reward, urgency reward and constraint satisfaction reward. The matching reward is obtained by calculating the cosine similarity between the work order skill requirements and the operation and maintenance personnel skill scores. The urgency reward is calculated based on the remaining processing time of the work order. The constraint satisfaction reward is calculated based on the satisfaction degree of the constraint conditions. The experience replay mechanism is used to store state transition samples, and training samples are selected based on priority sampling. The training samples are input into the dual Q network structure for training, and the network parameters of the dual Q network structure are updated by minimizing the temporal difference error. Substitute the updated network parameters of the double Q network structure into the ε-greedy strategy, explore and generate a set of candidate work order distribution solutions in the action space, and use the Pareto dominance relationship to screen the candidate work order distribution solution set to obtain an unowned solution set; build a reference point method model based on the unowned solution set, and calculate the utility function value based on the ideal point coordinates and target weights, and select the work order distribution solution with the largest utility function value as the optimal work order distribution solution; The performance of the optimal work order distribution plan is monitored. When the performance indicator change rate exceeds the preset change threshold, the dynamic adjustment mechanism is triggered, the target weight is updated based on the gradient descent method, and the updated target weight is fed back to the multi-objective reward function to continuously optimize the optimal work order distribution plan.

5. The method according to claim 4, characterized in that The Pareto dominance relationship is used to screen the candidate work order distribution scheme set to obtain the unowned solution set; a reference point method model is constructed based on the unowned solution set. The reference point method model calculates the utility function value based on the ideal point coordinates and the target weight, and selects the work order distribution scheme with the largest utility function value as the optimal work order distribution scheme, including: Based on the Pareto dominance relationship, the dominance and dominated degrees of each solution in the set of candidate work order distribution solutions are calculated, where the dominance degree indicates the number of other solutions dominated by the corresponding solution, and the dominated degree indicates the number of solutions that dominate the corresponding solution; The candidate work order distribution schemes are hierarchically screened according to the dominance and dominated degrees, and the schemes with a dominated degree of zero are divided into a multi-layer undominated solution set; The first layer of the masterless solution set in the multi-layer masterless solution set is selected to construct a reference point method model, and the coordinates of the ideal point are determined based on the function values ​​corresponding to the multi-objective reward functions of each solution in the multi-layer masterless solution set; The normalized target value of each solution in the multi-layer unowned solution set relative to the coordinates of the ideal point is calculated, and the utility function value is calculated based on the preset target weight. The work order distribution plan with the largest utility function value is selected as the optimal work order distribution plan.

6. The method according to claim 1, characterized in that The Pareto optimal solution algorithm is used to build a task collaborative scheduling model. The task collaborative scheduling model monitors the work order processing status of the target operation and maintenance personnel in real time. When the work order processing status is abnormal, the task collaborative scheduling model generates a task collaborative scheduling strategy based on the work order processing status and the work order processing task, including: The Pareto optimal solution algorithm is used to build a task collaborative scheduling model. The task collaborative scheduling model obtains the work order processing status of the target operation and maintenance personnel. The work order processing status includes the work order processing time, work order processing quality and resource utilization. The work order processing status is converted into a state feature matrix. The state feature matrix generates a historical state code through a bidirectional gated recurrent unit network. Based on historical state coding, the time series state feature data is obtained and multi-scale decomposition is performed. The learning rate is dynamically adjusted according to the state change rate. The dynamic statistics are calculated by combining the exponential moving average and median statistics. The dynamic statistics are weighted and fused through reliability evaluation and smoothing constraints are imposed to obtain the final statistical results. The dynamic mean and dynamic standard deviation of the state feature matrix are calculated using the final statistical results. The state feature matrix is ​​standardized by the dynamic mean and dynamic standard deviation to obtain a standardized state vector. A multidimensional reward and punishment function is constructed. The multidimensional reward and punishment function calculates the reward and punishment value based on the deviation of the work order processing time, the difference in processing quality, and the distance of resource utilization. When the standardized state vector exceeds the reward and punishment value, it is determined that the work order processing state is abnormal. When an abnormal work order processing status is detected, the real-time work order processing progress and the deployable operation and maintenance resource information are obtained, and candidate scheduling strategies are generated based on the real-time work order processing progress and the deployable operation and maintenance resource information; the candidate scheduling strategies are input into the task collaborative scheduling model, and the evaluation score of each strategy is calculated using a multi-dimensional reward and punishment function. The non-dominated solution is selected as the task collaborative scheduling strategy through the Pareto optimal solution algorithm.

7. The method according to claim 6, characterized in that Based on the historical state coding, the time series state feature data is obtained and multi-scale decomposition is performed. The learning rate is dynamically adjusted according to the state change rate. The dynamic statistics are calculated by combining the exponential moving average and median statistics. The dynamic statistics are weighted and fused through reliability evaluation and smooth constraints are imposed to obtain the final statistical results including: Extract the time series state feature data based on the historical state coding, perform wavelet transform on the time series state feature data to obtain the multi-scale decomposition coefficient; calculate the difference of the time series state feature data at adjacent moments to obtain the instantaneous change rate, perform exponential smoothing on the instantaneous change rate to obtain the cumulative change rate; Based on the cumulative change rate, an adaptive learning rate is generated through Sigmoid function mapping, and the value range of the adaptive learning rate is limited by the preset upper bound and lower bound of the learning rate; the adaptive learning rate is used to perform exponential moving average operation on the time series state feature data to obtain a dynamic mean, and the dynamic standard deviation of the time series state feature data is calculated based on the dynamic mean; The median of the time series state feature data is calculated within a sliding time window of a preset length to obtain a robust mean, and the median of the absolute deviation of the time series state feature data relative to the robust mean is calculated to obtain a robust standard deviation; the relative deviation between the dynamic mean and the robust mean is calculated, and the relative deviation is converted through an exponential function to obtain a reliability assessment score; The fusion weight is calculated based on the reliability assessment score, and the dynamic mean and the robust mean are weightedly fused using the fusion weight to obtain the final mean. The dynamic standard deviation and the robust standard deviation are weightedly fused to obtain the final standard deviation. A time series smoothing constraint is imposed on the final mean to limit the amplitude of change between adjacent moments, and a value range constraint is imposed on the final standard deviation to ensure its stability. The final mean and the final standard deviation are taken as the final statistical results.

8. An intelligent collaborative management system for network security operation and maintenance, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain network security operation and maintenance work order information, which includes work order priority and work order processing time limit, convert the network security operation and maintenance work order information into a work order feature vector, and input the work order feature vector into a multi-objective optimization model trained based on historical work order processing data; The second unit is used for the multi-objective optimization model to calculate the work order matching score based on the skill score of the operation and maintenance personnel, the historical work order completion rate and the current workload, and to calculate the work order urgency score based on the work order priority and the work order processing time limit; Input the work order matching score and the work order urgency score into the Pareto optimal solution algorithm to generate a multi-objective optimized work order distribution plan; The third unit is used to use a multi-layer perceptron to state-code the work order feature vector, operation and maintenance personnel characteristics and constraints, use the dual Q network structure to extract the spatial correlation of the state code and calculate the target Q value, train the parameters of the dual Q network based on the multi-objective reward function, combine the Pareto dominance relationship and the reference point method to select the optimal solution, and continuously optimize the work order distribution plan through a dynamic adjustment mechanism; The fourth unit is used to determine the target operation and maintenance personnel according to the work order distribution plan, and issue the work order processing tasks to the target operation and maintenance personnel; the Pareto optimal solution algorithm is used to build a task collaborative scheduling model, and the task collaborative scheduling model monitors the work order processing status of the target operation and maintenance personnel in real time; when the work order processing status is abnormal, the task collaborative scheduling model generates a task collaborative scheduling strategy based on the work order processing status and the work order processing task.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Work order processing method and device, equipment and storage medium

    CN114004597A

  • Mobile crowdsourcing strategy optimization method and system based on multi-objective optimization

    CN119204625A

  • Constrained reinforcement learning neural network systems using pareto front optimization

    EP4475037A2

Cited By

  • Base-level grid member-oriented task pushing method and system

    CN120509691A

  • Work order screening management method and system based on dynamic sorting

    CN121352424A

  • A work order screening management method and system based on dynamic sequencing

    CN121352424B

  • Multi-objective collaborative optimization pipe network automatic transmission and distribution control method

    CN121364636A

  • Intelligent work order scheduling method and system for water supply network

    CN121390620A