A multi-task parallel power material detection verification method

By employing a multi-task parallel power material inspection method, which dynamically allocates inspection tasks using a deep learning model and combines time series analysis, feature comparison, and semantic parsing sub-units, the problem of low efficiency and insufficient accuracy in power material inspection is solved, achieving efficient and accurate inspection and verification.

CN121117548BActive Publication Date: 2026-02-13JIANGSU ELECTRIC POWER INFORMATION TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511650166.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-13
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

Existing power material testing technologies are inefficient, and their accuracy is greatly affected by human factors, making it difficult to meet the requirements of rapid response and accurate capture of numerical drift, standard misuse, and potential tampering in verification results data.

Method used

A multi-task parallel power material inspection method is adopted, which uses a deep learning model to dynamically allocate inspection and verification tasks. It combines time series analysis, feature comparison and semantic parsing sub-units, and optimizes task allocation through the DQN model to achieve multi-modal feature fusion and structured conclusion output.

Benefits of technology

It improved the utilization rate of testing resources, shortened the testing cycle, enhanced testing efficiency and accuracy, provided comprehensive testing and verification conclusions, and met the real-time requirements of power material testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121117548B_ABST
    Figure CN121117548B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-task parallel electric power material detection verification method and device, comprising: obtaining multi-source electric power material detection data and the current state of each analysis unit, respectively give integrated dataset and state characteristics;From integrated dataset, the task data of each detection verification task is extracted, and task characteristics are given;Based on the reward of fusion rhythm compliance, the state characteristics and task characteristics are processed, and the task parallel processing allocation result is given;Each analysis unit is based on the task parallel processing allocation result to analyze and process the task data of detection verification task, and the verification result corresponding to the detection verification task is given;Multi-modal features in verification result are extracted, and combined with the preset structured instruction, the detection verification conclusion is given.The analysis unit state, task characteristics, rhythm compliance reward are fused, and the optimal allocation result of detection verification task is dynamically given, to improve the detection verification efficiency, detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of electric power material detection, and particularly relates to a multi-task parallel electric power material detection verification method and device. BACKGROUND

[0002] With the rapid development of the electric power industry, the quality and safety of electric power materials are directly related to the stable operation of the power grid and power supply reliability. Electric power material detection, as an important link to ensure the quality of electric power materials, its efficiency and accuracy are of great significance to improve the overall operation level of the power grid.

[0003] The existing electric power material detection technology mainly has the following shortcomings: low detection efficiency, since the serial processing mode is adopted, each detection link needs to wait for the completion of the previous link before proceeding, resulting in a long overall detection period, which cannot meet the demand for rapid response; the detection accuracy is greatly affected by human factors, manual operation is easy to introduce errors, and the processing capacity for complex data is limited, making it difficult to accurately capture the value drift, standard misuse and potential tampering in the verification result data.

[0004] Patent CN112330014A discloses a task scheduling method, device, equipment and storage medium. According to the number of task detection platforms and the number of tasks to be detected, a task scheduling variable is constructed; according to the task scheduling variable, the task test time and downtime of the task detection platform, a target function is constructed; according to the target function and the constraint condition of the task to be detected, the task scheduling value of the task scheduling variable is determined; according to the task scheduling value, the task to be detected is scheduled. Under the condition of high detection resource utilization, the detection verification task is reasonably allocated and scheduled, realizing efficient material detection, and providing a new idea for simultaneous detection of multiple materials on multiple platforms. However, in this method, the detection verification task type is not distinguished in detail; nor is a matching detection algorithm designed for the characteristics of different detection verification tasks; the detection project results cannot be comprehensively evaluated.

[0005] Therefore, how to design an electric power material detection verification method, introduce a multi-task parallel processing mechanism, and design a targeted detection algorithm for different detection verification tasks, and comprehensively evaluate the verification results of different detection verification tasks to improve the detection accuracy, improve the detection verification efficiency, shorten the detection period, and make the detection verification conclusion more comprehensive, is a problem to be solved by the skilled in the art. SUMMARY

[0006] In view of the defects in the prior art, the application provides a multi-task parallel power material detection verification method and device, which comprises the following steps: obtaining multi-source power material detection data and the current state of each analysis unit, and giving an integrated data set and state features; extracting task data of each detection verification task from the integrated data set, and giving task features; processing the state features and the task features based on a deep learning model that fuses beat compliance rewards, and giving a task parallel processing allocation result; each analysis unit analyzes and processes the task data of the detection verification task based on the task parallel processing allocation result, and gives a verification result corresponding to the detection verification task; multi-modal features in the verification result are extracted, and a detection verification conclusion is given by combining a preset structured instruction, and the verification of the power material detection is completed. By fusing the analysis unit state, the task features, and the beat compliance rewards, the optimal allocation result of the detection verification task is dynamically given, the detection resource utilization is optimized, the real-time requirement is met, the detection verification task is executed in parallel, the detection cycle is greatly shortened, the multi-modal fusion is performed, the structured detection verification conclusion is output, and the decision efficiency is improved.

[0007] In a first aspect, the application provides a multi-task parallel power material detection verification method, which is applied to a verification device and comprises an analysis unit, and specifically comprises the following steps:

[0008] Obtaining multi-source power material detection data and the current state of each analysis unit, and giving an integrated data set and state features;

[0009] Extracting task data of each detection verification task from the integrated data set, and giving task features;

[0010] Processing the state features and the task features based on a deep learning model that fuses beat compliance rewards, and giving a task parallel processing allocation result;

[0011] Each analysis unit analyzes and processes the task data of the detection verification task based on the task parallel processing allocation result, and gives a verification result corresponding to the detection verification task;

[0012] Extracting multi-modal features in the verification result, and giving a detection verification conclusion by combining a preset structured instruction, and completing the verification of the power material detection.

[0013] Further, the detection verification task comprises a drift verification task, a tampering identification task, and a standard comparison task.

[0014] Further, the analysis unit comprises at least one of a time sequence analysis subunit, a feature comparison subunit, and a semantic analysis subunit.

[0015] The time series analysis subunit adopts a time series model with the optimization objective of maximizing drift detection accuracy, takes the power material time series data set labeled with numerical drift events as training data, and iteratively trains the time series model to convergence to obtain;

[0016] The feature comparison subunit adopts a twin network model with the optimization objective of maximizing tampering recognition F1-score, takes the paired data set containing original power material detection features and fake tampered features as training data, and iteratively trains the twin network model to convergence to obtain;

[0017] The semantic analysis subunit adopts a pre-trained language model with the optimization objective of maximizing the semantic matching accuracy of standard clauses and detection reports, takes the paired data set of power standards and detection reports labeled with standard clause matching results as training data, and iteratively trains the pre-trained language model to convergence to obtain.

[0018] Further, the state features include task matching degree, load state, historical processing time, and current task remaining completion time, and the task features include task type, data volume, and priority.

[0019] Further, the current state of each analysis unit is obtained, and the state features are given, including:

[0020] Based on the current state of each analysis unit, the current task backlog of each analysis unit, and the current task remaining completion time are determined;

[0021] The best processing task type and historical processing time corresponding to each analysis unit are obtained;

[0022] Based on the best processing task type corresponding to each analysis unit, and combined with the current detection verification task type, the task matching degree is given;

[0023] Based on the current task backlog of each analysis unit, the load state is determined;

[0024] The task matching degree, load state, current task remaining completion time, and historical processing time are normalized to give the state features.

[0025] Further, the task data of each detection verification task is extracted from the integrated data set, and the task features are given, including:

[0026] According to the task data in the integrated data set, the data type is determined;

[0027] Based on the data type, the corresponding task type is given;

[0028] According to the given task type, combined with the preset conditions, the priority of different types of tasks is determined;

[0029] Extract each type of data in the integrated data set to give the data volume corresponding to each task type;

[0030] Based on the task type, data volume and priority, the task characteristics are given.

[0031] Further, the deep learning model is a DQN model, and the deep learning model pre-constructed based on the beat compliance reward specifically includes the following steps:

[0032] Build a DQN model, define a global state space, take all feasible allocation results of parallel processing of detection and verification tasks as an action space, and construct a reward function based on detection accuracy, detection efficiency, failure penalty, load penalty and beat compliance reward;

[0033] Obtain the state features of the analysis unit and the task features and beat constraint features of the detection and verification tasks to give the current global state;

[0034] Based on the current global state, an exploration and utilization strategy is used to select an action, and the reward after executing the action and the next global state are given;

[0035] The current global state, action, reward, and next global state are stored as experience in an experience replay pool;

[0036] Repeat the process of selecting an action, executing an action, and storing experience until the capacity of the experience replay pool exceeds a preset value;

[0037] Randomly sample from the experience replay pool to train the main network and target network of the DQN model, and give the final deep learning model.

[0038] Further, randomly sample from the experience replay pool to train the main network and target network of the DQN model, specifically including the following steps:

[0039] Randomly sample a batch of experiences from the experience replay pool;

[0040] Give the Q value of the current global state by combining the extracted experience through the main network of the DQN model;

[0041] Give the maximum Q value of the next global state by combining the reward function through the target network of the DQN model;

[0042] Determine the target loss of the Q value of the current global state and the maximum Q value of the next global state, update the main network parameters in reverse, and update the target network parameters regularly.

[0043] Further, the detection and verification task includes a drift verification task, and the task data corresponding to the detection and verification task includes time series data;

[0044] The task data of the detection verification task is analyzed and processed to give a verification result corresponding to the detection verification task, specifically including:

[0045] The analysis unit processes the time series data based on the weight function of the fused environmental factors to give a local weighted regression smoothing value at each time;

[0046] Based on the local weighted regression smoothing value, the mean and standard deviation within the sliding window are calculated, and a dynamic threshold is given in combination with the environmental self-adaptive adjustment coefficient;

[0047] Based on the local weighted regression smoothing value at the current time point, in combination with the dynamic threshold, a drift verification result is given.

[0048] Further, based on the weight function of the fused environmental factors, the time series data is processed to give a local weighted regression smoothing value at each time, specifically including the following steps:

[0049] Obtain multiple environmental factors and construct a weight function in combination with a multivariate Gaussian kernel function, determine the weight of each time data within the target time corresponding local window, and give a weight diagonal matrix;

[0050] Determine the time series data within the target time corresponding local window, construct a design matrix, and give a detection value vector within the local window;

[0051] In combination with the detection value vector and the weight diagonal matrix, a weighted least squares model is constructed to solve the regression coefficient;

[0052] In combination with the design vector of the target time and the regression coefficient, a local weighted regression smoothing value at each target time is given.

[0053] Further, based on the detection accuracy, detection efficiency, failure penalty, load penalty and beat compliance reward, a reward function is constructed, specifically including the following steps:

[0054] Based on the basic accuracy, task matching degree, task priority, and accuracy basic weight, the detection accuracy is given;

[0055] Based on the beat time, historical processing time, data volume, and remaining beat time, the efficiency factor and beat emergency coefficient are determined, and in combination with the efficiency basic weight, the detection efficiency is given;

[0056] According to the task type, the task type weight is obtained, and in combination with the task priority, failure penalty basic weight and failure flag, the failure penalty is given;

[0057] Based on the current task remaining completion time, the remaining time coefficient is obtained, and in combination with the load penalty basic weight and load state, the load penalty is given;

[0058] Based on the data volume, a data volume coefficient is acquired, and in combination with the priority, the takt compliance flag, and the takt compliance base weight, a takt compliance reward is given.

[0059] Based on the influence coefficients of the detection accuracy, the detection efficiency, the failure penalty, the load penalty, and the takt compliance reward on the multi-task parallel processing, the construction of the reward function is completed.

[0060] Further, the reward function is specifically expressed as:

[0061]

[0062] In the formula, r represents the reward function, Accuracy represents the detection accuracy, Efficiency represents the detection efficiency, FailP represents the failure penalty, LoadP represents the load penalty, and TaktA represents the takt compliance reward.

[0063] The detection accuracy is specifically expressed as:

[0064]

[0065] In the formula, Accuracy represents the detection accuracy, is the accuracy base weight, TaskMatch represents the task matching degree, Prio represents the priority, and BaseAcr represents the base accuracy.

[0066] The detection efficiency is specifically expressed as:

[0067]

[0068] In the formula, Efficiency represents the detection efficiency, is the efficiency base weight, TaktTime represents the takt time, RemainTime represents the remaining takt time, HistTime represents the historical processing time, DataVolume represents the data volume, and MaxDataVolume represents the maximum data volume threshold.

[0069] The failure penalty is specifically expressed as:

[0070]

[0071] In the formula, FailP represents the failure penalty, is the failure penalty base weight, is the task type weight, Prio represents the priority, and Failure represents the failure flag.

[0072] The load penalty is specifically expressed as:

[0073]

[0074] wherein LoadP is a load penalty, Load is a load weight, TaskRemainTime is a current task remaining completion time, and MaxRemainTime is a maximum remaining beat time threshold value;

[0075] TaktA is a beat adherence reward,

[0076]

[0077] TaktA is a beat adherence reward, Prio is a priority, DataVolume is a data volume, MaxDataVolume is a maximum data volume threshold value, and TaktComp is a beat adherence flag.

[0078] In a second aspect, the present application further provides a multi-task parallel power material detection verification device, which adopts the multi-task parallel power material detection verification method, and specifically comprises:

[0079] A preprocessing unit is configured to obtain multi-source power material detection data and current states of each analysis unit, and give integrated data sets and state features, respectively; and extract task data of each detection verification task from the integrated data sets, and give task features.

[0080] A task allocation unit is configured to process the state features and the task features based on the deep learning model fused with the beat adherence reward, and give a task parallel processing allocation result.

[0081] An analysis unit is configured to analyze and process the task data of the detection verification task based on the task parallel processing allocation result, and give a verification result of the corresponding detection verification task.

[0082] A conclusion output unit is configured to extract multi-modal features in the verification result, combine preset structured instructions, give a detection verification conclusion, and complete the verification of the power material detection.

[0083] The multi-task parallel power material detection verification method and device provided by the present application have at least the following beneficial effects:

[0084] (1) Through the multi-task parallel and task dynamic allocation strategy based on the deep learning model, the multi-class detection verification tasks are synchronously executed, the detection resource utilization rate is improved, the detection verification period is shortened, and the power material detection efficiency is improved.

[0085] (2) The global state space of the DQN model fuses state features, task features, and rhythm compliance rewards, can perceive environmental changes in real time, dynamically adjusts task parallel processing allocation results, improves the dynamic adaptability of task allocation, optimizes resource utilization, and meets real-time requirements.

[0086] (3) The reward function of the DQN model takes into account detection accuracy, detection efficiency, failure penalty, load balancing, rhythm compliance, and other objectives, guiding the DQN to learn an accurate, efficient, and stable allocation strategy.

[0087] (4) Corresponding analysis units can be constructed for different detection and verification tasks, and corresponding detection and verification algorithms are used to improve the accuracy and pertinence of task processing. Comprehensive detection and verification conclusions of multiple detection and verification tasks are given to provide reliable decision basis for power material detection scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0088] Figure 1 A flowchart of a multi-task parallel power material detection verification method provided by the present application is provided.

[0089] Figure 2 A pre-training process diagram of an analysis unit of an embodiment provided by the present application is provided.

[0090] Figure 3 A construction flowchart of a deep learning model of an embodiment provided by the present application is provided.

[0091] Figure 4 A DQN model architecture diagram of an embodiment provided by the present application is provided.

[0092] Figure 5 A construction flowchart of a reward function of an embodiment provided by the present application is provided.

[0093] Figure 6 A flowchart of an embodiment provided by the present application is provided to give the verification result of the corresponding detection and verification task.

[0094] Figure 7 A flowchart of an embodiment provided by the present application is provided to give the local weighted regression smoothing value at each time.

[0095] Figure 8 A diagram of an embodiment provided by the present application is provided to give the detection and verification conclusion.

[0096] Figure 9 A multi-task parallel power material detection verification device structure diagram provided by the present application is provided. DETAILED DESCRIPTION

[0097] For better understanding of the above technical solutions, the above technical solutions will be described in detail below in combination with the drawings of the specification and specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor belong to the scope of protection of the present application.

[0098] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Multiple" generally includes at least two.

[0099] It should also be noted that the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the goods or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such goods or devices. Without more limitation, the element defined by the sentence "including a" does not exclude the presence of other identical elements in the goods or devices including the element.

[0100] The quality of electric power materials is directly related to the stable operation of the power grid, and the quality detection data of production, transportation, acceptance and other links are stored in a scattered manner, lacking a unified correlation mechanism. The traditional method verifies and processes the detection data in a serial manner, which is low in efficiency and difficult to guarantee the detection accuracy. Therefore, the present application provides a multi-task parallel electric power material detection verification method and device, which designs a dynamic allocation mechanism for detection verification tasks with the help of a deep learning model, executes the detection verification tasks in parallel, integrates the verification results, and gives the verification conclusion.

[0101] As shown in Figure 1 A multi-task parallel electric power material detection verification method applied to a verification device, the verification device comprising an analysis unit, characterized in that it specifically comprises the following steps:

[0102] Obtain multi-source electric power material detection data and the current state of each analysis unit, and give integrated data sets and state features respectively;

[0103] Determine a plurality of detection verification tasks, and extract task data of each detection verification task from the integrated data set to give task features;

[0104] Based on the deep learning model of fusion beat compliance reward, the state features and task features are processed to give task parallel processing allocation results;

[0105] According to the task parallel processing distribution result, each analysis unit executes a task data analysis process corresponding to the detection verification task in parallel, and gives a verification result corresponding to the detection verification task;

[0106] The multi-modal features in the verification result are extracted, and a detection verification conclusion is given in combination with a preset structured instruction, so as to complete the verification of the power material detection.

[0107] Obtaining multi-source power material detection data, giving an integrated data set, including: aligning the multi-source power material detection data to give an integrated data set. Among them, the multi-source power material detection data includes the factory detection report of the power material, the sampling detection data of the third-party detection agency and the retest data before the field installation, etc. Taking the material code + detection item + batch number as the key, the multi-source data is grouped by the KeyBy operation of Flink, so that the same batch of material detection data scattered in multiple sources is gathered together. An adaptive sliding window algorithm is used to dynamically adjust the window size to ensure that the detection data in the same window belongs to the same time interval. For the detection data in each window, a feature vector is extracted, and a feature weight is calculated. For each pair of detection data in the window, the cosine similarity is calculated, and the detection data with a similarity higher than a set threshold is associated and integrated as the detection data of the same batch of power materials, completing the alignment of the power material detection data. Finally, the integrated data set given is a structured table containing unique keys, timestamps and multi-source features, including time series data (i.e. continuous time point detection data of the same batch and the same detection item), original detection data (such as factory detection report), real-time detection data (such as retest data before field installation), standard clauses (power industry standard clauses) and other types of data.

[0108] Further, the detection verification task includes a drift verification task, a tampering identification task and a standard comparison task. The drift verification task refers to monitoring the dynamic change of the power material detection data over time to determine whether it is beyond the normal fluctuation range; the tampering identification task refers to checking whether the power material detection data has been modified or forged by human; the standard comparison task refers to verifying whether the power material detection result meets the industry standard or enterprise standard.

[0109] The analysis unit is an execution component of the detection verification task, responsible for processing specific tasks in the power material detection verification. The analysis unit includes at least one of a time series analysis subunit, a feature comparison subunit and a semantic analysis subunit. Each analysis subunit is a data processing model obtained by pre-training, which has professional ability to process corresponding tasks (such as time series analysis, feature comparison and semantic matching), executes detection algorithm and outputs verification result.

[0110] As shown in Figure 2 , the pre-training of the analysis unit includes:

[0111] The time series model is adopted, the power material time series dataset marked with numerical drift events is taken as training data, the drift detection accuracy maximization is taken as the optimization target, and the iteration training is performed until convergence to obtain the time series analysis subunit;

[0112] The twin network model is adopted, the paired dataset containing original power material detection features and forged and tampered features is taken as training data, the F1-score maximization of tampering identification is taken as the optimization target, and the iteration training is performed until convergence to obtain the feature comparison subunit;

[0113] The pre-trained language model is adopted, the power standard and detection report paired dataset marked with standard clause matching results is taken as training data, the semantic matching accuracy maximization of standard clauses and detection reports is taken as the optimization target, and the iteration training is performed until convergence to obtain the semantic analysis subunit.

[0114] Specifically, the time series model can adopt an isolated forest model, an LSTM model, etc. Numerical drift events are marked on the power material time series dataset as training data. The drift detection accuracy refers to the proportion of correctly identified drift events. During the training process, the learning rate and other model parameters are adjusted, and the model is iterated for a certain number of training rounds until convergence, i.e., the time series analysis subunit is obtained. The time series analysis subunit can analyze the dynamic changes of power material detection data (such as cable insulation resistance and transformer oil breakdown voltage) over time and identify numerical drifts that exceed the normal fluctuation range.

[0115] The twin network model includes an input layer, a contribution encoder, a distance measurement layer, and an output layer. The input layer receives paired feature vectors, i.e., original data feature vectors and forged and tampered data feature vectors; the encoder compresses high-dimensional feature vectors into low-dimensional vectors, which can adopt a fully connected neural network; the distance measurement layer adopts cosine distance to calculate the similarity of two low-dimensional vectors; the output layer adopts a threshold method to convert the similarity into a tampering judgment result combined with a preset threshold. During training, positive and negative sample pairs are constructed, and a contrast loss is used as the loss function; the loss is calculated, and the network parameters are updated using stochastic gradient descent (SGD) or Adam optimizer, and the iteration training is performed until the model converges, i.e., the feature comparison subunit is obtained. The F1-score of tampering identification is an index that combines the precision index and recall rate. The feature comparison subunit can compare the feature differences between original detection records (such as factory reports of manufacturers) and real-time detection data (such as on-site retest data) to identify data tampering.

[0116] The pre-trained language model can adopt BERT, RoBERTa, Sentence-BERT, etc., and the paired text of the standard clause and the detection report as the training data. During the training process, the learning rate, sentence vector dimension and other parameters are adjusted to maximize the semantic matching accuracy, and the model converges through iteration, i.e. the semantic analysis subunit is obtained. The semantic analysis subunit can analyze the semantic association between the power standard clause and the detection report, and determine whether the detection data meets the standard.

[0117] Further, the state features include task matching degree, load state, historical processing time and current task remaining completion time, and the task features include task type, data volume and priority.

[0118] The current state of each analysis unit is obtained, and the state features are given, which can include: based on the current state of each analysis unit, determining the current task backlog of each analysis unit, the current task remaining completion time; obtaining the best processing task type and historical processing time corresponding to each analysis unit; based on the best processing task type corresponding to each analysis unit, and combining the current detection verification task type, giving the task matching degree; based on the current task backlog of each analysis unit, determining the load state; normalizing the task matching degree, the load state, the current task remaining completion time and the historical processing time, and giving the state features.

[0119] It can be understood that the current state of the analysis unit refers to the running state of the analysis unit at the current time, reflecting the task processing capability, busy degree and current task progress of the unit. In specific implementation, a state monitoring module can be set for each analysis unit to obtain its state information through real-time monitoring and historical data query. The task matching degree refers to the matching degree of the task type that the analysis unit is good at and the current to-be-assigned task type, and the analysis unit can be pre-labeled with the task type that it is good at, i.e. the task matching degree of each task; the load state is the task backlog degree of the analysis unit at present, which can be obtained by monitoring the task queue of the analysis unit; the historical processing time is the average time of the analysis unit processing the same type of task as the detection verification task; the current task remaining completion time refers to how long the task being processed by the analysis unit will take to complete. The above state data is converted into a numerical value and normalized to integrate into a state feature vector.

[0120] The task data of each detection verification task is extracted from the integrated data set, and the task features are given, which can include: determining the data type according to the task data in the integrated data set; giving the corresponding task type based on the data type; determining the priority of different types of tasks according to the given task type and the preset conditions; extracting each type of data in the integrated data set to give the data volume corresponding to each task type; and giving the task features based on the task type, data volume and priority.

[0121] It needs to be understood that, due to the different task characteristics and processing data of each detection verification task, task data strongly related to the task characteristics needs to be extracted from the integrated data set to ensure the effectiveness of the corresponding detection algorithm and improve detection efficiency and accuracy. For the drift verification task, time series data of the same material, the same detection item and the same batch are extracted; for the tamper identification task, feature vectors of original detection data records (such as factory detection data) and real-time detection data (such as on-site detection data) are extracted; for the standard comparison task, semantic information (such as detection report description) of real-time detection data and standard clauses are extracted. Task features are vectors that describe task attributes. Each detection task can be labeled with task type using one-hot encoding; data volume is a scale indicator of task data, which directly affects the processing time and resource consumption of the analysis unit, and can be measured by the number of time points, the number of text characters and the number of clauses; priority refers to the degree of urgency or importance of the task, which is used to determine the processing order of the task, and can be determined according to business rules (preset conditions).

[0122] In a specific example, the deep learning model uses a DQN model, as shown in Figure 3 The pre-construction of the deep learning model that fuses the beat compliance reward specifically includes the following steps:

[0123] An initial DQN model is built, and a global state space is defined, with all feasible allocation results of the detection verification tasks processed in parallel as the action space. A reward function is constructed based on detection accuracy, detection efficiency, failure penalty, load penalty and beat compliance reward, wherein the DQN model includes a main network and a target network;

[0124] The main network parameters and target network parameters of the DQN model are initialized, and an experience replay pool is initialized;

[0125] The current global state is given by obtaining and combining state features, task features and beat constraint features;

[0126] Based on the current global state, an exploration-exploitation strategy is used to select an action;

[0127] The action is executed, and the reward after the action is executed and the next global state are recorded;

[0128] The current global state, action, reward and next global state are stored in the experience replay pool as experience;

[0129] The process of selecting an action, executing an action and storing experience is repeated until the capacity of the experience replay pool exceeds a preset value;

[0130] Random sampling is performed from the experience replay pool to train the main network and the target network of the DQN model, and the final deep learning model is given.

[0131] The cycle time constraint features include cycle time, used cycle time, and remaining cycle time. These features are essentially time constraints on task execution, used to quantify the time pressure of a task. Cycle time refers to the maximum allowed time to complete a single detection / verification task or a batch of detection tasks, pre-set according to business needs. Used cycle time refers to the time elapsed from the start of the task to the current moment, calculated using the system clock or timestamp. Remaining cycle time is the remaining available time after subtracting used cycle time from the cycle time, reflecting the upper limit of the task's acceptable latency. These cycle time constraint features allow deep learning models to perceive the time pressure of a task, thereby adjusting and optimizing task allocation strategies to meet the real-time requirements of power material detection.

[0132] Randomly sampling from the experience replay pool to train the main and target networks of the DQN model can include the following steps:

[0133] Randomly select a batch of experiences from the experience replay pool;

[0134] By combining the main network of the DQN model with the extracted experience, the Q value of the current global state is given;

[0135] By combining the target network of the DQN model with the reward function, the maximum Q value of the next global state is given;

[0136] Determine the Q-value of the current global state and the target loss with the maximum Q-value of the next global state, update the main network parameters in reverse, and periodically update the target network parameters.

[0137] Specifically, the DQN model refers to the Deep Q-Network model, which uses a deep neural network to approximate the Q-function and outputs the Q-value for each action. In this example, such as... Figure 4 As shown, the DQN model includes a main network and a target network. The main network structure includes an input layer, multiple hidden layers, and an output layer. The input layer receives a global state vector fused from state features, task features, and beat constraint features; the input layer dimension is the same as the global state space dimension. Hidden layers can be 2 or 3 layers, with a dimension larger than the input layer dimension. They are used to extract nonlinear features from the global state vector, capturing the matching relationship between the detection / verification task and the analysis unit state. ReLU is used as the activation function to avoid gradient vanishing. The output layer outputs the Q-value of each action; a larger Q-value indicates a higher value for the action. A linear activation function is used, and the output layer dimension is the action space dimension. The target network structure is the same as the main network structure.

[0138] Further, the global state space is a feature space after quantifying and splicing the unit state feature, task feature, and beat constraint feature, and is a collection of all state variables that may affect task allocation. The detection and verification task is all possible allocation results for the action space, which is a three-dimensional discrete space in this example, i.e., the detection and verification task is allocated to the time sequence analysis subunit, the feature comparison subunit, or the semantic analysis subunit.

[0139] It can be understood that the reward function design needs to consider detection accuracy, detection efficiency, load balancing, and beat compliance to guide the DQN model to learn an accurate, efficient, balanced, and compliant task allocation strategy.

[0140] After completing the initial DQN model, the main network weight is randomly initialized, the target network replicates the main network parameters, and the experience replay pool capacity is initialized. The exploration rate, minimum exploration rate, learning rate, discount factor, batch size, target network update frequency, and other hyperparameters are set.

[0141] The current global state, i.e., the global state vector fused by the state feature, task feature, and beat constraint feature, is obtained; the action selection strategy is as follows: The strategy (with The probability of selecting a random action, The probability of selecting an action with the maximum Q value, is a very small constant) selects an action; the action is executed, and the task is allocated to the corresponding analysis unit; the reward is calculated according to the reward function, and the next global state is obtained; the current global state, action, reward, and next global state are stored as an experience in the experience replay pool, and the exploration rate is decayed. The action selection, action execution, and experience storage process is repeated.

[0142] When the experience replay pool capacity exceeds the batch size, random sampling begins, i.e., a batch of experiences is randomly sampled from the experience replay pool. The target Q value, i.e., the maximum Q value of the next global state, is calculated using the target network. The target Q value is specifically represented as:

[0143]

[0144] In the formula, y t is the target Q value, r t is the immediate reward obtained by the reward function, is the discount factor, and the value range is [0, 1], is the Q value of the optimal action in the next global state S t+1 , which represents the maximum expected cumulative reward that can be obtained by executing the optimal action a' in the action space A, is the target network parameter.

[0145] The Q value of the current global state is calculated by the main network as the current Q value. The current Q value is specifically represented as:

[0146]

[0147] In the formula, q t is the current Q value, represents the expected cumulative reward that can be obtained by performing action a t in the current global state S t . is the main network parameter.

[0148] The difference between the target Q value and the current Q value is calculated as the target loss by using the mean square error. The target loss is specifically represented as:

[0149]

[0150] In the formula, L is the target loss, N is the batch size, the number of samples sampled from the experience replay pool each time, is the target Q value of the i-th sample at time step t, is the current Q value of the i-th sample at time step t.

[0151] The loss is minimized by using the Adam optimizer to update the main network parameters in reverse. It is specifically represented as:

[0152]

[0153] In the formula, represents the main network parameter, is the learning rate, is the gradient of the target loss with respect to the main network parameter.

[0154] According to the target network update frequency (in this example, 100 steps), the parameters of the target network are synchronized to the parameters of the main network to avoid fluctuations in Q value estimation. The target network parameter synchronization is specifically represented as:

[0155]

[0156] In the formula, represents the target network parameter, represents the main network parameter.

[0157] Iterative training is performed until the average reward of the model on the validation set no longer improves or the exploration rate decays to the minimum. In this example, the convergence condition is that the average reward improvement of 10 consecutive rounds is less than 0.01 or the exploration rate decays to 0.1.

[0158] In a specific example, as Figure 5As shown, the reward function is constructed based on detection accuracy, detection efficiency, failure penalty, load penalty and takt adherence reward, specifically including the following steps:

[0159] Based on the basic accuracy, task matching degree, task priority, and accuracy basic weight, the detection accuracy is given;

[0160] Based on the takt time, historical processing time, data volume, and remaining takt time, the efficiency factor and takt emergency coefficient are determined, and combined with the efficiency basic weight, the detection efficiency is given;

[0161] According to the task type, the task type weight is obtained, and combined with the task priority, the failure penalty basic weight and the failure flag, the failure penalty is given;

[0162] Based on the current task remaining completion time, the remaining time coefficient is obtained, and combined with the load penalty basic weight and the load state, the load penalty is given;

[0163] Based on the data volume, the data volume coefficient is obtained, combined with the priority, takt adherence flag, and takt adherence basic weight, the takt adherence reward is given;

[0164] Based on the influence coefficient of detection accuracy, detection efficiency, failure penalty, load penalty and takt adherence reward on multi-task parallel processing, the construction of the reward function is completed.

[0165] It can be understood that the influence coefficient of multi-task parallel processing refers to the positive or negative influence on task processing. In this example, the influence coefficient of detection accuracy is "+1", the influence coefficient of detection efficiency is "+1", the influence coefficient of failure penalty is "-1", the influence coefficient of load penalty is "-1", and the influence coefficient of takt adherence reward is "+1".

[0166] Further, the reward function is specifically represented as:

[0167]

[0168] In the formula, r represents the reward function, Accuracy is the detection accuracy, Efficiency is the detection efficiency, FailP is the failure penalty, LoadP is the load penalty, and TaktA is the takt adherence reward.

[0169] In which, the detection accuracy is specifically represented as:

[0170]

[0171] In the formula, Accuracy is the detection accuracy; For the accuracy base weight, cross-validation can be used to determine, and in this example, 0.5 is taken; TaskMatch is the task matching degree, Prio is the priority, and BaseAcr is the base accuracy, which represents the accuracy of the analysis unit in processing a task under ideal conditions. The detection accuracy combines the task matching degree and the priority, guiding the model to preferentially select the analysis unit with high matching degree, improving the detection scene adaptability, and ensuring the detection verification task accuracy.

[0172] The detection efficiency is specifically represented as:

[0173]

[0174] In the formula, Efficiency is the detection efficiency; is the efficiency base weight, and in this example, 0.3 is taken; TaktTime is the tact time, RemainTime is the remaining tact time, HistTime is the historical processing time, DataVolume is the data volume, and MaxDataVolume is the maximum data volume threshold, which can be pre-set; is the tact emergency coefficient, and the smaller the remaining tact time, the larger the tact emergency coefficient, indicating that the task is more urgent; is the efficiency factor, and the shorter the historical processing time and the smaller the data volume, the larger the efficiency factor, indicating that the efficiency is higher. The detection efficiency combines the tact emergency coefficient and the efficiency factor, guiding the model to preferentially select the analysis unit with high efficiency.

[0175] The failure penalty is specifically represented as:

[0176]

[0177] In the formula, FailP is the failure penalty, is the failure penalty base weight, and in this example, 0.2 is taken; is the task type weight, and in this example, 1 is taken for drift verification tasks, 2 is taken for tamper identification tasks, and 1.5 is taken for standard comparison tasks; Prio is the priority, and Failure is the failure flag, which is represented by 1 in this example. The consequences of failure of high-priority tasks (such as tamper identification) are more serious, and heavier penalties are needed; the more dangerous the task type (such as tamper identification), the heavier the failure penalty. The failure penalty guides the model to avoid assigning high-risk tasks to low-accuracy analysis units.

[0178] The load penalty is specifically represented as:

[0179]

[0180] In the formula, LoadP is the load penalty, Load is a load weight, TaskRemainTime is a current task remaining completion time, and MaxRemainTime is a maximum remaining beat time threshold, which can be set in advance; The remaining time coefficient is used. The load penalty guide model reduces task allocation when the analysis unit load is high, reduces the task backlog rate of the high-load analysis unit, and can also shorten the processing time of long-time sequence tasks.

[0181] The beat compliance reward is specifically represented as:

[0182]

[0183] In the formula, TaktA is the beat compliance reward, The beat compliance base weight is specifically represented as: Prio is a priority, DataVolume is a data volume, MaxDataVolume is a maximum data volume threshold, which can be set in advance, and TaktComp is a beat compliance flag. The data volume coefficient is used. The larger the data volume, the larger the data volume coefficient, indicating that the difficulty of complying with the beat is higher. The beat compliance reward guide model preferentially allocates tasks to analysis units that can complete quickly.

[0184] By using the DQN model, expanding the state space, optimizing the reward function, and designing a task allocation strategy driven by task features, state features, and beat constraint features, the detection efficiency, detection resource utilization rate, and detection accuracy are improved.

[0185] In another specific example, as shown in Figure 6 When the detection and verification task is a drift verification task, the corresponding task data is time series data, the detection algorithm for performing the corresponding detection and verification task is used to analyze and process the task data of the detection and verification task, and the verification result of the corresponding detection and verification task is given, which specifically includes:

[0186] Based on the weight function of the fused environmental factors, the local weighted regression smoothing value at each time point is given by processing the time series data;

[0187] Based on the local weighted regression smoothing value, the mean and standard deviation in the sliding window are calculated, and a dynamic threshold is given in combination with the environmental adaptive adjustment coefficient;

[0188] Based on the local weighted regression smoothing value at the current time point, in combination with the dynamic threshold, a drift verification result is given.

[0189] Further, as shown in Figure 7 Based on the weight function of the fused environmental factors, the local weighted regression smoothing value at each time point is given by processing the time series data, which specifically includes the following steps:

[0190] Multiple environmental factors are obtained and normalized. A weighting function is constructed using a multivariate Gaussian kernel function. The weights of the data at each time point within the local window corresponding to the target time are obtained, and a weight diagonal matrix is ​​given.

[0191] For the target time, collect time series data within a local window, construct a design matrix, and give the detection value vector within the local window;

[0192] By combining the detection value vector and the weight diagonal matrix, a weighted least squares model is constructed to solve for the regression coefficients;

[0193] By combining the design vector and regression coefficients at the target time, a locally weighted regression smoothing value is given for each target time.

[0194] To address the environmental dependence (e.g., the temperature-dependent nature of cable insulation resistance) and dynamic characteristic changes (e.g., oil aging) of power materials, an environmentally adaptive dynamic threshold drift detection algorithm is employed. After acquiring time-series detection data of the power materials and corresponding environmental factors (e.g., temperature, humidity), a smoothed value for each time step is calculated using Locally Weighted Regression (LWLR) to eliminate the interference of environmental factors on the detection values. Since environmental factors have different dimensions, they must first be normalized to eliminate dimensional influence. A multivariate Gaussian kernel function is used to construct the weighting function, specifically expressed as:

[0195]

[0196] In the formula, For each t within the local window corresponding to the target time i The weights of the time-series data, where t0 is the target time, i.e., the time at which the smoothed value needs to be calculated, and t i h represents the time of the i-th data point within the local window. t h is the time bandwidth parameter. j Let x be the bandwidth parameter for the j-th environmental factor, m be the number of environmental factors, and x be the bandwidth parameter for the j-th environmental factor. j (t) i ) represents the j-th environmental factor at time t i The normalized value at time x j (t0) is the normalized value of the j-th environmental factor at time t0.

[0197] The weights of the data at each time step relative to the target time step are represented as a diagonal matrix, i.e., a diagonal weight matrix. For several data points surrounding the target time step, a design matrix is ​​constructed, and a detection value vector is provided. Each row of the design matrix is ​​a feature vector of a data point, and the detection value vector contains the detection values ​​within the local window.

[0198] To ensure that data points with higher weights have a greater impact on the regression results, weighted least squares is used to solve for the regression coefficients, i.e.:

[0199]

[0200] wherein, is the regression coefficient, X is the design matrix, X T is the transpose of the design matrix, W is the weight diagonal matrix, and y is the detection value vector.

[0201] The design vector at the target time is multiplied by the regression coefficient to obtain the smoothed value at the target time, i.e.

[0202]

[0203] wherein, is the smoothed value of the detection value at the target time t0, is the regression coefficient, and x0 is the detection value at the target time.

[0204] After smoothing the time series data, the mean and standard deviation in the sliding window are calculated, and the dynamic threshold is given by combining the environmental adaptive adjustment coefficient.

[0205] The upper limit of the dynamic threshold is specifically represented as:

[0206]

[0207] wherein, Upper(t) represents the upper limit of the dynamic threshold, is the mean in the sliding window, is the standard deviation in the sliding window, is the adaptive adjustment coefficient of the jth environmental factor, which can be obtained by optimization through a particle swarm algorithm, m is the number of environmental factors, and t refers to the time t;

[0208] The lower limit of the dynamic threshold is specifically represented as:

[0209]

[0210] wherein, Lower(t) represents the lower limit of the dynamic threshold, is the mean in the sliding window, is the standard deviation in the sliding window, is the adaptive adjustment coefficient of the jth environmental factor, which can be obtained by optimization through a particle swarm algorithm, m is the number of environmental factors, and t refers to the time t. In this example, when the environmental temperature is 30°, when the environmental temperature is 20°, The smoothed value at the current time point is compared with the upper and lower limits of the dynamic threshold to determine whether there is a numerical drift, and the drift verification result is given.

[0211] When the detection verification task is tamper detection, the corresponding task data is the original report data and the real-time detection data, the detection algorithm corresponding to the detection verification task is executed to analyze the task data, and the verification result of the corresponding detection verification task is given, which specifically includes:

[0212] Hash values are calculated for the original report data and the real-time detection data, and hash differences are given;

[0213] The mean difference and the variance ratio are calculated, and the statistical difference is given in combination with the weight coefficient;

[0214] The hash difference and the statistical difference are fused by using a logistic regression model to obtain a tamper probability;

[0215] In combination with a preset probability threshold and the tamper probability, it is determined whether there is tampering, and a tamper identification result is given.

[0216] For each corresponding feature of the original report data and the real-time detection data, the feature value is converted into a string format, and the hash value of each feature is calculated by a hash function. For each feature, the Hamming distance between the corresponding hash values is calculated and normalized, i.e. the feature hash difference is obtained. The feature hash differences of all features are weighted and averaged to obtain the hash difference. The hash difference is used to detect whether the features of the original detection data and the real-time detection data are maliciously modified (such as tampering with the detection value or forging the timestamp).

[0217] The statistical difference is used to determine whether the distribution of the original detection data and the real-time detection data is abnormal (such as a sudden decrease in the mean value of the detection value or a sudden increase in the variance), and the statistical difference is obtained by calculating the mean difference and the variance ratio.

[0218] The hash difference and the statistical difference are normalized to the interval [0, 1] and combined:

[0219]

[0220] In the formula, p is the hash difference and the statistical difference combined feature.

[0221] The combined feature is input into the trained logistic regression model, and the tamper probability is output, which is specifically represented as:

[0222]

[0223] In the formula, represents the probability that the feature combination p is tampered with; is a bias term, is a weight feature, which can be obtained by model training. The tamper probability is compared with the preset probability threshold to determine whether there is tampering, and a tamper verification result is given.

[0224] When the detection verification task is a standard comparison task, a hybrid standard comparison algorithm is used, which first performs semantic analysis and then rule-based reasoning.

[0225] Specifically, for the power material standard clauses, semantic analysis is first performed through natural language processing technology to extract key semantic information and constraint conditions, and a standard semantic graph is constructed. For real-time detection data, semantic analysis is performed to obtain data semantics The similarity between the standard semantics and the data semantics is calculated An attention mechanism-based cosine similarity calculation method is used, and the formula is:

[0226]

[0227] In the formula, is the semantic similarity between the standard semantics and the data semantics, are the semantics of the standard clauses and real-time data, respectively, and n and m are the number of semantic vectors in the standard semantics and data semantics, respectively, are the semantic vectors in the standard semantics and data semantics, respectively, is the attention weight.

[0228] A rule-based reasoning engine is used to determine the matching of the detection data and the rules, and a rule matching score is given. The semantic similarity and the rule matching score are combined in a weighted fusion manner to give a final score, and the standard comparison result is obtained. The final score is specifically represented as:

[0229]

[0230] In the formula, Score is the final score, Sim is the semantic similarity, Mat is the rule matching score; k1 and k2 are weight coefficients, which can be determined by training through power standard comparison cases (such as 10,000 standard clause and detection data matching cases).

[0231] By using the hybrid comparison mechanism of semantic analysis and rule matching, the standard matching accuracy is improved, the processing capability of the semantic analysis subunit for multi-level standards is improved, the specific clauses in the referenced standards are accurately matched, and missed judgments are avoided.

[0232] The analysis unit performs the same material and the same batch of detection verification tasks, and the verification results include numerical data and semantic data. The multi-modal features in the verification results are extracted, and the detection verification conclusion is given in combination with the pre-set structured instructions, which specifically includes:

[0233] The dynamic threshold, drift type, and environmental factors in the drift verification conclusion are extracted, and a drift feature vector is generated;

[0234] The tampering probability, hash difference, statistical difference, and tampering type in the tampering verification conclusion are extracted, and a tampering feature vector is generated;

[0235] The semantic similarity, rule matching score, deviation clause, and standard type in the standard comparison conclusion are extracted, and a standard comparison feature vector is generated;

[0236] The drift feature vector, tampering feature vector, and standard comparison feature vector are fused by using an attention mechanism to generate a comprehensive feature vector;

[0237] The comprehensive feature vector is subjected to nonlinear transformation, and a structured detection verification conclusion is generated in combination with a structured instruction.

[0238] In specific implementation, the multi-modal feature extraction can select Flan-T5, LLaMA-2, and other instruction fine-tuning models supporting multi-modal feature input and having instruction understanding capability, including a feature extraction layer, a feature fusion layer, a fully connected layer, and an output layer. The feature extraction layer performs feature extraction and normalization processing on the input numerical data (such as drift degree, tampering probability, semantic similarity, etc.) and semantic data (such as deviation clause), and converts the verification result data of the analysis unit into a feature vector that can be processed by the model. In this example, for numerical data, principal component analysis is used for dimensionality reduction, and the principal components with cumulative variance contribution rate ≥95% are retained; for semantic data, a pre-trained language model can be used to extract semantic features. The feature fusion layer uses an attention mechanism to fuse the drift feature vector, tampering feature vector, and standard comparison feature vector to obtain a comprehensive feature vector. The number of neurons in the fully connected layer is determined according to the task complexity, and ReLU is used as the activation function to perform nonlinear transformation on the comprehensive feature vector and convert it into a feature representation more suitable for detection verification conclusion generation. The output layer receives the output of the fully connected layer, and through multi-task learning, the output format of the model is constrained by a pre-set structured instruction to generate a structured detection verification conclusion containing three types of information: whether it is qualified, problem explanation, and executable suggestions. The task execution and detection verification conclusion generation process is as shown in Figure 8 In this example, the structured instruction is (“1. Qualified determination; 2. Problem explanation; 3. Rectification suggestion”). Specifically, the Softmax activation function is used to output the probabilities of “qualified” and “unqualified”, and the category with the larger probability is taken as the determination result (for example, “ P 合格 = 0.1, P 不合格=0.9", the determination result is "unqualified"); a natural language description question explanation (such as "violating GB / T 7595-2017 "Transformer Oil Quality Standard in Operation" article 4.2", "cable insulation resistance value (1000MΩ) exceeds the environmental adaptive dynamic threshold (800MΩ) by 20%, temperature 30℃, humidity 60%") is generated by using the Transformer decoder; and executable suggestions (such as "re-filtering the oil sample and re-detecting") are generated by using the Transformer decoder. The training of the instruction fine-tuning model adopts a cross-entropy loss function, and the optimization target is to minimize the prediction error of the detection verification conclusion. The model parameters are updated through back propagation, and the model is iteratively trained until the model converges, thereby improving the accuracy of the generated detection verification conclusion.

[0239] As shown in Figure 9 The application also provides a multi-task parallel power material detection verification device, which adopts the multi-task parallel power material detection verification method, and specifically comprises:

[0240] A preprocessing unit is configured to obtain multi-source power material detection data and current states of each analysis unit, and give integrated data sets and state features, respectively; and extract task data of each detection verification task from the integrated data sets, and give task features;

[0241] A task allocation unit is configured to process the state features and the task features based on a deep learning model complying with a fusion rhythm reward, and give a task parallel processing allocation result;

[0242] An analysis unit is configured to analyze and process the task data of the detection verification task based on the task parallel processing allocation result, and give a verification result corresponding to the detection verification task;

[0243] A conclusion output unit is configured to extract multi-modal features in the verification result, combine preset structured instructions, and give a detection verification conclusion, thereby completing the verification of the power material detection.

[0244] Although the preferred embodiments of the application have been described, those skilled in the art, once they know the basic creative concept, can make additional changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the application. Obviously, those skilled in the art can make various modifications and variations to the application without departing from the spirit and scope of the application. Thus, if these modifications and variations of the application fall within the scope of the claims of the application and their equivalents, the application also intends to include these modifications and variations.

Claims

1. A method for verifying multi-task parallel electric power material detection, applied to a verification device, the verification device comprising an analysis unit, characterized in that, Specifically comprising the following steps: Obtain multi-source electric power material detection data and the current state of each analysis unit, and give integrated data set and state characteristics respectively; Extract task data of each detection verification task from the integrated data set, and give task characteristics; Build a DQN model, define a global state space, take all feasible allocation results of parallel processing of detection verification tasks as an action space, give detection accuracy based on basic accuracy, task matching degree, task priority, and accuracy basic weight; Determine efficiency factor and beat emergency coefficient based on beat time, historical processing time, data volume, and remaining beat time, and give detection efficiency combined with efficiency basic weight, wherein the beat time refers to the maximum time limit allowed for completing a single detection verification task or a batch of detection tasks, the remaining beat time refers to the remaining available time after the beat time minus the used beat time, and the used beat time refers to the time spent from the start of the task execution to the current time; obtain task type weight according to task type, and give failure penalty combined with task priority, failure penalty basic weight, and failure flag; obtain remaining time coefficient based on the current task remaining completion time, and give load penalty combined with load penalty basic weight and load state; obtain data volume coefficient based on data volume, and give beat compliance reward combined with priority, beat compliance flag, and beat compliance basic weight; build a reward function based on the influence coefficients of detection accuracy, detection efficiency, failure penalty, load penalty, and beat compliance reward on multi-task parallel processing; Obtain state characteristics of the analysis unit, task characteristics of the detection verification task, and beat constraint characteristics, and give the current global state, wherein the beat constraint characteristics include beat time, used beat time, and remaining beat time; select an action based on the current global state using an exploration strategy, and give the reward after executing the action and the next global state; store the current global state, action, reward, and next global state as experience in an experience replay pool; repeat the process of selecting an action, executing an action, and storing experience until the capacity of the experience replay pool exceeds a preset value; randomly sample from the experience replay pool, train the main network and target network of the DQN model, and give the final deep learning model; Process the state characteristics and task characteristics based on the deep learning model fused with the beat compliance reward, and give the task parallel processing allocation result; Each analysis unit analyzes and processes the task data of the detection verification task based on the task parallel processing allocation result, and gives the verification result of the corresponding detection verification task; Extract multi-modal features from the verification result, and give detection verification conclusions combined with the preset structured instructions, and complete the verification of electric power material detection.

2. The method of claim 1, wherein the method further comprises: Obtain the current state of each analysis unit, and give state characteristics, including: Determine the current task backlog and the current task remaining completion time of each analysis unit based on the current state of each analysis unit; Obtain the best processing task type and historical processing time corresponding to each analysis unit; Give the task matching degree based on the best processing task type corresponding to each analysis unit and the current detection verification task type. determine a load state based on a current task backlog of each analysis unit; normalize the task matching degree, the load state, the current task remaining completion time, and the historical processing time to give a state feature.

3. The method of claim 1, wherein the method further comprises: extract task data of each detection verification task from the integrated dataset to give a task feature, including: determine a data type according to the task data in the integrated dataset; give a corresponding task type based on the data type; determine the priority of different types of tasks according to the given task type and the preset conditions; extract each type of data in the integrated dataset to give the data volume corresponding to each task type; give a task feature based on the task type, the data volume, and the priority.

4. The method of claim 1, wherein the method further comprises: randomly sample from the experience replay pool to train the main network and the target network of the DQN model, specifically including the following steps: randomly extract a batch of experiences from the experience replay pool; give the Q value of the current global state through the main network of the DQN model combined with the extracted experiences; give the maximum Q value of the next global state through the target network of the DQN model combined with the reward function; determine the target loss of the Q value of the current global state and the maximum Q value of the next global state, update the parameters of the main network in reverse, and update the parameters of the target network regularly.

5. The method of claim 1, wherein the method further comprises: The reward function is specifically represented as: ; In the formula, r represents the reward function, Accuracy is the detection accuracy, Efficiency is the detection efficiency, FailP is the failure penalty, LoadP is the load penalty, and TaktA is the beat adherence reward.

6. The method of claim 1, wherein the method further comprises: The detection verification task includes a drift verification task, and the task data corresponding to the detection verification task includes time series data; analyze and process the task data of the detection verification task to give the verification result of the corresponding detection verification task, specifically including: process the time series data based on the weight function of the fused environmental factors to give the local weighted regression smoothing value of each time point; based on the local weighted regression smoothing value, calculate the mean and standard deviation in the sliding window, and give the dynamic threshold combined with the environmental self-adaptive adjustment coefficient; based on the local weighted regression smoothing value of the current time point, give the drift verification result combined with the dynamic threshold.

7. The method of claim 6, wherein the method further comprises: Process the time series data based on the weight function of the fused environmental factors to give the local weighted regression smoothing value of each time point, specifically including the following steps: obtain multiple environmental factors and construct a weight function combined with a multivariate Gaussian kernel function to determine the weight of each time point data in the local window corresponding to the target time point, and give the weight diagonal matrix; determine the time series data in the local window corresponding to the target time point, construct a design matrix, and give the detection value vector in the local window; construct a weighted least squares model combined with the detection value vector and the weight diagonal matrix to solve the regression coefficient; give the local weighted regression smoothing value of each target time point combined with the design vector of the target time point and the regression coefficient.

8. A multi-task parallel power material detection verification device, characterized in that, Use the multi-task parallel power material detection verification method of any one of claims 1-7, comprising: The preprocessing unit is configured to obtain multi-source electric power material detection data and current states of each analysis unit, and give an integrated data set and a state feature respectively; extract task data of each detection verification task from the integrated data set, and give a task feature; The task allocation unit is configured to process the state feature and the task feature based on a deep learning model complying with a fusion beat reward, and give a task parallel processing allocation result; The analysis unit is configured to analyze and process the task data of the detection verification task based on the task parallel processing allocation result, and give a verification result corresponding to the detection verification task; The conclusion output unit is configured to extract multi-modal features in the verification result, combine a preset structured instruction, and give a detection verification conclusion to complete verification of electric power material detection.

Citation Information

Patent Citations

  • Task scheduling method and device, electronic equipment and storage medium

    CN112330014A

  • DQN-based intelligent factory job shop scheduling method and system

    CN119398426A

  • Multi-mechanical-arm flexible production line scheduling method and system based on visual language model

    CN120611244A