Data maintenance method and system based on big data and artificial intelligence

Through data maintenance methods based on big data and artificial intelligence, data maintenance discriminant agents are used to perform multiple data maintenance and discriminant operations, solving the problems of low data maintenance efficiency and difficulty in training in the existing technology, and achieving efficient and real-time data maintenance and training in the intelligent body.

CN119670018BActive Publication Date: 2025-05-13SICHUAN BODA ZHENGHENG INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510148465.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-13
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

When facing complex and diverse data, the existing data maintenance methods have low operation efficiency and are difficult to meet the needs of large-scale data real-time processing; traditional data processing agents have difficulties in training and debugging, lack flexibility and versatility, and are slow in training, so they cannot quickly converge to the optimal solution.

Method used

The data maintenance method based on big data and artificial intelligence is adopted, and the target terminal maintenance detection data is obtained, and a data maintenance discrimination agent is used to perform a variety of data maintenance and discrimination operations to generate a data maintenance discrimination agent. The method includes acquiring several target data processing agents, fusing their characterization information processing networks, generating a first data processing agent, and debugging them to obtain a data maintenance discriminant agent.

Benefits of technology

It improves the real-time and efficiency of data maintenance, enhances the training efficiency and performance of data maintenance and discrimination agents, helps to efficiently use the agent for data maintenance and discrimination operations, and ensures the timeliness and accuracy of data maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670018B_ABST
    Figure CN119670018B_ABST
Patent Text Reader

Abstract

The present invention provides a data maintenance method and system based on big data and artificial intelligence. By acquiring several target data processing agents, any target data processing agent includes multiple representation information processing networks, and the network module architecture composed of the multiple representation information processing networks included in each target data processing agent is consistent. Any target data processing agent is used to perform a data maintenance discrimination operation on the target terminal maintenance detection data; the representation information processing networks in the same architecture node in the network module architecture of the several target data processing agents are fused, and the representation information processing networks of a node are fused to obtain a first fused network; based on each first fused network, a first data processing agent is generated; and the first data processing agent is debugged to obtain a data maintenance discrimination agent. The present invention can efficiently complete the training of the data maintenance discrimination agent to ensure the timeliness of the data maintenance discrimination operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a data maintenance method and system based on big data and artificial intelligence. Background Art

[0002] In today's digital age, the importance of data is self-evident. It is widely used in various fields, including finance, medical care, communications, industrial manufacturing, etc. With the rapid development of information technology, the amount of data has exploded, and data maintenance and management have become crucial tasks. Existing data maintenance methods have many shortcomings when facing complex and diverse data. On the one hand, the efficiency of data maintenance and discrimination operations is low, and it is difficult to meet the needs of real-time processing of large-scale data. Traditional methods often require a lot of time and computing resources to complete operations such as data accuracy identification, integrity identification, and timeliness identification, which greatly reduces the timeliness of data maintenance and affects the normal operation of the system and the timeliness of decision-making. On the other hand, existing data processing agents have difficulties in training and debugging. Traditional data processing agents lack flexibility and versatility when facing different types of data maintenance tasks, and it is difficult to quickly adapt to new data characteristics and maintenance requirements. Moreover, the adjustment of network weights and biases during training is complex, resulting in slow training of agents and failure to quickly converge to the optimal solution, which increases development costs and time cycles. In addition, it is difficult for existing data maintenance and discrimination agents to effectively integrate the advantages of different data processing agents. Different data processing agents are usually designed for specific data maintenance tasks and are relatively independent of each other. They cannot fully utilize the advantages of each agent in different aspects, thus affecting the overall effect and accuracy of data maintenance. Summary of the invention

[0003] In view of this, the present invention provides a data maintenance method and system based on big data and artificial intelligence. The technical solution of the present invention is implemented as follows:

[0004] On the one hand, the present invention provides a data maintenance method based on big data and artificial intelligence, the method comprising: obtaining target terminal maintenance detection data; performing at least one data maintenance discrimination operation on the target terminal maintenance detection data based on a data maintenance discrimination agent to obtain each data maintenance discrimination result. The data maintenance discrimination agent is debugged by the following steps to obtain: obtaining a number of target data processing agents, any target data processing agent includes a plurality of representation information processing networks, and the network module architecture composed of the plurality of representation information processing networks included in each target data processing agent is consistent, and any target data processing agent is used to perform a data maintenance discrimination operation on the target terminal maintenance detection data; fusing the representation information processing networks at the same architecture node in the network module architecture of the plurality of target data processing agents, and fusing the representation information processing networks of a node to obtain a first fused network; generating a first data processing agent based on each first fused network; debugging the first data processing agent to obtain a data maintenance discrimination agent, and the data maintenance discrimination agent is used to perform at least one data maintenance discrimination operation on the target terminal maintenance detection data among the data maintenance discrimination operations corresponding to each target data processing agent.

[0005] In a second aspect, the present invention provides a computer system, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps in the above-described method when executing the program.

[0006] Technical effect of the present invention: In the embodiment of the present invention, each target data processing agent includes multiple representation information processing networks, and the network module architecture composed of the multiple representation information processing networks included in each target data processing agent is consistent. Based on the fusion of each representation information processing network at the same architecture node in the network module architecture, each first fusion network is obtained, and the network weights and biases of the representation information processing network are transferred to the first fusion network, thereby promoting the training speed of the first fusion network. Based on each first fusion network, a first data processing agent is generated, and the first data processing agent is debugged to obtain a data maintenance judgment agent, so that the efficiency of training to obtain a data maintenance judgment agent is improved, which helps to efficiently use the data maintenance judgment agent to perform data maintenance judgment operations and ensure the real-time nature of data maintenance. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present invention and, together with the specification, are used to explain the technical solutions of the present invention.

[0008] Figure 1A schematic diagram of the implementation flow of a data maintenance method based on big data and artificial intelligence provided in an embodiment of the present invention.

[0009] Figure 2 A debugging flow chart of a data maintenance judgment agent provided in an embodiment of the present invention.

[0010] Figure 3 A hardware entity schematic diagram of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0011] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present invention. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present invention.

[0012] The embodiment of the present invention provides a data maintenance method based on big data and artificial intelligence, which can be executed by a processor of a computer system. The computer system can refer to a device with data processing capabilities such as a server, a laptop, a tablet computer, and a desktop computer.

[0013] Figure 1 A schematic diagram of an implementation flow of a data maintenance method based on big data and artificial intelligence provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method includes:

[0014] Step S1: Obtain target terminal maintenance detection data.

[0015] The target terminal maintenance detection data is the data detected by the remote operation and maintenance terminal during the operation and maintenance work for government and enterprise users. The operation and maintenance scenarios of government and enterprise users are diverse and complex, covering various business systems and equipment, which constitute the scope of target terminals. For example, in government departments, it may involve government office systems, data storage servers, etc.; in the enterprise field, it may include enterprise resource planning (ERP) systems, production equipment monitoring terminals, etc. The computer system obtains maintenance and detection data related to these different types of target terminals.

[0016] Taking the government office system as an example, the computer system can acquire data by deploying data acquisition modules at the key nodes of the system. These key nodes include the server interface, the database interaction port, etc. The acquisition module can monitor the system's operating status parameters in real time, such as CPU usage, memory usage, network bandwidth, etc. These data all belong to the category of target terminal maintenance detection data. Specifically, for CPU usage, the acquisition module can read the relevant indicator data provided by the system at regular time intervals (such as one minute). The technical means is to obtain data by calling the system information query interface provided by the operating system, such as using the top command in the Linux system or the relevant API interface of the task manager in the Windows system.

[0017] For the enterprise's ERP system, the computer system can be connected to the ERP system with the help of a specially developed adapter. The adapter can extract specific data from the ERP system database according to predetermined rules, such as transaction records in the order processing process, goods in and out of the warehouse data in the inventory management module, etc. These data reflect the business processing status of the ERP system during operation, which is of great significance for maintaining and optimizing the system.

[0018] Step S2: Based on the data maintenance judgment agent, at least one data maintenance judgment operation is performed on the target terminal maintenance detection data to obtain various data maintenance judgment results.

[0019] When the computer system executes step S2, it performs at least one data maintenance judgment operation on the target terminal maintenance detection data based on the data maintenance judgment agent, and then obtains various data maintenance judgment results. As a machine learning model, the data maintenance judgment agent has the ability to judge whether the data needs maintenance from different dimensions. This process is achieved through the prediction part of the agent.

[0020] The target terminal maintenance detection data covers all kinds of information detected by the remote operation and maintenance terminals of government and enterprise users during operation and maintenance work. For example, for a large government and enterprise unit, its remote operation and maintenance terminal can monitor various parameters of the server in real time, including hardware-related data such as CPU usage, memory usage, disk I / O rate, network bandwidth utilization, as well as software-level data such as application response time, transaction processing success rate, and error logs. These data form the basis for the computer system to perform data maintenance and judgment operations.

[0021] The data maintenance judgment operation is the process of judging whether the data needs to be maintained from different dimensions, which relies on the prediction part of the data maintenance judgment agent. Taking the server CPU usage rate as an example, the agent will combine historical data and a preset reasonable range to judge whether the current data is in a normal state. If the CPU usage rate is higher than 80% for a long time, the agent may determine that the data needs maintenance, because the continuous high load may cause the server performance to degrade or even fail. For another example, for the response time of the application, if the agent finds that it exceeds the historical average and fluctuates greatly, it can be considered that this part of the data needs maintenance, because this may affect the user experience and indicate that there are potential problems in the system.

[0022] The prediction part of the data maintenance discrimination agent is based on the feature processing results. In actual operation, the agent first extracts and converts the features of the target terminal maintenance detection data. For example, for the disk I / O rate data of the server, the agent can calculate its average value, standard deviation, change trend and other features in different time periods. These features will be used as the input of the agent's prediction. The agent uses complex algorithms and pre-trained parameters to determine whether the data needs maintenance. Assuming that the agent has been trained to understand that when the standard deviation of the disk I / O rate exceeds a certain threshold, the stability of the data may be affected and maintenance operations are required, then when the new disk I / O rate data is received and the corresponding features are extracted, the agent can make accurate predictions based on these features.

[0023] During the agent training phase, a large amount of historical target terminal maintenance detection data is used as a training set. These data are divided into training data and validation data. The training data is used to adjust the parameters of the agent, and the validation data is used to evaluate the performance of the agent. During the training process, the agent will perform feature processing based on the input data, and then output the prediction results through the prediction part. The prediction results are compared with the known correct labels (that is, the actual situation of whether the data needs maintenance), and the loss function (such as the cross entropy loss function) is used to measure the difference between the prediction results and the true labels. Then, the optimization algorithm (such as the stochastic gradient descent algorithm) is used to adjust the agent parameters to reduce the value of the loss function and improve the prediction accuracy of the agent.

[0024] When performing data maintenance judgment operations, the computer system inputs the target terminal maintenance detection data obtained in real time into the trained data maintenance judgment agent. The agent first performs feature processing on the data, which may involve normalization of the data to ensure that data with different features are processed on the same scale. Then, the feature-processed data is passed to the prediction part of the agent, which outputs each data maintenance judgment result through a series of calculations and mappings (such as matrix multiplication and nonlinear transformation). For example, for a set of target terminal maintenance detection data containing multiple parameters of the server, the agent may output multiple judgment results after processing, such as CPU usage data needs to be maintained, memory usage data does not need to be maintained temporarily, etc.

[0025] These data maintenance and discrimination results are of great guiding significance for the operation and maintenance work of government and enterprise users. Operation and maintenance personnel can take corresponding measures in a timely manner based on these results. For example, for CPU usage data that needs to be maintained, the CPU load can be reduced by optimizing application algorithms, increasing server resources, etc.; for data that does not need to be maintained, the monitoring status can continue to be maintained. In this way, the computer system uses the data maintenance discrimination agent to perform discrimination operations on the target terminal maintenance detection data, which can effectively ensure the efficiency and stability of the operation and maintenance work of government and enterprise users, ensure the normal operation of the system, reduce the probability of potential failures, and improve the overall operation and maintenance management level.

[0026] Please refer to Figure 2 , the data maintenance discrimination agent is debugged using the following steps to obtain:

[0027] Step S10: Acquire several target data processing agents, any target data processing agent includes multiple representation information processing networks, and the network module architecture composed of the multiple representation information processing networks included in each target data processing agent is consistent. Any target data processing agent is used to perform a data maintenance judgment operation on the target terminal maintenance detection data.

[0028] Any target data processing agent contains multiple representation information processing networks. The representation information processing network is an important component for feature extraction and conversion of target terminal maintenance detection data, aiming to convert raw and complex data into more representative and valuable feature representations for subsequent accurate data maintenance and discrimination operations. Taking the remote operation and maintenance terminal detection data for government and enterprise users as an example, it is assumed that the detection data contains multiple information such as equipment operation status parameters, network connection information, and business processing traffic. Among them, the equipment operation status parameters include CPU usage, memory usage, disk read and write speed, etc.; network connection information includes network bandwidth, connection stability, packet loss rate, etc.; business processing traffic involves the number of business requests per unit time, response time, etc. The representation information processing network will convert these different types of data into vectors or matrices that can reflect the essential characteristics of the data through specific network structures and algorithms. For example, for CPU usage data, the representation information processing network can generate a feature vector that can represent the CPU usage by calculating the average value, change trend, and comparison relationship with historical data over a period of time.

[0029] The network module architecture composed of multiple representation information processing networks contained in each target data processing agent is consistent. This feature provides convenient conditions for the subsequent agent fusion. For example, suppose there are three target data processing agents M1, M2 and M3. M1 is used to maintain and judge the data of the device hardware status, M2 focuses on data maintenance and judgment in network security, and M3 mainly handles data maintenance and judgment of business processing efficiency. Although their judgment focuses are different, the network module architecture composed of the internal representation information processing network is similar. They all have an input layer, an intermediate hidden layer and an output layer. The input layer is responsible for receiving the target terminal maintenance detection data, the intermediate hidden layer performs deep processing on the data through various nonlinear transformations and feature combinations, and the output layer outputs the processed feature representation. Taking the input layer as an example, it will unify and preprocess the different types of detection data so that it can adapt to the processing of the subsequent network layer; the intermediate hidden layer can use the convolution operation in the convolutional neural network (CNN) or the loop structure in the recurrent neural network (RNN) to extract and fuse the data; the output layer outputs one or more feature vectors based on the processing results of the intermediate layer as the feature representation of the input data of the agent.

[0030] Each target data processing agent is used to perform a data maintenance judgment operation on the target terminal maintenance detection data. For example, agent M1 will use its own multiple representation information processing networks to process the target terminal maintenance detection data related to the device hardware status. It will comprehensively analyze indicators such as CPU usage, memory usage, disk read and write speed, and determine whether the device hardware status is normal through specific network weights and bias settings, and then determine whether the relevant data needs maintenance. If the CPU usage continues to be higher than a certain threshold, or the memory usage reaches a certain proportion and remains high for a long time, agent M1 may determine that this part of the data needs maintenance to avoid performance problems of the device. Agent M2 will analyze network connection information and network security-related data, and identify whether there is a network attack or abnormal behavior by detecting network traffic patterns, port access conditions, etc., so as to determine whether network security data needs maintenance. Agent M3 will focus on business processing traffic data, evaluate business processing efficiency by analyzing indicators such as the number of business requests and response time, and determine whether relevant data needs maintenance.

[0031] The computer system can obtain the target terminal maintenance detection data from the remote operation and maintenance terminal through the data acquisition module and store it in the database. For the acquisition of the target data processing agent, the computer system can select a suitable agent from the pre-built agent library. These target data processing agents in the agent library are obtained by training and optimization through a large amount of historical data. During the training process, supervised learning or unsupervised learning methods can be used. For supervised learning, labeled training data is used, such as for the equipment hardware status data, whether the data needs maintenance is marked; for unsupervised learning, training is performed through the intrinsic structure and pattern of the data. In terms of agent construction, deep learning frameworks such as TensorFlow or PyTorch can be used. For example, a target data processing agent based on CNN is constructed, and the network structure is constructed using convolutional layers, pooling layers, and fully connected layers. The network weights and biases are adjusted through optimization algorithms such as back propagation algorithms and stochastic gradient descent, so that the agent can accurately perform corresponding data maintenance discrimination operations on the data. At the same time, in order to ensure the generalization ability and stability of the agent, regularization techniques such as L1 and L2 regularization can be used.

[0032] Step S20: Fusing the representation information processing networks in the same architecture node in the network module architecture of several target data processing agents, and fusing the representation information processing networks of one node to obtain a first fused network.

[0033] In multiple target data processing agents, since their network module architectures are consistent, the network layers at corresponding positions are the same architecture nodes. For example, suppose there are three target data processing agents A, B, and C, all of which have similar network structures, including input layers, middle layers, and output layers, and the middle layers are further subdivided into several sublayers of the same level. Then in these three agents, the representation information processing network at the second sublayer position of the middle layer belongs to the same architecture node. Although these representation information processing networks are in different target data processing agents, they have certain similarities in structure and function, and all perform specific feature extraction and conversion operations on the data passed from the previous layer.

[0034] The computer system fuses the representation information processing networks of these nodes with the same architecture. The purpose of fusion is to combine the advantages of each agent at the same node to obtain a more comprehensive and representative feature representation. For example, agent A may have an advantage in processing device hardware status data, and its representation information processing network at a node with the same architecture can extract features of CPU usage more accurately; agent B performs well in network security data processing, and the representation information processing network of this node has a strong ability to recognize abnormal patterns in network traffic; agent C is good at analyzing data related to business processing efficiency, and the representation information processing network of this node can effectively extract key features of business request response time. Through fusion, the computer system can integrate the advantages of these different agents at the same node, so that the first fused network has more powerful feature processing capabilities.

[0035] Taking the specific fusion operation as an example, assume that the representation information processing network of the same architecture node extracts features from the data and outputs a feature vector. When fusion, the computer system can use the weighted average method. Assume that there are three representation information processing networks N1, N2, and N3 in the same architecture node, which come from the target data processing agents A, B, and C respectively. After processing the same set of target terminal maintenance detection data, the output feature vectors are V1, V2, and V3 respectively. The computer system assigns different weights w1, w2, and w3 to each representation information processing network according to its performance or importance in past tasks (the sum of the weights is 1, that is, w1+w2+w3=1). The fused feature vector V is calculated by the formula V=w1*V1+w2*V2+w3*V3, and this feature vector will be used as the output of the first fusion network at this node. This weighted average method can reasonably integrate their information according to the performance differences of each representation information processing network, making the fusion result more reliable and effective. In addition to the weighted average method, splicing can also be used for fusion. The computer system stitches together the feature vectors output by each representation information processing network in order to form a new feature vector with a higher dimension. For example, the dimension of the feature vector V1 output by N1 is 10, the dimension of V2 output by N2 is 15, and the dimension of V3 output by N3 is 12. The dimension of the new feature vector obtained after stitching is 10+15+12=37. This stitching method can retain the original feature information of each representation information processing network and provide a richer data source for subsequent intelligent agent processing.

[0036] Step S30: Generate a first data processing agent according to each first fusion network.

[0037] When the computer system executes step S30, the core task is to generate a first data processing agent based on each first fusion network. The first fusion network is obtained by fusing the representation information processing networks of the same architecture nodes in multiple target data processing agents. It integrates the advantage information of different agents at specific nodes, providing a basis for generating a first data processing agent with more powerful functions and better performance.

[0038] First, for any first fusion network, the computer system generates a first reconstruction network based on it and the initial adjustment network. The initial adjustment network plays a key adjustment role in this process. It is used to adjust the output of the first reconstruction network so that the network output is more in line with actual needs and data characteristics. For example, when processing the remote operation and maintenance terminal data of government and enterprise users, it is assumed that one of the first fusion networks is responsible for processing the data characteristics related to the hardware status of the device. The first fusion network may have effectively integrated information such as CPU usage and memory occupancy. The initial adjustment network will adjust the output of the first reconstruction network according to the dynamic changes of the data and the preset rules. For example, when it is detected that the device is in a high-load operation stage, the initial adjustment network can enhance certain feature weights that reflect the pressure of the device, so that the output of the first reconstruction network can more accurately reflect the data representation of the current state of the device. Specifically, the process of generating the first reconstruction network is a process of organically combining the first fusion network with the initial adjustment network. The computer system can use a variety of methods to achieve this combination, such as realizing information fusion through matrix multiplication and addition operations. Assume that the feature matrix output by the first fusion network is F, and the adjustment matrix generated by the initial adjustment network is R. Through matrix multiplication, F and R are multiplied to obtain a new matrix M, and then M and F are added through addition (which can be expressed as F'=M+F), and the result F' is used as the output of the first reconstruction network. This operation method allows the initial adjustment network to effectively adjust the output of the first fusion network, making the information output by the first reconstruction network richer and more accurate.

[0039] If F is a matrix representing the characteristics of the device hardware status, each row represents a different hardware indicator (such as CPU usage, memory usage, etc.), and each column represents a different time point or sample. The R matrix is ​​generated based on the context information, historical data, and preset strategies of the device operation, and is used to adjust the weights of each element in the F matrix. Through the above operations, the F' matrix output by the first reconstruction network can better reflect the real-time status of the device hardware status and potential problems.

[0040] Next, the computer system generates a first data processing agent based on each first reconstruction network. This process is to organically integrate multiple first reconstruction networks to build a unified agent that can handle multiple data maintenance and discrimination tasks. For example, in addition to the first reconstruction network that processes the hardware status data of the device, there are also first reconstruction networks that process network security data, business processing efficiency data, and other different aspects. The computer system will combine these first reconstruction networks with different functions according to a certain logical structure.

[0041] One feasible combination method is to use a cascade method. Connect each first reconstruction network in sequence, and use the output of the previous first reconstruction network as the input of the next first reconstruction network. Through this cascade structure, the first data processing agent can gradually conduct in-depth mining and processing of data features in different aspects. For example, the first reconstruction network that processes the device hardware status data first performs preliminary processing on the relevant data and outputs a feature vector that reflects the device hardware status. Then, this feature vector is input into the first reconstruction network that processes network security data. The network will further fuse and transform the input feature vector in combination with the network security information processed by itself, and output a more comprehensive feature representation. And so on, after cascade processing of multiple first reconstruction networks, a first data processing agent that can comprehensively reflect multiple data features is finally generated.

[0042] Step S40: Debug the first data processing agent to obtain a data maintenance determination agent, which is used to perform at least one data maintenance determination operation corresponding to each target data processing agent on the target terminal maintenance detection data.

[0043] When the computer system executes step S40, the first data processing agent is debugged to obtain a data maintenance judgment agent, which is used to perform at least one data maintenance judgment operation corresponding to each target data processing agent on the target terminal maintenance detection data. The debugging process is intended to optimize the performance of the agent so that it can process and analyze data more accurately and provide reliable support for subsequent data maintenance work.

[0044] During the debugging process, the computer system uses the terminal maintenance detection training data and prior labels of the target data processing agent. The terminal maintenance detection training data is a large amount of historical data that has been collected and organized. These data cover various status information of the target terminal in different operation and maintenance scenarios. For example, for the remote operation and maintenance terminals of government and enterprise users, the training data may include data records of the server's CPU usage, memory usage, network bandwidth consumption, application response time, etc. in different time periods. Prior labels are known results obtained after performing data maintenance judgment operations on these training data, which are used to annotate information such as whether the data needs maintenance and the type or direction of maintenance. For example, if the server CPU usage rate continues to be too high for a period of time, resulting in a decrease in system performance, the corresponding prior label will clearly indicate that the data needs maintenance, and may prompt the need to optimize related applications or increase server resources and other maintenance measures.

[0045] The computer system trains and optimizes the first data processing agent based on these training data and prior labels. Specifically, by inputting the terminal maintenance detection training data into the first data processing agent, the agent processes the data and generates reasoning results. This reasoning process extracts, converts and analyzes the input data based on the network structure and parameter settings within the agent, thereby making a judgment on whether the data needs maintenance. For example, the agent may judge whether the current server's memory usage is within a normal range and whether memory cleaning or expansion operations are required based on the patterns and rules learned from the training data. Then, the computer system compares the reasoning results generated by the agent with the prior labels and evaluates the performance of the agent by calculating the difference between the two. Commonly used evaluation indicators include accuracy, recall rate, F1 value, etc.

[0046] In order to adjust the agent, the computer system uses an optimization algorithm to update the parameters of the agent. Possible optimization algorithms include stochastic gradient descent (SGD) and its variants Adagrad, Adadelta, Adam, etc. Taking the stochastic gradient descent algorithm as an example, its core idea is to randomly select a small batch of training data in each iteration, calculate the gradient of the loss function on the small batch of data, and then update the parameters of the agent according to the gradient direction. Through continuous iteration, the value of the loss function is gradually reduced, so that the parameters of the agent gradually converge to a better solution, thereby improving the performance of the agent.

[0047] During the debugging process, the computer system can also use cross-validation and other techniques to ensure the generalization ability of the agent. For example, the training data is divided into multiple subsets, and a part of the subsets are used as the training set each time, and the rest of the subsets are used as the validation set. By training and evaluating on different combinations of training sets and validation sets, we can have a more comprehensive understanding of the performance of the agent under different data distributions and avoid overfitting or underfitting of the agent.

[0048] In addition, the computer system can also adjust and optimize the structure of the agent. For example, increase or decrease the number of neurons in certain hidden layers, adjust the number of layers of the network, or try different activation functions. The activation function is used to introduce nonlinear transformations so that the agent can learn complex patterns in the data. Viable activation functions include ReLU (Rectified Linear Unit), Sigmoid, Tanh, etc. By constantly trying different structures and parameter settings, the computer system gradually finds the agent configuration that is most suitable for the data maintenance and discrimination task, and finally obtains a data maintenance and discrimination agent with good performance. This data maintenance and discrimination agent can accurately perform a variety of data maintenance and discrimination operations on the target terminal maintenance and detection data, providing strong support for the operation and maintenance work of government and enterprise users, helping them to promptly discover problems in the data and take corresponding maintenance measures to ensure the stable operation of the system and the accuracy of the data.

[0049] As an implementation mode, any representation information processing network includes several sub-processing networks, any sub-processing network includes network weights and biases of multiple feature paths, and a first fusion network includes several sub-fusion networks; in step S20, each representation information processing network of a node is fused to obtain a first fusion network, including:

[0050] Step S21: for any sub-processing network included in any characterization information processing network, determine the network weight and bias of the selected characteristic path from the network weights and biases of each characteristic path included in any sub-processing network;

[0051] Step S22: for any sub-processing network included in each characterization information processing network of a node, a sub-fusion network is generated according to the network weights and biases of the selected feature paths corresponding to any of the sub-processing networks.

[0052] In step S21, the computer system determines the network weight and bias of the selected feature path from the network weights and biases of each feature path of any sub-processing network included in any characterization information processing network. This operation is the basis of the subsequent fusion process, and its purpose is to screen out the key parts for data characterization from many feature paths.

[0053] First, the computer system performs a standardization operation on the network weights and biases of each feature path included in any sub-processing network to obtain the standardized network weights and biases of each feature path. A feature path is a feature channel. The standardization operation is to eliminate the dimensional differences of different feature paths so that they can be compared and analyzed on the same scale. Taking a sub-processing network that processes the temperature data of remote operation and maintenance terminal equipment of government and enterprise users as an example, the sub-processing network may contain multiple feature paths, each of which corresponds to different sensor data or data processing dimensions. Assume that one of the feature paths focuses on the temperature change rate of a certain area of ​​the device, and the other feature path focuses on the average value of the overall temperature of the device. The numerical ranges and physical meanings of these two feature paths are different. The computer system uses a standardization formula, such as the Z-score standardization formula. Through this formula, the network weights and biases of different feature paths are converted into standard data with a mean of 0 and a standard deviation of 1, so that they can be subsequently operated on a unified scale.

[0054] Next, based on the standardized network weights and biases of each feature path, the computer system determines the significance measure of each feature path. The significance measure can be understood as a score for the importance of each feature path in data representation. This process is implemented through a specific algorithm or agent. For example, a method based on information gain can be used to calculate the significance measure. Information gain measures the amount of information that a feature can add to the classification task. For a binary classification problem (such as determining whether the temperature of the device is normal), the information gain calculation formula is: IG(X, Y)=H(X) - H(X|Y), where H(X) is the entropy of feature X, and H(X|Y) is the conditional entropy of feature X under the condition of known category Y. The entropy calculation formula is: ,in, is the probability of feature X taking the ith value. Using these formulas, the computer system can calculate the information gain of each feature path as a measure of significance. In the processing of device temperature data, those feature paths that can more accurately distinguish between normal and abnormal device temperature states will have higher information gains, which means that their significance measures are greater.

[0055] Finally, in each feature path, the computer system determines the selected feature paths whose significance metric is greater than or equal to the preset value, and obtains the network weights and biases of these selected feature paths. The preset value is a threshold determined based on experience or previous experiments, which is used to screen out the most important feature paths. For example, after calculation, the preset value is set to 0.5. If the significance metric of a feature path is 0.6, which is greater than the preset value, the feature path is selected as the selected feature path, and its corresponding network weights and biases will be extracted by the computer system. The network weights and biases of these selected feature paths will become the key information for the subsequent generation of the sub-fusion network.

[0056] In step S22, for any sub-processing network included in each characterization information processing network of a node, the computer system generates a sub-fusion network based on the network weights and biases of the selected feature paths corresponding to any of the sub-processing networks.

[0057] Specifically, the computer system integrates the network weights and biases of the selected characteristic paths in the sub-processing networks from different representation information processing networks but belonging to the same node. For example, in a node involving the maintenance of network traffic data for government and enterprise users, there are three representation information processing networks. Each sub-processing network of the representation information processing network has screened out its own selected characteristic paths and their network weights and biases through step S21. Assume that the sub-processing network of the first representation information processing network has selected the characteristic path that focuses on the network bandwidth peak, and its network weight is w 11 , with a bias of b 11 ; The second sub-processing network representing the information processing network selects the characteristic path that focuses on abnormal fluctuations in network traffic, and its network weight is w 21 , with a bias of b 21 ; The third sub-processing network representing the information processing network selects the characteristic path that focuses on the traffic of a specific port, and its network weight is w 31 , with a bias of b 31 .

[0058] The computer system can generate the sub-fusion network in a variety of ways. One possible method is the weighted average method. For the bias, the weighted average is also calculated. The weights and biases of the fused network obtained through these calculations constitute part of the sub-fusion network.

[0059] Another method is the concatenation method, in which the computer system concatenates the network weights and biases of the selected feature paths in a certain order.

[0060] Through the above method, the computer system generates a sub-fusion network for each sub-processing network. These sub-fusion networks integrate the key information of multiple characterization information processing networks at the same node, and can extract and convert data features more comprehensively and accurately.

[0061] As an implementation mode, in step S21, among the network weights and biases of each feature path included in any sub-processing network, determining the network weight and bias of the selected feature path includes:

[0062] Step S211: performing a normalization operation on the network weights and biases of each feature path included in any sub-processing network to obtain the normalized network weights and biases of each feature path;

[0063] Step S212: determining the significance measure of each feature path according to the standardized network weight and bias of each feature path;

[0064] Step S213: Determine, among the various feature paths, a selected feature path whose significance metric is greater than or equal to a preset value, and obtain a network weight and bias of the selected feature path.

[0065] In step S211, the computer system performs a standardization operation on the network weights and biases of each feature path included in any sub-processing network, thereby obtaining the standardized network weights and biases of each feature path. The core purpose of the standardization operation is to eliminate the differences in the numerical range and dimension of different feature paths, so that they can be compared and analyzed fairly and effectively at the same scale. Take a sub-processing network responsible for processing server performance data in a remote operation and maintenance terminal for government and enterprise users as an example. The sub-processing network contains multiple feature paths. Assume that one of the feature paths focuses on the change in the CPU usage of the server, and the value range of its network weight and bias may be between 0 and 1; another feature path focuses on the allocation of server memory, and its corresponding network weight and bias may range from 0 to 1000; and there is another feature path for the server disk I / O operation frequency, and its value range may be between 1 and 100. The data ranges and dimensions of these different feature paths are different. If they are not standardized, in subsequent comparisons and analyses, feature paths with a large value range can dominate the calculation, thereby masking the information of other feature paths with a smaller value range but may be equally important.

[0066] In step S212, based on the standardized network weights and biases of each feature path, the computer system determines the significance measure of each feature path, that is, calculates an importance score for each feature path to measure the role of the feature path in data representation and intelligent agent decision-making process.

[0067] There are many ways to determine the significance measure. Here we take the information gain-based method as an example. Information gain is an indicator that measures the amount of information that a feature can add to the classification task. In the data processing scenario of remote operation and maintenance terminals for government and enterprise users, suppose our task is to determine whether the server is in normal operation based on the various performance indicators of the server (that is, the data represented by the feature path) (this is a binary classification problem, normal or abnormal). Through such calculations, the computer system calculates an information gain value for each feature path, which is the significance measure of the feature path. The larger the information gain value, the more information the feature path provides in distinguishing between normal and abnormal operating states of the server, which means that it is more important in data representation and intelligent agent decision-making.

[0068] In addition to information gain, there are other methods to calculate significance measures, such as methods based on correlation analysis. The computer system can calculate the correlation coefficient between each feature path and the target variable (such as the server running status). The larger the absolute value of the correlation coefficient, the closer the relationship between the feature path and the target variable, and the higher its significance measure. Commonly used correlation coefficient calculation methods such as the Pearson correlation coefficient, the formula is: ,in and are the characteristic path data and the target variable data, and are their means respectively. Through these methods, the computer system can accurately determine the significance measure for each characteristic path, providing a quantitative basis for the subsequent screening of the selected characteristic paths.

[0069] In step S213, the computer system determines selected feature paths whose significance metrics are greater than or equal to a preset value in each feature path, and obtains network weights and biases of these selected feature paths.

[0070] The preset value is a threshold value determined by the computer system based on a large amount of experimental data, experience, and actual application requirements. This threshold value is used to screen out those feature paths that are sufficiently important for data representation and intelligent agent decision-making. For example, after many experiments and analyses of remote operation and maintenance terminal data of government and enterprise users, the computer system determined the preset value to be 0.2. For the feature path representing the server CPU usage rate with an information gain of 0.224 calculated previously, since its significance measure is greater than the preset value of 0.2, this feature path will be determined as the selected feature path. The computer system extracts and saves the network weights and biases of these selected feature paths. The network weights and biases of these selected feature paths will become the key components of the subsequent generation of the sub-fusion network. For example, for the above-selected feature path representing the server CPU usage rate, its corresponding network weights and biases will be separated from the parameters of the entire sub-processing network by the computer system for use in generating the sub-fusion network.

[0071] Through steps S211 - S213, the computer system selects selected feature paths that play a key role in data characterization from the numerous feature paths of any sub-processing network, and obtains their network weights and biases. These operations not only provide core data for the subsequent generation of more effective sub-fusion networks, but also improve the agent's ability to extract data features and capture important information through technical means such as standardization, significance measurement calculation and threshold screening, thereby helping to improve the performance and accuracy of the entire data maintenance and discrimination agent, so that it can better support the remote operation and maintenance terminal data maintenance work of government and enterprise users, and ensure the stable operation of the system and the reliability of data.

[0072] As an implementation mode, step S30, generating a first data processing agent according to each first fusion network, includes:

[0073] Step S31: for any first fusion network, generating a first reconstruction network according to any first fusion network and an initial adjustment network, wherein the initial adjustment network is used to adjust the output of the first reconstruction network;

[0074] Step S32: Generate a first data processing agent based on each first reconstruction network.

[0075] When the computer system executes steps S31-S32 in the implementation method of step S30, it is a coherent process from constructing a first reconstruction network to generating a first data processing intelligent entity, which aims to integrate and optimize the information of multiple first fusion networks to form a more powerful data processing intelligent entity that can more effectively process the target terminal maintenance detection data.

[0076] In step S31, for any first fusion network, the computer system generates a first reconstruction network based on the first fusion network and the initial adjustment network, wherein the initial adjustment network is used to adjust the output of the first reconstruction network. This process is to enable the first reconstruction network to flexibly adjust its output according to different data characteristics and task requirements, so as to better adapt to subsequent data processing tasks.

[0077] Taking the server performance data of the remote operation and maintenance terminal of government and enterprise users as an example, it is assumed that the first fusion network has integrated the data features of the server's CPU usage, memory usage, disk I / O, etc., and output a comprehensive feature representation. This feature representation contains multi-dimensional information about the current performance status of the server, but it may need further adjustment to more accurately reflect the essential characteristics of the data and meet the specific requirements of the discrimination task.

[0078] The initial adjustment network can be regarded as an intelligent adjustment module that can dynamically adjust the output of the first fusion network according to the characteristics of the data and the goals of the agent. For example, the initial adjustment network can be a structure composed of a multi-layer neural network, which receives the output of the first fusion network as input. Assume that the feature vector output by the first fusion network is F=[f1, f2,…, fn], where fi represents the feature values ​​of different dimensions. The initial adjustment network adjusts these feature values ​​through a series of calculations.

[0079] In the specific calculation process, the initial adjustment network can first perform a weighted sum operation on the feature vector F. The weights are learned through training, and they can adjust the contribution of each feature dimension in subsequent calculations based on the importance and relevance of the data. For example, if in a specific operation and maintenance scenario, the CPU usage has a greater impact on judging server performance, then the corresponding weight will be relatively large, making the CPU usage feature occupy a more important position in the weighted sum calculation.

[0080] Next, the initial conditioning network can apply a nonlinear transformation function, such as the ReLU (Rectified Linear Unit) function: y=max(0, x). The weighted sum S is transformed to obtain T=ReLU(S). The purpose of this nonlinear transformation is to introduce nonlinear features so that the network can learn more complex data patterns. For example, in server performance data, there may be nonlinear relationships between different performance indicators. The ReLU function can effectively capture these relationships and transform the weighted sum S into a more representative feature value T.

[0081] Then, the initial conditioning network can be further normalized to ensure that the adjusted feature values ​​are within a suitable range. For example, the Softmax normalization function is used: , where z is the input vector, j represents the dimension of the vector, and K is the total number of dimensions of the vector. Assume that a new vector Z is obtained after the previous calculation. After normalizing it through the Softmax function, the resulting vector Each element in is between 0 and 1, and the sum of all elements is 1. This makes the values ​​between different feature dimensions comparable and can highlight the relative importance of important features.

[0082] Through these calculations and operations, the initial adjustment network generates an adjusted output vector R. This vector R will be used to adjust the output of the first reconstruction network. The computer system combines the output of the first fusion network with the adjusted vector R in some form, such as element-by-element multiplication or concatenation. Assuming that the output of the first fusion network is O, the new output is obtained by element-by-element multiplication. ,in This new output O' is the output of the first reconstruction network. After fine-tuning by the initial adjustment network, it can more accurately reflect the characteristics of the data and adapt to subsequent data processing tasks.

[0083] In step S32, the computer system generates a first data processing agent based on each first reconstruction network, which is a process of integrating multiple adjusted and optimized first reconstruction networks to form a unified data processing agent.

[0084] Continuing with the above-mentioned server performance data processing as an example, assume that there are multiple first reconstruction networks, each of which focuses on different aspects of server performance data, such as some first reconstruction networks mainly process CPU-related data, some process memory-related data, and some process disk I / O-related data.

[0085] The computer system first needs to determine the combination method of these first reconstruction networks. One feasible method is the cascade method. That is, these first reconstruction networks are connected in a certain order, and the output of the previous first reconstruction network is used as the input of the next first reconstruction network. For example, the output O1 of the first reconstruction network R1 that processes CPU-related data is first used as the input of the first reconstruction network R2 that processes memory-related data. R2 further integrates and processes O1 and the memory data processed by itself, and then outputs O2. Then O2 is used as the input of the first reconstruction network R3 that processes disk I / O-related data, and so on. Through this cascade method, each first reconstruction network can further mine and integrate the characteristics of the data on the basis of the previous network, so that the first data processing agent finally generated can comprehensively consider various factors and process the server performance data more comprehensively and accurately.

[0086] During the cascading process, the computer system also ensures the compatibility and synergy between the various first reconstruction networks. This includes matching of network structures, consistency of data dimensions, etc. For example, if the output dimension of R1 is m, and the input dimension of R2 is required to be n, the computer system may need to adjust the output dimension of R1 to n through some operations (such as linear transformation) to ensure that the data can be smoothly transmitted and processed between the various first reconstruction networks.

[0087] In addition to the cascade method, the computer system can also combine the first reconstruction networks in a parallel manner. In the parallel method, each first reconstruction network processes the input data at the same time, and then merges their outputs. For example, each first reconstruction network processes different aspects of server performance data, and then splices or weighted sums their output vectors to obtain a comprehensive output vector. Assume that R1, R2, and R3 process CPU, memory, and disk I / O data respectively, and their outputs are O1, O2, and O3 respectively. The computer system can splice them into a new vector O=[O1; O2; O3] (where ";" represents a vector splicing operation), or through a weighted summation method: O=w1 O1+w2O2+w3 O3, where w1, w2, and w3 are weights assigned according to the importance of each first reconstruction network, and w1+w2+w3=1.

[0088] Through these combinations, the computer system integrates multiple first reconstruction networks into a first data processing intelligent agent, which can comprehensively utilize the advantages of each first reconstruction network to perform comprehensive and in-depth processing on the target terminal maintenance detection data.

[0089] Through steps S31-S32, the computer system constructs a first reconstruction network from a single first fusion network and an initial adjustment network, and then integrates multiple first reconstruction networks into a first data processing agent, completing a process from local optimization to overall integration. This process enables the first data processing agent to better process the target terminal maintenance detection data through fine adjustment and reasonable integration, providing more powerful and accurate support for subsequent data maintenance judgment operations, helping to improve the efficiency and quality of data maintenance of remote operation and maintenance terminals for government and enterprise users, and ensuring the stable operation of the system.

[0090] As an implementation mode, step S40, debugging the first data processing agent to obtain a data maintenance determination agent, includes:

[0091] Step S41: obtaining terminal maintenance detection training data and priori marks of each target data processing agent, the terminal maintenance detection training data corresponding to any target data processing agent is used to train any target data processing agent, and the priori mark of any target data processing agent is used to annotate the maintenance result obtained by performing data maintenance discrimination operation on the terminal maintenance detection training data;

[0092] Step S42: for the terminal maintenance detection training data of any target data processing agent, based on the first data processing agent, a corresponding data maintenance discrimination operation is performed on the terminal maintenance detection training data of any target data processing agent to obtain an inference result corresponding to any target data processing agent;

[0093] Step S43: Based on the inference results and prior labels corresponding to each target data processing agent, the first data processing agent is trained to obtain a data maintenance discrimination agent.

[0094] In step S41, the computer system obtains the terminal maintenance detection training data and prior labels of each target data processing agent. Terminal maintenance detection training data is the cornerstone of agent training. It covers a large amount of data collected by remote operation and maintenance terminals of government and enterprise users in various actual operation and maintenance scenarios. These data are rich in diversity and complexity. For example, for a large government and enterprise unit, its remote operation and maintenance terminal may monitor the operating status of many servers in real time. The training data may include detailed information on the CPU usage, memory occupancy, disk I / O read and write rates, real-time consumption of network bandwidth, and application response time of the server during different business peaks and troughs. These data not only record the parameters of the device under normal operating conditions, but also contain data samples under various abnormal conditions, such as abnormal surges in CPU usage when server hardware fails, and continuous increases in memory occupancy due to memory leaks.

[0095] Priori marks are known results obtained after professional analysis and judgment of these training data, and are used to accurately annotate the maintenance results obtained by performing data maintenance and judgment operations on the terminal maintenance detection training data. For example, for a set of server operation data, after in-depth analysis and judgment by professional operation and maintenance personnel, if it is found that the CPU usage rate exceeds 85% for a long time and the memory usage is close to the physical memory limit, it can cause a serious decline in system performance or even a risk of crash. At this time, the priori mark will clearly mark that the data requires urgent maintenance, and it is recommended to take specific measures such as optimizing application algorithms and increasing server resources. For another example, if the network bandwidth fluctuates abnormally within a specific time period, the priori mark will point out this data anomaly and may prompt to check whether the network equipment is faulty or under network attack.

[0096] In actual operation, the computer system collects these training data from various monitoring points of the remote operation and maintenance terminal through a special data acquisition module and stores them in a carefully designed database. At the same time, a professional operation and maintenance team combines rich experience and expertise to conduct a detailed analysis of each set of training data and generate accurate prior labels. These prior labels correspond one-to-one with the training data to form a complete data set, which provides clear guidance and standards for the subsequent training of the intelligent agent.

[0097] Entering step S42, the computer system performs terminal maintenance detection training data for any target data processing agent, and performs corresponding data maintenance discrimination operations on it based on the first data processing agent, thereby obtaining the corresponding reasoning result of the target data processing agent. This process is the core link for the agent to analyze and judge the input data, and demonstrates the processing ability and decision logic of the first data processing agent for different types of data.

[0098] Assume that there is a target data processing agent that is specifically used to analyze the stability of the server's network connection. The computer system inputs the terminal maintenance detection training data corresponding to the agent into the first data processing agent. These training data may contain real-time status information of the network connection, such as network delay, packet loss rate, number of connection interruptions, etc. After receiving these data, the first data processing agent will first extract and convert the features. For example, through a specific algorithm and network structure, the network delay data is converted into a feature vector that reflects the delay fluctuation trend, and the packet loss rate data is mapped into a numerical feature that can reflect the network reliability.

[0099] Next, the first data processing agent uses its internal decision-making mechanism to conduct a comprehensive analysis and judgment of these features. Assume that the agent has been trained and has learned that when the fluctuation of network delay exceeds a certain threshold and the packet loss rate exceeds a certain set value for multiple consecutive sampling points, the network connection may be at risk of instability. Then, when the current training data is input, the agent will infer the data based on these learned rules and patterns. If the current data shows that the standard deviation of the fluctuation of network delay has reached 0.5 (assuming the threshold is 0.3), and the packet loss rate exceeds 5% on average in the last 10 sampling points (assuming the set value is 3%), the agent will infer that there is a stability problem with the network connection, which is the inference result corresponding to the target data processing agent.

[0100] The computer system uses the predefined network structure and parameter settings within the first data processing agent to perform the above operations. These network structures may include multi-layer neural networks, such as convolutional neural networks (CNN) for extracting spatial features of data, and recurrent neural networks (RNN) or their variants (such as LSTM, GRU) for processing data with time series characteristics. Through the layer-by-layer calculation and conversion of these network structures, the input training data is gradually converted into feature representations that the agent can understand and make decisions. At the same time, the computer system uses programming languages ​​and deep learning frameworks (such as Python combined with TensorFlow or PyTorch) to implement data input, agent calling, and the generation of reasoning results.

[0101] After completing step S42 and obtaining the inference results corresponding to each target data processing agent, the computer system enters step S43. In this step, the first data processing agent is trained based on the inference results and prior labels corresponding to each target data processing agent, thereby obtaining a data maintenance discrimination agent. This training process is a process of continuously optimizing the agent parameters so that the agent's inference results are as close as possible to the prior labels, thereby minimizing the difference between the two to improve the accuracy and reliability of the agent.

[0102] Specifically, the computer system trains the first data processing agent based on the first training cost and the second training cost corresponding to each target data processing agent. The first training cost corresponding to any target data processing agent is determined based on the nonlinear transformation results of each feature path corresponding to each first reconstruction network. For example, in a target data processing agent that processes server performance data, the first reconstruction network processes features such as server CPU usage and memory occupancy. During the processing, these features are converted into more representative feature representations through nonlinear transformations (such as ReLU function: y=max(0, x)). Assume that after the nonlinear transformation, the feature value corresponding to a feature path is [0.2, 0.5, 0.8]. The first training cost can be determined by measuring the difference between these nonlinear transformation results and the expected feature representation. A feasible calculation method is to use the mean square error (MSE) loss function. By calculating the MSE, the computer system can obtain the first training cost of the target data processing agent, which reflects the degree of error of the agent in the feature conversion process.

[0103] The second training cost corresponding to any target data processing agent is determined according to the reasoning result and prior mark corresponding to the target data processing agent. For example, for the target data processing agent that analyzes the stability of the server network connection, its reasoning result is that the network connection has a stability problem, and the prior mark indicates that the network connection is stable in the current state. At this time, the computer system can use the cross entropy loss function to calculate the second training cost. This loss function measures the degree of difference between the agent's reasoning result and the prior mark.

[0104] The computer system comprehensively considers the first training cost and the second training cost, and uses an optimization algorithm (such as a stochastic gradient descent algorithm) to adjust the parameters of the first data processing agent. Through continuous iterative training, that is, inputting the training data of different target data processing agents into the agent multiple times, calculating the cost and adjusting the parameters, the difference between the reasoning result of the first data processing agent and the prior label is gradually reduced.

[0105] During the training process, the computer system can also use some techniques to improve the training effect and the generalization ability of the agent. For example, using data enhancement technology, randomly transforming the training data (such as translating and scaling the time series in the server performance data) to increase the diversity of the data and prevent the agent from overfitting. At the same time, using regularization methods (such as L1 and L2 regularization to constrain the complexity of the agent, to avoid the agent learning the noise in the data and overfitting.

[0106] As the training progresses, the performance of the first data processing agent gradually improves, and the consistency of its reasoning results with the prior labels becomes higher and higher. When the performance of the agent reaches a certain level of satisfaction, the computer system stops training. At this time, the data maintenance judgment agent obtained has been fully optimized and debugged, and can accurately perform various data maintenance judgment operations on the target terminal maintenance detection data. For example, for newly input server performance data or network connection data, the data maintenance judgment agent can quickly and accurately determine whether the data needs maintenance, as well as determine the direction and focus of maintenance, providing strong support for the remote operation and maintenance work of government and enterprise users, helping them to promptly discover potential problems and take effective measures to ensure the stable operation of the system and the reliability of data.

[0107] As an implementation mode, each first reconstruction network is cascaded, and in step S42, based on the first data processing agent, a corresponding data maintenance discrimination operation is performed on the terminal maintenance detection training data of any target data processing agent to obtain an inference result corresponding to any target data processing agent, including:

[0108] Step S421: for the first first reconstruction network, extracting representation information from the terminal maintenance detection training data of any target data processing agent based on the first first reconstruction network to obtain a data representation vector output by the first first reconstruction network;

[0109] Step S422: for the non-first first reconstruction network, extracting representation information from a data representation vector output by a previous first reconstruction network of the non-first first reconstruction network based on the non-first first reconstruction network, to obtain a data representation vector output by the non-first first reconstruction network;

[0110] Step S423: Determine the reasoning result corresponding to any target data processing agent based on the data representation vector output by the last first reconstruction network.

[0111] In step S421, for the first first reconstruction network, the computer system extracts representation information from the terminal maintenance detection training data of any target data processing agent based on the network, thereby obtaining the data representation vector output by the first first reconstruction network. The first first reconstruction network plays a key role in the entire data processing process. It is responsible for the preliminary feature extraction and conversion of the original terminal maintenance detection training data, converting the data into a more representative and valuable feature representation, and providing a basis for subsequent processing.

[0112] Taking the server performance data of remote operation and maintenance terminals for government and enterprise users as an example, assuming that the target data processing agent is concerned with the overall health status assessment of the server, its terminal maintenance detection training data contains information such as the server's CPU usage, memory usage, disk I / O rate, and network bandwidth over a period of time. After receiving these raw data, the first reconstruction network begins to extract representation information.

[0113] The association perception unit (which can be understood as the attention layer) in the network first performs association perception operations on the data, that is, attention processing. For example, the association perception unit analyzes the potential correlation between various indicators in the data. In the server performance data, there may be a certain correlation between CPU usage and memory usage. When the CPU usage increases, the memory usage may also increase accordingly. The association perception unit identifies these associations through specific algorithms and weight settings, and assigns different attention weights to different parts of the data according to the importance and relevance of the data. Assuming that it is found through calculation that the CPU usage has a higher importance in the current evaluation, the association perception unit will assign a larger attention weight to it, so that the agent pays more attention to this indicator in subsequent processing. This process can be expressed by a mathematical formula. Assuming that the input data is a vector X=[x1, x2, …, xn] containing multiple indicators, the association perception unit calculates the attention weight matrix A and performs matrix multiplication with the input data to obtain the weighted input data X'=A·X, where the element a in A is ij represents the attention weight of the i-th indicator on the j-th dimension.

[0114] Next, the intermediate processing unit (that is, the hidden layer) performs feature fusion processing on the associated perception results. The intermediate processing unit usually contains multiple neuron layers, and performs deep processing on the associated perception results through nonlinear transformation and matrix operations. It combines and transforms the features of different indicators to mine more complex patterns and features in the data. For example, the intermediate processing unit can fuse features such as CPU usage, memory usage, and disk I / O rate, and generate a new and more representative feature vector by calculating the nonlinear relationship between them. Assume that the intermediate processing unit adopts a multi-layer perceptron (MLP) structure, and the input is the associated perception result X'. After a series of weight matrices W1, W2 and nonlinear activation functions (such as ReLU function: y=max(0, x)), the feature fusion processing result H is obtained. The specific calculation process is: Z1=W1·X', H=ReLU(Z1), Z2=W2·H, H'=ReLU(Z2), and the final H' is the feature fusion processing result.

[0115] Based on the initial adjustment network, the feature fusion processing result is adjusted according to the associated perception result to obtain the adjustment result. The initial adjustment network includes a feature conversion unit (a mapping layer) and a nonlinear transformation unit. The feature conversion unit performs feature conversion on the associated perception result, for example, mapping certain features in the associated perception result to different feature spaces to better adapt to subsequent processing. Assuming that the associated perception result is A, the feature conversion unit converts A to A'=M·A through a mapping matrix M. The nonlinear transformation unit performs nonlinear transformation on the feature conversion results of each feature path to obtain the nonlinear transformation results of each feature path, which are used to characterize the activity of the feature path. For example, the nonlinear transformation unit uses the Sigmoid function to transform each element in the feature conversion result A' to obtain the nonlinear transformation result S. Through these nonlinear transformations, important features in the data can be highlighted and unimportant features can be suppressed. Finally, the adjustment result is determined based on the nonlinear transformation results of each feature path and the path processing results of multiple feature paths.

[0116] The output result is determined based on the adjustment result based on the result mapping unit. The result mapping unit (which can be understood as the output layer) maps the adjustment result to the final output space to obtain the output result of the first reconstruction network. For example, the result mapping unit linearly transforms the adjustment result R through a weight matrix O and a bias vector b, and then passes through an activation function (such as the Softmax function) to obtain the output result Y=Softmax(O·R+b). This output result Y is the data representation vector output by the first reconstruction network. It integrates various feature information of the original data and has been processed through multiple links such as association perception, feature fusion, and adjustment. It can more accurately represent the input terminal maintenance detection training data.

[0117] After completing step S421, the computer system proceeds to step S422. For the non-first first reconstruction network, the computer system extracts the representation information of the data representation vector output by the previous first reconstruction network based on the non-first first reconstruction network to obtain the data representation vector output by the non-first first reconstruction network. This step enables the data to be transmitted and processed in sequence in multiple first reconstruction networks, and continuously digs deeper into the characteristics of the data.

[0118] Continuing with the above-mentioned server performance data processing as an example, assume that there is a second first reconstruction network, which is responsible for further analyzing the potential trends and anomalies in the server performance data. The second first reconstruction network receives the data representation vector output by the first first reconstruction network as input. Similarly, its association perception unit will perform attention processing on the input data representation vector and analyze the new association relationship between different features. For example, it can focus on the changing trend of the server performance indicators reflected in the data representation vector over time, and assign higher attention weights to those features related to the trend.

[0119] The intermediate processing unit performs feature fusion processing on the associated perception results to further mine the complex features in the data. It can combine the features output by the first reconstruction network with the new information it receives to perform more in-depth feature combination and conversion. For example, the features of the current performance status of the server output by the first reconstruction network are fused with the features obtained by the performance trend analysis of the second reconstruction network itself to generate a more comprehensive feature vector that better reflects the dynamic changes in server performance.

[0120] Based on the initial adjustment network, the feature fusion processing result is adjusted to obtain the adjustment result. Similar to the first reconstruction network, the initial adjustment network adjusts the feature fusion processing result through the feature conversion unit and the nonlinear transformation unit to highlight important features and suppress unimportant features. For example, the feature is mapped to a more suitable space through the feature conversion unit, and then processed by the nonlinear transformation unit, so that the adjustment result can more accurately reflect the essential characteristics of the data. Finally, the result mapping unit determines the output result based on the adjustment result, and this output result is the data representation vector output by the second first reconstruction network.

[0121] This process is repeated in multiple non-first first reconstruction networks. Each non-first first reconstruction network conducts a deeper analysis and feature extraction of the data based on the previous network, continuously enriching and optimizing the representation of the data. In terms of technical implementation, the construction and operation of each non-first first reconstruction network is similar to the first first reconstruction network. The network structure is defined and the parameters are set through a deep learning framework to ensure that the data can be smoothly transmitted and processed between the networks.

[0122] Finally, in step S423, the computer system determines the reasoning result corresponding to any target data processing agent based on the data representation vector output by the last first reconstruction network. The data representation vector output by the last first reconstruction network is obtained after being processed layer by layer by multiple first reconstruction networks, and it contains rich feature information about the target terminal maintenance detection training data.

[0123] The computer system determines the inference result based on these characteristic information and predefined decision rules or intelligent agent structures. For example, in the processing of server performance data, if the data representation vector output by the last first reconstruction network shows that the CPU usage of the server has been high for a long time and has an upward trend, the memory usage has continued to increase, and the disk I / O rate has abnormal fluctuations, according to the pre-set rules, the computer system can infer that there is a problem with the current performance of the server and maintenance operations are required, such as optimizing applications, increasing server resources, etc.

[0124] Specifically, the computer system can use a classifier or regression agent to reason based on the data representation vector. For example, a simple linear classifier is used to take the data representation vector V output by the last first reconstruction network as input, and a linear transformation is performed through the weight matrix W and the bias vector b to obtain Z=W·V+b. Then, the reasoning result is determined by a threshold function (for example, for a binary classification problem, when Z is greater than a certain threshold, it is determined to require maintenance; when Z is less than the threshold, it is determined not to require maintenance). For more complex situations, a neural network classifier or regression agent can be used to obtain more accurate reasoning results through multiple layers of nonlinear transformations and decision-making mechanisms.

[0125] Through steps S421 - S423, the computer system can conduct a comprehensive and in-depth analysis of the terminal maintenance detection training data of the target data processing agent based on the first data processing agent, gradually extract the characteristics of the data, and finally obtain accurate reasoning results. This process not only demonstrates the agent's powerful data processing capabilities, but also provides a basis for subsequent training of the first data processing agent based on reasoning results and prior labels, which helps to continuously optimize the agent's performance and improve the accuracy and reliability of the data maintenance discrimination agent, thereby better supporting the remote operation and maintenance work of government and enterprise users.

[0126] As an implementation mode, the first first reconstruction network includes multiple first reconstruction networks, and the first fusion network in the first reconstruction network includes an associated perception unit, an intermediate processing unit, and a result mapping unit; step S421, based on the first first reconstruction network, extracts representation information from the terminal maintenance detection training data of any target data processing agent to obtain a data representation vector output by the first first reconstruction network, including:

[0127] Step S4211: for any first reconstruction network included in the first first reconstruction network, based on the association perception unit, perform an association perception operation on the terminal maintenance detection training data of any target data processing agent to obtain an association perception result; based on the intermediate processing unit, perform feature fusion processing on the association perception result to obtain a feature fusion processing result; based on the initial adjustment network, adjust the feature fusion processing result according to the association perception result to obtain an adjustment result; based on the result mapping unit, determine the output result according to the adjustment result;

[0128] Step S4212: Determine a data representation vector output by the first reconstruction network according to the output results of each first reconstruction network included in the first reconstruction network.

[0129] In step S4211, for any first reconstruction network included in the first first reconstruction network, the computer system performs an association perception operation on the terminal maintenance detection training data of any target data processing agent based on the association perception unit to obtain an association perception result; performs feature fusion processing on the association perception result based on the intermediate processing unit to obtain a feature fusion processing result; adjusts the feature fusion processing result based on the initial adjustment network according to the association perception result to obtain an adjustment result; and determines the output result based on the adjustment result based on the result mapping unit.

[0130] First, the association perception unit performs association perception operations on the terminal maintenance detection training data of any target data processing agent, such as performing attention processing. Then, the intermediate processing unit performs feature fusion processing on the association perception results. The intermediate processing unit is usually composed of multiple neuron layers and has powerful nonlinear transformation capabilities. Its task is to organically combine the different features in the association perception results to mine more complex and representative features. Continuing with the above-mentioned server performance data as an example, the intermediate processing unit will comprehensively process the features such as CPU usage, memory usage, disk I / O rate, network bandwidth, etc. after association perception weighting.

[0131] Assuming that the intermediate processing unit adopts a multi-layer perceptron (MLP) structure, it first performs a linear transformation on the associated perception result X'. Then, using a nonlinear activation function, such as the ReLU function: y=max(0, x), Z1 is nonlinearly transformed to obtain H1=ReLU(Z1). The role of this nonlinear transformation is to introduce nonlinear features so that the agent can learn more complex patterns in the data. Then, more layers of similar operations can be performed, such as linear transformation through the weight matrix W2 and the bias vector b2 again: Z2=W2·H1+b2, and then transformed by the ReLU function to obtain H2=ReLU(Z2). After these layers of processing, features of different dimensions are deeply fused to form a new and more representative feature vector, that is, the feature fusion processing result H2. This feature fusion processing result is no longer a simple combination of raw data, but a comprehensive representation of the complex relationship between data, which can more comprehensively reflect the actual situation of server performance.

[0132] Based on the initial adjustment network, the feature fusion processing result is adjusted according to the associated perception result to obtain the adjustment result. The initial adjustment network includes a feature conversion unit and a nonlinear transformation unit. The feature conversion unit performs feature conversion on the associated perception result, and its purpose is to map the features in the associated perception result to a more appropriate feature space so as to better combine it with the feature fusion processing result. Assuming that the associated perception result is X', the feature conversion unit converts it through a specific mapping matrix M to obtain X''=M·X'.

[0133] The nonlinear transformation unit performs a nonlinear transformation on the feature conversion results of each feature path to obtain the nonlinear transformation results of each feature path, which are used to characterize the activity of the feature path. For example, using the Sigmoid function, a nonlinear transformation is performed on each element in X''. After the Sigmoid function transformation, the value of each element is compressed to between 0 and 1. The closer the value is to 1, the higher the activity of the feature path, and the closer the value is to 0, the lower the activity. In this way, the nonlinear transformation unit can highlight important feature paths and suppress unimportant feature paths.

[0134] Finally, the adjustment result is determined based on the nonlinear transformation results of each characteristic path and the path processing results of multiple characteristic paths (ie, the feature fusion processing result H2). For example, the result after nonlinear transformation and the feature fusion processing result are weighted and summed.

[0135] The output result is determined based on the result mapping unit according to the adjustment result. The result mapping unit maps the adjustment result to the final output space to obtain the output result of the first reconstruction network in the first reconstruction network. Assume that the result mapping unit performs a linear transformation through a weight matrix O and a bias vector b3: Y=O·R+b3. Then, according to the specific task requirements, it can be processed through an activation function (such as a Softmax function) to obtain the final output result Y'. This output result Y' integrates the processing information of multiple links such as association perception, feature fusion, and adjustment, and is the final representation of the first reconstruction network in the first reconstruction network after processing the input data.

[0136] After completing the processing of each first reconstruction network in step S4211, proceed to step S4212. The computer system determines the data representation vector output by the first first reconstruction network based on the output results of each first reconstruction network included in the first first reconstruction network. Since the first first reconstruction network may include multiple first reconstruction networks, each first reconstruction network processes the input data at different angles and depths to obtain its own output results.

[0137] Continuing with the example of server performance data processing, assume that the first first reconstruction network includes three first reconstruction networks, which process the server performance data from different aspects. The first first reconstruction network focuses on the processing of CPU and memory related features, the second focuses on the processing of disk I / O and network related features, and the third focuses on the overall trend of server performance and anomaly detection. After the three first reconstruction networks are processed in step S4211, the results Y1, Y2, and Y3 are output respectively.

[0138] The computer system can use a variety of methods to determine the data representation vector output by the first reconstruction network. One feasible method is the splicing operation. The computer system splices the three output results in a certain order to form a new, higher-dimensional data vector. For example, Y1, Y2, and Y3 are sequentially spliced ​​into a new data vector V=[Y1;Y2;Y3] (where ";" represents a vector splicing operation). This spliced ​​vector V contains the information obtained by the three first reconstruction networks from processing data from different angles, and can more comprehensively represent the input server performance data.

[0139] Another method is weighted summation, in which the computer system assigns different weights to the output results of each first reconstruction network according to its importance in data processing.

[0140] Through steps S4211 - S4212, the computer system performs comprehensive and in-depth processing on the target terminal maintenance detection training data in the first first reconstruction network. From the association perception unit capturing data associations, to the intermediate processing unit fusing features, to the initial adjustment network adjusting results, and finally determining the overall data representation vector based on the output of each first reconstruction network, each step is closely coordinated to gradually mine the key information in the data. This not only improves the data representation capability, but also provides a high-quality feature foundation for subsequent intelligent agent processing and data maintenance discrimination operations, which helps to improve the accuracy and reliability of the entire data maintenance discrimination intelligent agent, thereby better providing support for the remote operation and maintenance work of government and enterprise users.

[0141] As an implementation mode, the initial adjustment network includes a feature conversion unit and a nonlinear transformation unit, and the feature fusion processing result includes path processing results of multiple feature paths; step S4211, based on the initial adjustment network, the feature fusion processing result is adjusted according to the associated perception result to obtain the adjustment result, including:

[0142] Step S42111: performing feature conversion on the associated perception result based on the feature conversion unit to obtain feature conversion results of each feature path;

[0143] Step S42112: performing nonlinear transformation on the feature conversion results of each feature path based on the nonlinear transformation unit to obtain the nonlinear transformation results of each feature path, and the nonlinear transformation results of the feature path are used to characterize the activity of the feature path;

[0144] Step S42113: Determine the adjustment result based on the nonlinear transformation result of each characteristic path and the path processing results of multiple characteristic paths.

[0145] In step S42111, the feature conversion unit performs feature conversion on the associated perception result to obtain the feature conversion results of each feature path. As an important part of the initial adjustment network, the feature conversion unit undertakes the key task of mapping the features in the associated perception result to a more appropriate feature space. This process is designed to enable the features to play a better role in subsequent processing and enhance the agent's ability to understand and process data.

[0146] Taking the server performance data of the remote operation and maintenance terminal of government and enterprise users as an example, it is assumed that the associated perception result is a vector A=[a1,a2, a3, a4] containing multi-dimensional information such as server CPU usage, memory usage, disk I / O rate, and network bandwidth. The feature conversion unit realizes feature conversion through a specific mapping matrix M. The mapping matrix M is a matrix that is pre-defined or obtained through training according to the characteristics of the data and the requirements of the intelligent agent, and its dimension matches the associated perception result vector A. For example, if A is a 4-dimensional vector, M may be a 4×4 matrix.

[0147] Through matrix multiplication operations, the computer system multiplies the association perception result A with the mapping matrix M to obtain the feature conversion result A' of each feature path. The specific calculation formula is: A'=M·A. Among them, each element a_i' in A' is the result of linear combination of each feature in the original association perception result, which represents the representation of each feature path in the new feature space. Through this feature conversion, the features that may be difficult to be effectively used by the intelligent agent in the original space are converted to a space that is more conducive to the learning and processing of the intelligent agent. For example, in server performance data, there may be some complex nonlinear relationship between CPU utilization and memory usage. Through feature conversion, this potential relationship may be more clearly displayed, allowing the intelligent agent to better capture the key information in the data.

[0148] After completing the feature conversion, the computer system enters step S42112. In this step, the feature conversion results of each feature path are nonlinearly transformed based on the nonlinear transformation unit to obtain the nonlinear transformation results of each feature path, which are used to characterize the activity of the feature path. The nonlinear transformation unit processes the feature conversion results through a specific nonlinear function, breaking the limitations of the linear agent and enabling the agent to learn more complex patterns and relationships in the data.

[0149] Through this nonlinear transformation, the activity of the feature path can be quantitatively represented. For example, if the value of σ(a1') is close to 1, it means that under the current data state, the feature path related to a1' has a high activity and may play an important role in data representation and agent decision-making; conversely, if the value of σ(a1') is close to 0, it means that the activity of the feature path is low and its impact on the overall data is relatively small.

[0150] After obtaining the nonlinear transformation results of each characteristic path, the computer system enters step S42113. In this step, the adjustment result is determined based on the nonlinear transformation results of each characteristic path and the path processing results of multiple characteristic paths. This step is to organically combine the characteristic path activity information after nonlinear transformation with the previous feature fusion processing results, so as to obtain an adjustment result that comprehensively considers data characteristics and activity.

[0151] Assume that the nonlinear transformation results of each feature path obtained in step S42112 form a vector S=[s1,s2, s3, s4], where si represents the nonlinear transformation result of the i-th feature path, that is, the activity; the path processing results of multiple feature paths (that is, the feature fusion processing results) are vectors H=[h1, h2, h3, h4]. The computer system determines the adjustment result through a comprehensive calculation method.

[0152] One feasible method is weighted summation. The computer system assigns weights to the corresponding elements of the nonlinear transformation results and feature fusion processing results of each feature path, and then performs weighted summation calculation.

[0153] Through this weighted summation method, the computer system can flexibly adjust the adjustment results according to the activity of the feature path and the importance of the feature fusion processing results. For example, if a feature path has a high activity (i.e., a large si value) and is also important in the feature fusion processing results (i.e., a large hi value), then in the adjustment result, the element ri corresponding to the feature will get a larger value, thus occupying a more important position in the overall adjustment result.

[0154] Another method is fusion based on the attention mechanism. The computer system dynamically assigns attention weights according to the activity of each feature path, and then applies these attention weights to the feature fusion processing results. For example, first calculate the attention weight vector A'=[a1', a2', a3', a4'], where ai' is the weight obtained by a certain attention calculation method based on si. Then, the adjustment result R is obtained by the formula R=A'·H (where "·" means element-by-element multiplication). This method based on the attention mechanism can more flexibly adjust the feature fusion processing results according to the activity of the feature path, so that the adjustment results can better reflect the actual situation of the data.

[0155] Through steps S42111-S42113, the computer system processes the associated perception results comprehensively and carefully. From the feature conversion unit mapping the features to the new space, to the nonlinear transformation unit quantifying the activity of the feature path, to determining the adjustment results based on these results, each step is closely coordinated to gradually optimize the data representation. This not only improves the feature expression ability of the data, but also enables the agent to better understand and process the data, and provides high-quality input for the subsequent result mapping unit to determine the output results, thereby helping to improve the performance and accuracy of the entire data maintenance and judgment agent, so that it can more accurately provide maintenance and judgment support for the remote operation and maintenance terminal data of government and enterprise users.

[0156] As an implementation mode, step S43, based on the reasoning results and priori labels corresponding to each target data processing agent, training the first data processing agent to obtain a data maintenance discrimination agent includes:

[0157] Step S431: Based on the first training cost and the second training cost corresponding to each target data processing agent, the first data processing agent is trained to obtain a data maintenance discrimination agent; wherein the first training cost corresponding to any target data processing agent is determined based on the nonlinear transformation results of each feature path corresponding to each first reconstruction network; the second training cost corresponding to any target data processing agent is determined according to the inference result and prior label corresponding to any target data processing agent.

[0158] When the computer system executes step S431 in the implementation of step S43, the core task is to train the first data processing agent based on the first training cost and the second training cost corresponding to each target data processing agent, so as to obtain a data maintenance discrimination agent. This process is a key link in the entire agent training. By comprehensively considering the two training costs, the parameters of the first data processing agent are continuously optimized, so that it is gradually transformed into a data maintenance discrimination agent with better performance.

[0159] The first training cost corresponding to any target data processing agent is determined based on the nonlinear transformation results of each feature path corresponding to each first reconstruction network. For example, when processing the server performance data of the remote operation and maintenance terminal of government and enterprise users, it is assumed that there is a target data processing agent used to determine whether the server has performance problems. One of the first reconstruction networks is responsible for processing relevant features such as server CPU usage and memory occupancy. After nonlinear transformation, these feature paths obtain corresponding nonlinear transformation results. Assume that the value obtained after nonlinear transformation of the CPU usage feature path is [0.2, 0.3, 0.4], and the nonlinear transformation result of the memory occupancy feature path is [0.5, 0.6, 0.7]. The first training cost can be determined by measuring the difference between these nonlinear transformation results and the expected feature representation. Such calculations are performed on all feature paths of each first reconstruction network, and then summarized to obtain the first training cost of the target data processing agent.

[0160] The second training cost corresponding to any target data processing agent is determined based on the inference result and prior mark corresponding to the target data processing agent. Continuing with the above server performance data as an example, assume that the target data processing agent infers that the current performance of the server is normal based on the processing of the first data processing agent, but the prior mark indicates that the server has performance problems (such as high CPU usage and continuous increase in memory usage). For this binary classification problem, the cross entropy loss function can be used to calculate the second training cost.

[0161] The computer system comprehensively considers the first training cost and the second training cost, and uses an optimization algorithm to train the first data processing agent. For example, a stochastic gradient descent algorithm is used to adjust the parameters of the first data processing agent. The computer system traverses the training data of all target data processing agents, calculates the corresponding first training cost and second training cost for each set of data, and then calculates the total loss function J and its gradient according to the above formula.

[0162] By repeating this process, i.e., calculating and updating the parameters of each set of training data of the target data processing agent, the computer system gradually adjusts the parameters of the first data processing agent so that the value of the total loss function J continues to decrease. As the training progresses, the first data processing agent's ability to process data continues to improve, and the difference between its reasoning results and the prior labels gradually decreases. At the same time, the nonlinear transformation results of the characteristic paths of each first reconstruction network are also closer to the expected representation. Finally, when the total loss function J converges to a smaller value, or meets certain stopping conditions, the computer system believes that the first data processing agent has been fully trained, and the data maintenance discrimination agent obtained at this time can more accurately perform data maintenance discrimination operations on the target terminal maintenance detection data.

[0163] As an implementation mode, step S40, debugging the first data processing agent to obtain a data maintenance determination agent, includes:

[0164] Step S40a: training each initial adjustment network in the first data processing agent to obtain each first adjustment network;

[0165] Step S40b: Perform rotation training on each first fusion network and each first adjustment network to obtain a data maintenance discrimination agent.

[0166] In step S40a, the computer system trains each initial adjustment network in the first data processing agent to obtain each first adjustment network. The initial adjustment network plays an important role in the first data processing agent. It can dynamically adjust the output of the first reconstruction network according to the characteristics of the data and the task requirements to optimize the performance of the agent.

[0167] Taking the processing of server performance data of remote operation and maintenance terminals for government and enterprise users as an example, it is assumed that the first data processing agent includes multiple first reconstruction networks, and each first reconstruction network is equipped with an initial adjustment network. The initial adjustment network can be regarded as a small neural network structure, which receives the output of the first reconstruction network as input and adjusts it. For example, one of the first reconstruction networks mainly processes hardware performance-related data such as the server's CPU usage and memory usage, and its output is a vector containing these data features. The initial adjustment network corresponding to the first reconstruction network will further process this vector.

[0168] During the training process, the computer system uses a large amount of target terminal maintenance detection training data and corresponding priori marks. For the first reconstruction network and its initial adjustment network that process hardware performance data, the training data may include actual measured values ​​such as CPU usage and memory usage of the server under different time periods and different business loads, and the priori marks are determined based on professional operation and maintenance knowledge and historical experience, indicating whether these data are within the normal range and whether maintenance operations are required.

[0169] The computer system measures the difference between the output of the initial adjustment network and the expected result by defining a suitable loss function. In order to adjust the parameters of the initial adjustment network to reduce the value of the loss function, the computer system uses an optimization algorithm, such as the stochastic gradient descent (SGD) algorithm. The computer system determines the direction of parameter update by calculating the gradient, so that in each iteration, the parameters of the initial adjustment network are adjusted in the direction of reducing the value of the loss function. During the training process, the training data is continuously input into the initial adjustment network, the loss function and gradient are calculated, and the network parameters are updated until the loss function converges to a smaller value. At this point, the computer system obtains the trained first adjustment networks, which can more effectively adjust the output of the first reconstruction network, making the intelligent agent more accurate and flexible in processing data.

[0170] In step S40b, the computer system performs rotation training on each first fusion network and each first adjustment network to obtain a data maintenance discrimination agent. The purpose of the rotation training is to further optimize the overall performance of the agent, enable each component to work better together, and improve the agent's ability to discriminate the target terminal maintenance detection data.

[0171] Continuing with the above-mentioned server performance data processing as an example, after each first adjustment network is obtained through training in step S40a, the computer system begins to perform rotation training. In each training iteration, the computer system randomly selects a first fusion network and a first adjustment network corresponding thereto for training. For example, in a certain iteration, the first fusion network for processing server hardware performance data and its corresponding first adjustment network are selected.

[0172] During the training process, the computer system also uses the target terminal to maintain the detection training data and prior labels. The training data is input into the first fusion network, which performs feature fusion processing on the data and outputs a fused feature representation. Then, this feature representation is input into the corresponding first adjustment network, which adjusts it to obtain an adjusted output.

[0173] The computer system again uses the loss function to measure the difference between the adjusted output and the prior mark. For example, for server hardware performance data, the prior mark may indicate that the CPU usage of the current server is too high and maintenance is required. If the adjusted output does not accurately reflect this information, the computer system calculates the difference between the two, such as using the cross entropy loss function. According to the value of the loss function, the computer system uses an optimization algorithm to adjust the parameters of the first fusion network and the first adjustment network. In this process, not only the parameters of the first adjustment network should be adjusted, but also the parameters of the first fusion network should be fine-tuned according to the feedback of the first adjustment network. For example, the gradient is calculated by the back propagation algorithm to determine the direction and amplitude of the parameter update. The back propagation algorithm calculates the gradient of each neuron by back propagating the error from the output layer to the input layer, thereby updating the weights and biases of the network. Through continuous alternating training, the computer system gradually adjusts the parameters of each first fusion network and the first adjustment network to make them work more tacitly. During the training process, different first fusion networks and first adjustment networks will continuously optimize their own parameters according to different types of data and task requirements to better process and analyze data. For example, the first fusion network and the first regulation network that process network security data will gradually adapt to changes in network attack patterns during training and improve the ability to identify network security threats; the first fusion network and the first regulation network that process business processing efficiency data will continuously optimize the understanding of business processes and accurately judge whether business processing is efficient.

[0174] As the training continues, the agent's ability to discriminate the target terminal maintenance detection data continues to improve, and the value of the loss function gradually decreases. When the value of the loss function reaches a pre-set threshold or no longer decreases significantly in a certain number of iterations, the computer system believes that the agent has converged and the training is over. At this time, the data maintenance discrimination agent obtained has been fully optimized and adjusted. This data maintenance discrimination agent can accurately perform various data maintenance discrimination operations on the target terminal maintenance detection data, providing reliable support for the remote operation and maintenance work of government and enterprise users. For example, for new server performance data, the data maintenance discrimination agent can accurately determine whether there are performance problems. For network security data, it can promptly detect potential network attack threats, thereby helping government and enterprise users take timely measures to ensure the stable operation of the system and the security of data.

[0175] As an implementation method, the data maintenance identification agent includes a plurality of target reconstruction networks; step S2, based on the data maintenance identification agent, performs at least one data maintenance identification operation on the target terminal maintenance detection data to obtain various data maintenance identification results, including:

[0176] Step S201: for a first target reconstruction network, extracting representation information of target terminal maintenance detection data based on the first target reconstruction network to obtain a data representation vector output by the first target reconstruction network;

[0177] Step S202: for the non-first target reconstruction network, extracting representation information from a data representation vector output by a previous target reconstruction network of the non-first target reconstruction network based on the non-first target reconstruction network to obtain a data representation vector output by the non-first target reconstruction network;

[0178] Step S203: Determine each data maintenance judgment result according to the data representation vector output by the last target reconstruction network.

[0179] In step S201, for the first target reconstruction network, the computer system extracts characterization information from the target terminal maintenance detection data based on the network to obtain a data characterization vector output by the first target reconstruction network. The first target reconstruction network undertakes the important task of preliminary feature extraction and conversion of the original data in the entire data processing flow. It is like an information filter, extracting the most representative features from the complex target terminal maintenance detection data, and providing a basis for subsequent analysis and discrimination.

[0180] Taking the server performance data of remote operation and maintenance terminals for government and enterprise users as an example, it is assumed that the target terminal maintenance detection data contains multi-dimensional information such as the server's CPU usage, memory usage, disk I / O rate, network bandwidth, etc. over a period of time. After the first target reconstruction network receives these raw data, it begins to extract representation information. The network contains multiple functional layers, and first introduces data into the network through the input layer. For example, data such as CPU usage and memory usage are input into the network in the form of vectors. Next, the convolution layer in the network (if a convolutional neural network structure is used) performs convolution operations on the data. Assume that the convolution kernel is a size of For a two-dimensional server performance data (for example, CPU usage and memory usage at different time points are combined into a two-dimensional matrix), the convolution kernel is multiplied by the corresponding elements of the data matrix and summed to obtain a new feature matrix. This process can be expressed as: ,in is the feature matrix element after convolution, is the convolution kernel element, is the input data matrix element. Through this convolution operation, the first target reconstruction network is able to capture local patterns in the data. For example, in server performance data, it may be possible to find the co-variation pattern of CPU usage and memory usage in certain time periods.

[0181] The pooling layer then performs a pooling operation on the convolutional feature matrix. Possible pooling methods include maximum pooling and average pooling. The purpose of the pooling operation is to reduce the data dimension while retaining the most important features, reducing the amount of calculation and preventing overfitting. After operations such as convolution and pooling, the data is passed to the fully connected layer. The fully connected layer comprehensively processes the feature vectors output by the previous layer and maps the feature vectors to a new feature space through the operation of the weight matrix and the bias vector. Assuming that the feature vector output by the previous layer is X, the weight matrix of the fully connected layer is W, and the bias vector is b, the output Y of the fully connected layer can be calculated by the formula Y=W·X+b. The fully connected layer can learn more complex global features in the data. For example, in server performance data, the relationship between multiple factors such as CPU, memory, disk I / O and network bandwidth is comprehensively considered. Finally, after a series of processing, the first target reconstruction network outputs a data representation vector. This vector integrates various key features in the original data, such as the overall trend of server performance, the correlation between different performance indicators, and other information.

[0182] After completing step S201, the computer system enters step S202. For the non-first target reconstruction network, the computer system extracts the representation information of the data representation vector output by the previous target reconstruction network based on the non-first target reconstruction network to obtain the data representation vector output by the non-first target reconstruction network. This step enables the data to be sequentially transmitted and deeply processed in multiple target reconstruction networks, further mining the potential features in the data.

[0183] Taking server performance data processing as an example, assume that there is a second target reconstruction network. The second target reconstruction network receives the data representation vector output by the first target reconstruction network as input. This input vector already contains information after preliminary feature extraction and conversion, but the second target reconstruction network will further analyze it from different angles.

[0184] Similar to the first target reconstruction network, the second target reconstruction network also contains multiple functional layers. It first introduces the data representation vector output by the previous network into the network through the input layer. Then, the convolution layer can be used again to perform convolution operations on the data. Since the input data representation vector dimension and structure are different from the original data, the convolution operation here aims to explore deeper feature relationships in the data representation vector. For example, the second-order or high-order correlations between the features extracted by the first target reconstruction network can be discovered.

[0185] The pooling layer also performs pooling operations on the convolutional features to reduce the data dimension and retain key features. The fully connected layer performs comprehensive processing on the pooled features and maps the features to a new feature space through the operation of weight matrices and bias vectors.

[0186] Through these operations, the second target reconstruction network outputs a new data characterization vector. This vector not only contains the characteristics of the original data, but also integrates the information processed by the previous target reconstruction network, further enriching and refining the characterization of the data. In multiple non-first target reconstruction networks, this process will be repeated in sequence. Each non-first target reconstruction network conducts deeper mining and feature extraction on the data based on the previous network, and continuously improves the data characterization ability. Finally, in step S203, the computer system determines each data maintenance judgment result based on the data characterization vector output by the last target reconstruction network. The data characterization vector output by the last target reconstruction network is obtained after being processed layer by layer by multiple target reconstruction networks. It contains rich feature information about the target terminal maintenance detection data and is the key basis for determining the data maintenance judgment result.

[0187] The computer system determines the inference result based on the predefined rules or agent structure and the data representation vector output by the last target reconstruction network. For example, in the server performance data processing, if the data representation vector output by the last target reconstruction network indicates that the CPU usage of the server has been high for a long time and has an upward trend, the memory usage has also continued to increase, and the disk I / O rate has abnormal fluctuations, the computer system can infer that there is a problem with the current performance of the server according to the pre-set rules, and maintenance operations are required, such as optimizing applications, increasing server resources, etc. Specifically, the computer system can use a classifier or regression agent to infer based on the data representation vector. For example, a simple linear classifier is used to take the data representation vector V output by the last target reconstruction network as input, and a linear transformation is performed through the weight matrix W and the bias vector b to obtain Z=W·V+b. Then, the inference result is determined by a threshold function (for example, for a binary classification problem, when Z is greater than a certain threshold, it is determined to be maintenance-required; when Z is less than the threshold, it is determined to be maintenance-free). For more complex situations, a neural network classifier or regression agent can be used to obtain more accurate inference results through multi-layer nonlinear transformation and decision-making mechanisms.

[0188] Through steps S201-S203, the computer system conducts a comprehensive and in-depth analysis of the target terminal maintenance detection data based on the data maintenance judgment agent. From the initial feature extraction of the original data by the first target reconstruction network, to the sequential in-depth processing of multiple non-first target reconstruction networks, and then to the determination of the data maintenance judgment result based on the output of the last target reconstruction network, this process gradually mines the characteristics of the data, improves the data representation ability, and finally achieves accurate data maintenance judgment. This not only demonstrates the powerful data processing ability of the agent, but also provides a reliable basis for subsequent data maintenance work, which helps government and enterprise users to promptly discover problems in the data and take corresponding maintenance measures to ensure the stable operation of the system and the reliability of the data.

[0189] Figure 3 A hardware entity diagram of a computer system provided by an embodiment of the present invention is as follows Figure 3 As shown, the hardware entity of the computer system 1000 includes: a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can be run on the processor 1001, and the processor 1001 implements the steps in the method of any of the above embodiments when executing the program.

[0190] The above description is only an implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A data maintenance method based on big data and artificial intelligence, characterized in that: The method comprises: Obtain maintenance and detection data of the target terminal, including CPU usage, memory usage, and network bandwidth; Based on the data maintenance judgment agent, at least one data maintenance judgment operation is performed on the target terminal maintenance detection data to obtain various data maintenance judgment results; The data maintenance and discrimination agent is debugged by the following steps to obtain: Acquire a plurality of target data processing agents, any one of which includes a plurality of representation information processing networks, and the network module architecture composed of the plurality of representation information processing networks included in each target data processing agent is consistent, and any one of the target data processing agents is used to perform a data maintenance discrimination operation on the target terminal maintenance detection data; The representation information processing networks in the same architecture node in the network module architecture of the plurality of target data processing agents are merged, and each representation information processing network of a node is merged to obtain a first fused network; Generating a first data processing agent according to each first fusion network; Debugging the first data processing agent to obtain a data maintenance determination agent, wherein the data maintenance determination agent is used to perform at least one data maintenance determination operation among the data maintenance determination operations corresponding to each target data processing agent on the target terminal maintenance detection data; Wherein, any one of the representation information processing networks includes several sub-processing networks, any one of the sub-processing networks includes network weights and biases of multiple feature paths, and the one first fusion network includes several sub-fusion networks; the representation information processing networks of the one node are fused to obtain a first fusion network, including: For any sub-processing network included in any one of the representation information processing networks, among the network weights and biases of each feature path included in any one of the sub-processing networks, determine the network weight and bias of the selected feature path based on the significance measure of the feature path; For any sub-processing network included in each characterization information processing network of the node, based on the network weights and biases of the selected feature paths corresponding to any of the sub-processing networks, weighted averages are performed on multiple network weights and multiple biases to generate a sub-fusion network.

2. The method according to claim 1, characterized in that Determining the network weight and bias of the selected feature path from the network weights and biases of each feature path included in any one of the sub-processing networks comprises: Performing a normalization operation on the network weights and biases of each characteristic path included in any one of the sub-processing networks to obtain normalized network weights and biases of each characteristic path; Determining a significance measure for each feature path based on the standardized network weights and biases of each feature path; Determine a selected feature path whose significance metric is greater than or equal to a preset value among the feature paths, and obtain a network weight and bias of the selected feature path; The step of generating a first data processing agent according to each first fusion network includes: For any first fusion network, generating a first reconstruction network according to the any first fusion network and the initial adjustment network, wherein the initial adjustment network is used to adjust the output of the first reconstruction network; A first data processing agent is generated according to each first reconstruction network.

3. The method according to claim 1 or 2, characterized in that: The step of debugging the first data processing agent to obtain a data maintenance determination agent comprises: Acquire terminal maintenance detection training data and priori labels of each target data processing agent, the terminal maintenance detection training data corresponding to any target data processing agent is used to train the any target data processing agent, and the priori label of the any target data processing agent is used to annotate the maintenance result obtained by performing a data maintenance discrimination operation on the terminal maintenance detection training data; For the terminal maintenance detection training data of any target data processing agent, based on the first data processing agent, a corresponding data maintenance discrimination operation is performed on the terminal maintenance detection training data of any target data processing agent to obtain an inference result corresponding to the any target data processing agent; Based on the reasoning results and prior labels corresponding to each target data processing agent, the first data processing agent is trained to obtain a data maintenance discrimination agent.

4. The method according to claim 3, characterized in that The first reconstruction networks are cascaded, and the first data processing agent is used to perform a corresponding data maintenance discrimination operation on the terminal maintenance detection training data of any target data processing agent to obtain an inference result corresponding to any target data processing agent, including: For a first first reconstruction network, extracting representation information from the terminal maintenance detection training data of any target data processing agent based on the first first reconstruction network to obtain a data representation vector output by the first first reconstruction network; For a non-first first reconstruction network, extracting representation information from a data representation vector output by a previous first reconstruction network of the non-first first reconstruction network based on the non-first first reconstruction network, to obtain a data representation vector output by the non-first first reconstruction network; According to the data representation vector output by the last first reconstruction network, the reasoning result corresponding to any one of the target data processing agents is determined.

5. The method according to claim 4, characterized in that The first first reconstruction network includes a plurality of first reconstruction networks, wherein the first fusion network in the first reconstruction network includes an association perception unit, an intermediate processing unit, and a result mapping unit; the extracting of representation information from the terminal maintenance detection training data of any target data processing agent based on the first first reconstruction network to obtain a data representation vector output by the first first reconstruction network includes: For any first reconstruction network included in the first reconstruction network, based on the association perception unit, an association perception operation is performed on the terminal maintenance detection training data of any target data processing agent to obtain an association perception result; based on the intermediate processing unit, a feature fusion processing is performed on the association perception result to obtain a feature fusion processing result; based on the initial adjustment network, the feature fusion processing result is adjusted according to the association perception result to obtain an adjustment result; based on the result mapping unit, an output result is determined according to the adjustment result; According to the output results of each first reconstruction network included in the first reconstruction network, a data representation vector output by the first reconstruction network is determined.

6. The method according to claim 5, characterized in that The initial adjustment network includes a feature conversion unit and a nonlinear transformation unit, and the feature fusion processing result includes path processing results of multiple feature paths; The step of adjusting the feature fusion processing result based on the initial adjustment network according to the associated perception result to obtain the adjustment result includes: Performing feature conversion on the associated perception result based on the feature conversion unit to obtain feature conversion results of each feature path; Based on the nonlinear transformation unit, a nonlinear transformation is performed on the feature conversion results of each feature path to obtain a nonlinear transformation result of each feature path, wherein the nonlinear transformation result of the feature path is used to characterize the activity of the feature path; Determining an adjustment result according to the nonlinear transformation results of each characteristic path and the path processing results of the multiple characteristic paths; The step of training the first data processing agent based on the inference results and priori labels corresponding to each target data processing agent to obtain a data maintenance discrimination agent includes: Based on the first training cost and the second training cost corresponding to each target data processing agent, the first data processing agent is trained to obtain a data maintenance discrimination agent; wherein the first training cost corresponding to any target data processing agent is determined based on the nonlinear transformation results of each feature path corresponding to each first reconstruction network; the second training cost corresponding to any target data processing agent is determined according to the reasoning result and prior label corresponding to any target data processing agent.

7. The method according to claim 3, characterized in that The step of debugging the first data processing agent to obtain a data maintenance determination agent comprises: Training each initial adjustment network in the first data processing agent to obtain each first adjustment network; The first fusion networks and the first adjustment networks are trained alternately to obtain a data maintenance and discrimination agent.

8. The method according to claim 1, characterized in that The data maintenance discrimination agent includes a plurality of target reconstruction networks; the data maintenance discrimination agent performs at least one data maintenance discrimination operation on the target terminal maintenance detection data to obtain various data maintenance discrimination results, including: For a first target reconstruction network, extracting representation information of the target terminal maintenance detection data based on the first target reconstruction network to obtain a data representation vector output by the first target reconstruction network; For a non-first target reconstruction network, extracting representation information from a data representation vector output by a previous target reconstruction network of the non-first target reconstruction network based on the non-first target reconstruction network to obtain a data representation vector output by the non-first target reconstruction network; According to the data representation vector output by the last target reconstruction network, each data maintenance judgment result is determined.

9. A computer system comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Equipment defect prediction method and system based on time sequence fusion and neural network

    CN115221942A

  • Network flow monitoring method and system based on finite state of sudden change model

    CN116488912A