Data generation method, apparatus, device and medium for computing cluster
By building a data generation method for computing clusters, the problem of missing data simulation of computing clusters is solved, efficient and accurate data generation and simulation are achieved, and fault repair efficiency and system availability are improved.
Patent Information
- Application Number
- CN202111162285.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-09-30
AI Technical Summary
The prior art lacks methods specifically used for calculating cluster data simulation, resulting in limited execution of task scenarios such as fault detection and fault recovery.
By acquiring the pending equipment, building a computing cluster, generating a topology, collecting the initial state, operation data and final state of the device, generating training data, building an initial model based on the variational autoencoder, using training data to obtain a data generation model, collecting the pending data and building a computing cluster data group through the model.
It realizes efficient and accurate simulation of computing cluster data, and combines artificial intelligence to automatically generate data, avoids adverse effects on the normal work of the production line, and improves fault repair efficiency and system availability.
Smart Images

Figure CN113919420B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and medium for generating data of a computing cluster. Background Art
[0002] When performing fault detection, fault recovery and research and development of related systems for a computing cluster, a significant difficulty lies in that there is no appropriate environment for the research and development, testing, data collection and model training of related functions.
[0003] For example: The commonly adopted method is to carry out the research and development and testing of corresponding algorithms in the production line environment. Since the production line environment is relatively stable, the data used is real but relatively one-sided, mostly system data in a safe or stable state, and the amount of abnormal data is relatively small, which is difficult to be used for the training of models such as fault detection or fault location. At the same time, it may also lead to serious safety hazards and even cause production line accidents.
[0004] However, in the existing technical solutions, there is no special method for simulating the data of a computing cluster, resulting in limited execution of task scenarios such as fault detection and fault recovery. Summary of the Invention
[0005] In view of the above, it is necessary to provide a method, device, equipment and medium for generating data of a computing cluster, aiming to solve the problem of simulating the data of a computing cluster.
[0006] A method for generating data of a computing cluster, the method for generating data of the computing cluster includes:
[0007] Obtain a device to be processed, and construct a computing cluster according to the device to be processed;
[0008] Generate a topological structure of the computing cluster;
[0009] Collect the initial state of each device to be processed in the computing cluster, each operation data executed on the computing cluster, and the final state of each device to be processed after executing each operation data;
[0010] Generate training data according to the topological structure, the initial state of each device to be processed, each operation data and the final state of each device to be processed;
[0011] Construct an initial model based on a variational autoencoder;
[0012] Train the initial model by using the training data to obtain a data generation model;
[0013] Collect data to be processed, input the data to be processed into the data generation model, and construct a computing cluster data group according to the output of the data generation model.
[0014] According to a preferred embodiment of the present invention, the generation of the topological structure of the computing cluster includes:
[0015] Obtain the call relationship between the devices to be processed;
[0016] Determine each device to be processed as a node;
[0017] Construct a directed edge between the nodes according to the call relationship between the devices to be processed;
[0018] Number each node in the configured order to obtain the topological structure.
[0019] According to a preferred embodiment of the present invention, the generation of training data based on the topological structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed includes:
[0020] Encode each operation data to obtain an operation vector;
[0021] For each node in the topological structure, perform vectorization processing on each node according to the initial state of each device to be processed to generate an initial state signal for each node;
[0022] Construct a matrix based on the topological structure and the initial state signal of each node to obtain a set of first state data of the computing cluster; wherein, in the set of first state data, each first state data corresponds to each operation vector;
[0023] Perform vectorization processing on each node according to the final state of each device to be processed to generate a final state signal for each node;
[0024] Construct a matrix based on the topological structure and the final state signal of each node to obtain a set of second state data of the computing cluster; wherein, in the set of second state data, each second state data corresponds to each operation vector;
[0025] Divide the first state data, the second state data, and the same operation vector corresponding to the same operation vector into a group to obtain at least one data group;
[0026] Integrate the at least one data group to obtain the training data.
[0027] According to a preferred embodiment of the present invention, the construction of the initial model based on the variational autoencoder includes:
[0028] Obtain the output data of the encoder in the variational autoencoder, and obtain the mean signal and the variance signal from the output data;
[0029] Obtain random noise;
[0030] Fuse the mean signal and the variance signal with the random noise to obtain a noise code;
[0031] Add an operation coding layer to the variational autoencoder and deploy the operation vector on the operation coding layer;
[0032] Combine the noise code and the operation vector to obtain a hidden vector;
[0033] Input the hidden vector into the decoder in the variational autoencoder to obtain the initial model.
[0034] According to a preferred embodiment of the present invention, training the initial model using the training data to obtain a data generation model includes:
[0035] Construct a target loss function;
[0036] Use the target loss function and perform gradient descent training on the initial model based on the training data;
[0037] When the value of the target loss function is less than or equal to a preset threshold, stop training and determine the currently obtained model as the data generation model.
[0038] According to a preferred embodiment of the present invention, constructing the target loss function includes:
[0039] Calculate the output loss component using the following formula:
[0040]
[0041] where L output represents the output component loss, X' represents the second state data obtained from the training data, and X' output represents the second state data output by the model during training;
[0042] Calculate the hidden layer loss component using the following formula:
[0043]
[0044] where L latent represents the hidden layer loss component, I represents a constant matrix, Z log_var represents the output data of the variance output layer of the model, and Z mean represents the output data of the mean output layer of the model;
[0045] Calculate the sum of the output loss component and the hidden layer loss component to obtain the target loss function.
[0046] According to a preferred embodiment of the present invention, before inputting the data to be processed into the data generation model, the method further includes:
[0047] Identifying the current task scenario;
[0048] When the current task scenario requires generating abnormal data, obtaining backup processing devices from the devices to be processed;
[0049] Configuring the initial states of the backup processing devices to be abnormal.
[0050] A data generation device for a computing cluster, the data generation device for the computing cluster includes:
[0051] A construction unit, configured to obtain devices to be processed and construct a computing cluster according to the devices to be processed;
[0052] A generation unit, configured to generate a topology structure of the computing cluster;
[0053] An acquisition unit, configured to acquire the initial state of each device to be processed in the computing cluster, each operation data executed on the computing cluster, and the final state of each device to be processed after executing each operation data;
[0054] The generation unit is further configured to generate training data according to the topology structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed;
[0055] The construction unit is further configured to construct an initial model based on a variational autoencoder;
[0056] A training unit, configured to train the initial model by using the training data to obtain a data generation model;
[0057] The construction unit is further configured to acquire data to be processed, input the data to be processed into the data generation model, and construct a computing cluster data group according to the output of the data generation model.
[0058] A computer device, the computer device includes:
[0059] A memory, storing at least one instruction; and
[0060] A processor, executing the instruction stored in the memory to implement the data generation method for the computing cluster.
[0061] A computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in a computer device to implement the data generation method for the computing cluster.
[0062] As can be seen from the above technical solutions, the present invention can obtain a device to be processed, construct a computing cluster based on the device to be processed, generate a topological structure of the computing cluster, and the generated topological structure can clearly reflect the relationship between each device to be processed in the computing cluster. The initial state of each device to be processed, each operation data executed on the computing cluster, and the final state of each device to be processed after executing each operation data are collected from the computing cluster. Training data is generated according to the topological structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed. An initial model is constructed based on a variational autoencoder, and the initial model is trained using the training data to obtain a data generation model. The data to be processed is collected, input into the data generation model, and a computing cluster data group is constructed according to the output of the data generation model. It is possible to realize the simulation of data through the model, and then automatically generate data by combining artificial intelligence means, which is efficient and accurate. While generating a large amount of data, it effectively avoids having an adverse impact on the normal operation of the production line. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a flowchart of a preferred embodiment of the data generation method for the computing cluster of the present invention.
[0064] Figure 2 It is a functional module diagram of a preferred embodiment of the data generation device for the computing cluster of the present invention.
[0065] Figure 3 It is a schematic structural diagram of a computer device of a preferred embodiment for implementing the data generation method for the computing cluster of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0067] As Figure 1 shown, it is a flowchart of a preferred embodiment of the data generation method for the computing cluster of the present invention. According to different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted.
[0068] The data generation method of the computing cluster is applied to one or more computer devices, which are devices capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0069] The computer device can be any electronic product that can interact with users. For example, personal computers, tablets, smartphones, personal digital assistants (PDAs), game consoles, Internet Protocol Television (IPTV), smart wearable devices, etc.
[0070] The computer device may also include network devices and / or user devices. Among them, the network devices include, but are not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing (Cloud Computing).
[0071] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0072] Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0073] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0074] The network where the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0075] S10. Obtain the device to be processed, and construct a computing cluster according to the device to be processed.
[0076] In at least one embodiment of the present invention, the device to be processed may include multiple virtual machines or multiple servers, and the present invention does not limit this.
[0077] In at least one embodiment of the present invention, the devices to be processed are formed into a group to obtain the computing cluster.
[0078] S11. Generate the topological structure of the computing cluster.
[0079] It can be understood that there is a certain call relationship between the devices to be processed. Therefore, the computing cluster can be represented in the form of graph data.
[0080] In at least one embodiment of the present invention, generating the topological structure of the computing cluster includes:
[0081] Obtain the call relationship between the devices to be processed;
[0082] Determine each device to be processed as a node;
[0083] Construct directed edges between the nodes according to the call relationship between the devices to be processed;
[0084] Number each node in the configured order to obtain the topological structure.
[0085] In this embodiment, the configured order may be the working serial number of the device to be processed, or it may be custom-configured, and the present invention does not limit this.
[0086] For example: the generated topological structure may be in matrix form, denoted as: Where N is the total number of nodes.
[0087] Through the above implementation manner, the generated topological structure can clearly reflect the relationship between each device to be processed in the computing cluster.
[0088] S12. Collect the initial state of each device to be processed, each operation data executed on the computing cluster, and the final state of each device to be processed after executing each operation data from the computing cluster.
[0089] In at least one embodiment of the present invention, the initial state refers to the state of the corresponding device to be processed before a specified operation is performed, such as indicators like the CPU (central processing unit) usage rate, memory usage rate, and whether it is in the powered-on state of the corresponding device to be processed.
[0090] In at least one embodiment of the present invention, each operation data performed on the computing cluster refers to an operation on the computing cluster, such as powering on, powering off, stopping a computing task, etc.
[0091] In at least one embodiment of the present invention, the final state refers to the state of the corresponding device to be processed after a specified operation is performed, such as indicators like the CPU usage rate, memory usage rate, and whether it is in the powered-on state of the corresponding device to be processed after the power-on operation is performed.
[0092] S13. Generate training data according to the topology structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed.
[0093] In at least one embodiment of the present invention, generating training data according to the topology structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed includes:
[0094] Encode each operation data to obtain an operation vector;
[0095] For each node in the topology structure, perform vectorization processing on each node according to the initial state of each device to be processed to generate an initial state signal for each node;
[0096] Construct a matrix based on the topology structure and the initial state signals of each node to obtain a set of first state data of the computing cluster; wherein, in the set of first state data, each first state data corresponds to each operation vector;
[0097] Perform vectorization processing on each node according to the final state of each device to be processed to generate a final state signal for each node;
[0098] Construct a matrix based on the topology structure and the final state signals of each node to obtain a set of second state data of the computing cluster; wherein, in the set of second state data, each second state data corresponds to each operation vector;
[0099] Divide the first state data, the second state data, and the same operation vector corresponding to the same operation vector into a group to obtain at least one data group;
[0100] Integrate the at least one data group to obtain the training data.
[0101] For example, when vectorizing each node according to the initial state of each device to be processed, when it is determined that the CPU usage rate of node A is 20%, it can be recorded as 0.2; when the memory usage rate of node A is 50%, it can be recorded as 0.5; when node A is in the powered-on state, it can be recorded as 1; when node A is in the powered-off state, it can be recorded as 0. Further, horizontally splice the values corresponding to each state to obtain the initial state signal of node A.
[0102] Further, in the case where the numbers and topologies of the nodes in the computing cluster are determined, the first state data of the entire cluster can be represented as a set of initial state signals on all nodes, denoted as matrix X ∈ R N×D , where D is the signal dimension. For example, when there are 100 nodes and each node corresponds to 10 initial states, the value of the dimension D of the matrix corresponding to the first state data is 100 * 10. In the matrix, the i-th row represents the initial state signal of the i-th node.
[0103] In this embodiment, the one-hot encoding algorithm can be used to encode each operation data to obtain the operation vector.
[0104] For example: the encoding of power-on is represented as 001, the encoding of power-off is represented as 010, and the encoding of stopping the computing task is represented as 001.
[0105] Further, the operations performed on the entire cluster can be represented as a set of operations on all nodes, denoted as matrix A ∈ R N×K , where K is the total number of operation types. In the matrix, the element in the i-th row and j-th column represents whether the j-th type of operation is performed on the i-th node.
[0106] In this embodiment, after performing the corresponding operations on the computing cluster, the second state data of the computing cluster can be obtained, denoted as matrix X' ∈ R N×D .
[0107] Further, group the first state data, the second state data corresponding to the same operation vector, and the same operation vector together. At this time, a group of data obtained can be denoted as (X, A, X'), where X represents the first state data, A represents the operation vector, and X' represents the second state data.
[0108] S14. Construct an initial model based on the Variational Auto-Encoder (VAE).
[0109] Specifically, constructing the initial model based on the variational autoencoder includes:
[0110] Obtain the output data of the encoder in the variational autoencoder, and obtain the mean signal and variance signal from the output data;
[0111] Obtain random noise;
[0112] Fuse the mean signal and the variance signal with the random noise to obtain a noise code;
[0113] Add an operation coding layer in the variational autoencoder, and deploy the operation vector on the operation coding layer;
[0114] Combine the noise code and the operation vector to obtain a hidden vector;
[0115] Input the hidden vector into the decoder in the variational autoencoder to obtain the initial model.
[0116] Specifically, in the initial model, the encoder and decoder in the variational autoencoder can use a graph convolutional neural network as the basic structure.
[0117] Specifically, in the encoder, taking the graph convolutional neural network with a single hidden layer as an example, its main structure can be expressed as:
[0118] Input layer: Input the first state data X of the computing cluster;
[0119] Hidden layer:
[0120] Mean output layer:
[0121] Variance output layer:
[0122] Among them, W1, W2, and W3 are weight coefficients, and σ(·) is an activation function.
[0123] Furthermore, superimpose the noise signal and the operation signal:
[0124] Sample the Gaussian noise ε ∼ N(0, 1) from the Gaussian distribution as the random noise, and fuse it with the mean signal Z mean and the variance signal Z log_var to obtain the noise code as:
[0125]
[0126] Combine the noise code with the operation vector A ∈ R N×K, obtain the latent vector Z, representing the encoding of the system state - system operation:
[0127] Z = [Z noise |A],
[0128] Furthermore, in the decoder, taking the graph convolutional neural network with a single hidden layer as an example, its main structure can be expressed as:
[0129] Input layer: latent vector Z;
[0130] Hidden layer:
[0131] Output layer:
[0132] where W4 and W5 are weight coefficients.
[0133] Specifically, W1, W2, W3, W4, and W5 can be randomly initialized.
[0134] In the above embodiment, an operation encoding layer is added on the basis of the original variational auto - encoder. By changing the structure of the variational auto - encoder, the constructed initial model can be integrated with the actual operation.
[0135] S15. Use the training data to train the initial model to obtain a data generation model.
[0136] In at least one embodiment of the present invention, the using the training data to train the initial model to obtain a data generation model includes:
[0137] Construct a target loss function;
[0138] Use the target loss function and perform gradient - descent training on the initial model based on the training data;
[0139] When the value of the target loss function is less than or equal to a preset threshold, stop training and determine the currently obtained model as the data generation model.
[0140] where the preset threshold can be custom - configured.
[0141] Specifically, the constructing the target loss function includes:
[0142] Calculate the output loss component using the following formula:
[0143]
[0144] where L outputrepresents the loss of the output component, X' represents the second state data obtained from the training data, X' output represents the second state data output by the model during training;
[0145] The hidden layer loss component is calculated using the following formula:
[0146]
[0147] where L latent represents the hidden layer loss component, I represents a constant matrix, Z log_var represents the output data of the variance output layer of the model, Z mean represents the output data of the mean output layer of the model;
[0148] Calculate the sum of the output loss component and the hidden layer loss component to obtain the target loss function.
[0149] For example: The target loss function can be expressed as: L = L output + L latent .
[0150] where I is a constant matrix with all elements being 1 and having the same dimension as Z log_var and Z mean .
[0151] In the above formula, the operation SUM represents summing all elements of the matrix; square(·) represents squaring each element of the matrix respectively, and exp(·) represents calculating the exponential function for each element of the matrix.
[0152] Furthermore, when the training reaches that the loss function remains lower than the preset threshold, the model training ends, and the weight coefficients W1, W2, W3, W4, W5 of the optimized model are obtained.
[0153] Through the above embodiments, it is possible to comprehensively consider the output loss and the hidden layer loss during model training, making the accuracy of the trained model higher.
[0154] S16, collect the data to be processed, input the data to be processed into the data generation model, and construct a computing cluster data group according to the output of the data generation model.
[0155] In at least one embodiment of the present invention, the data to be processed can be randomly obtained from the training data, or data can be collected from the production line environment as the data to be processed.
[0156] Specifically, the step of inputting the data to be processed into the data generation model and constructing a computing cluster data group according to the output of the data generation model includes:
[0157] After the model training is completed, the generated model can be used to generate relevant data. The first state data can be randomly selected from the collected training dataset, or the first state data can be collected in real time from the real production line environment as the initial state of the simulation system (i.e., the initial model), denoted as matrix X origin ;
[0158] Input the initial state X of the simulation system origin into the encoder, and propagate forward layer by layer to obtain the system state encoding, i.e., the mean signal Z mean and the variance signal Z log_var .
[0159] Among them, the output of the hidden layer the output of the mean output layer the output of the variance output layer
[0160] Collect the Gaussian noise signal ε~N(0,1) from the normal distribution, and construct the noisy system state signal and merge it with the operation vector A to be tested test to obtain the superimposed signal Z = [Z noise |A test .
[0161] Input the system state-system operation signal Z into the decoder to obtain the second state data X' generate . At the same time, a set of simulation data is obtained, denoted as (X origin , A test , X' generate ).
[0162]
[0163]
[0164] Repeat the above steps to generate a series of datasets, denoted as {(X origin , A test , X' generate )}, thereby realizing data simulation and data generation.
[0165] Through the above implementation manner, it is possible to simulate and simulate data through the model, and then automatically generate data in combination with artificial intelligence means, which is efficient and accurate. While generating a large amount of data, it effectively avoids adverse effects on the normal operation of the production line.
[0166] In at least one embodiment of the present invention, before inputting the data to be processed into the data generation model, the method further includes:
[0167] Identify the current task scenario;
[0168] When the current task scenario requires generating abnormal data, obtain the devices to be processed that are backed up each other from the devices to be processed.
[0169] Configure the initial states of the devices to be processed that are backed up each other as abnormal.
[0170] For example: When devices A, B, and C are backup devices for each other, abnormal data will only be generated when all three are abnormal. Therefore, configure the initial states of devices A, B, and C as abnormal (such as all being in the shutdown state), then the simulation data generated by the data generation model is abnormal data.
[0171] In the above embodiment, a sufficient amount of abnormal data is generated through the model for subsequent use in fault detection, fault recovery issues, and the research and development of related systems. This not only does not pose hazards or risks to the real production line environment but also can obtain relatively rich and comprehensive abnormal data, further improving the fault repair efficiency of the computing cluster, shortening the fault location and repair time, and thus improving system availability.
[0172] It should be noted that, in order to further improve data security and prevent data from being maliciously tampered with, the data generation model can be stored in the blockchain node.
[0173] It can be seen from the above technical solutions that the present invention can obtain the devices to be processed, construct a computing cluster based on the devices to be processed, generate the topological structure of the computing cluster, and the generated topological structure can clearly reflect the relationship between each device to be processed in the computing cluster. Collect the initial state of each device to be processed, each operation data executed on the computing cluster, and the final state of each device to be processed after executing each operation data from the computing cluster. Generate training data based on the topological structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed. Construct an initial model based on the variational autoencoder, and use the training data to train the initial model to obtain a data generation model. Collect the data to be processed, input the data to be processed into the data generation model, and construct a computing cluster data group according to the output of the data generation model. It can simulate and imitate data through the model, and then automatically generate data in combination with artificial intelligence means, which is efficient and accurate. While generating a large amount of data, it effectively avoids having an adverse impact on the normal operation of the production line.
[0174] Such as Figure 2As shown in the figure, it is a functional module diagram of a preferred embodiment of the data generation device of the computing cluster of the present invention. The data generation device 11 of the computing cluster includes a construction unit 110, a generation unit 111, a collection unit 112, and a training unit 113. The modules / units referred to in the present invention refer to a series of computer program segments that can be executed by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0175] The construction unit 110 obtains the device to be processed and constructs a computing cluster according to the device to be processed.
[0176] In at least one embodiment of the present invention, the device to be processed may include multiple virtual machines or multiple servers, and the present invention does not limit this.
[0177] In at least one embodiment of the present invention, the devices to be processed are formed into a group to obtain the computing cluster.
[0178] The generation unit 111 generates the topological structure of the computing cluster.
[0179] It can be understood that there is a certain call relationship between the devices to be processed. Therefore, the computing cluster can be represented in the form of graph data.
[0180] In at least one embodiment of the present invention, the generation unit 111 generating the topological structure of the computing cluster includes:
[0181] Obtain the call relationship between the devices to be processed;
[0182] Determine each device to be processed as a node;
[0183] Construct a directed edge between the nodes according to the call relationship between the devices to be processed;
[0184] Number each node in the configured order to obtain the topological structure.
[0185] In this embodiment, the configured order may be the working serial number of the device to be processed, or it may be custom-configured, and the present invention does not limit this.
[0186] For example: the generated topological structure may be in matrix form, denoted as: Where N is the total number of nodes.
[0187] Through the above implementation manner, the generated topological structure can clearly reflect the relationship between each device to be processed in the computing cluster.
[0188] The acquisition unit 112 acquires the initial state of each device to be processed, each operation data executed on the computing cluster, and the final state of each device to be processed after executing each operation data from the computing cluster.
[0189] In at least one embodiment of the present invention, the initial state refers to the state of the corresponding device to be processed before executing the specified operation, such as: indicators such as the CPU (central processing unit) usage rate, memory usage rate, and whether it is in the powered-on state of the corresponding device to be processed.
[0190] In at least one embodiment of the present invention, each operation data executed on the computing cluster refers to an operation on the computing cluster, such as: power on, power off, stop computing tasks, etc.
[0191] In at least one embodiment of the present invention, the final state refers to the state of the corresponding device to be processed after executing the specified operation, such as: indicators such as the CPU usage rate, memory usage rate, and whether it is in the powered-on state of the corresponding device to be processed after executing the power-on operation.
[0192] The generation unit 111 generates training data according to the topological structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed.
[0193] In at least one embodiment of the present invention, the generation unit 111 generating training data according to the topological structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed includes:
[0194] Encoding each operation data to obtain an operation vector;
[0195] For each node in the topological structure, vectorizing each node according to the initial state of each device to be processed to generate an initial state signal for each node;
[0196] Based on the topological structure and the initial state signals of each node, constructing a matrix to obtain a set of first state data of the computing cluster; wherein, in the set of first state data, each first state data corresponds to each operation vector;
[0197] Vectorizing each node according to the final state of each device to be processed to generate a final state signal for each node;
[0198] Based on the topological structure and the final state signals of each node, constructing a matrix to obtain a set of second state data of the computing cluster; wherein, in the set of second state data, each second state data corresponds to each operation vector;
[0199] Divide the first state data, the second state data, and the same operation vector corresponding to the same operation vector into a group to obtain at least one data group;
[0200] Integrate the at least one data group to obtain the training data.
[0201] For example, when performing vectorization processing on each node according to the initial state of each device to be processed, when it is determined that the CPU usage rate of node A is 20%, it can be recorded as 0.2; when the memory usage rate of node A is 50%, it can be recorded as 0.5; when node A is in the powered-on state, it can be recorded as 1; when node A is in the powered-off state, it can be recorded as 0. Further, horizontally splice the values corresponding to each state to obtain the initial state signal of node A.
[0202] Further, when the numbers and topologies of the nodes in the computing cluster are determined, the first state data of the entire cluster can be represented as a set of initial state signals on all nodes, denoted as matrix X ∈ R N×D , where D is the signal dimension. For example, when there are 100 nodes and each node corresponds to 10 initial states, the value of the dimension D of the matrix corresponding to the first state data is 100 * 10. In the matrix, the i-th row represents the initial state signal of the i-th node.
[0203] In this embodiment, the one-hot encoding algorithm can be used to encode each operation data to obtain the operation vector.
[0204] For example: the encoding of power-on is represented as 001, the encoding of power-off is represented as 010, and the encoding of stopping the computing task is represented as 001.
[0205] Further, the operations performed on the entire cluster can be represented as a set of operations on all nodes, denoted as matrix A ∈ R N×K , where K is the total number of operation types. In the matrix, the element in the i-th row and j-th column represents whether the j-th type of operation is performed on the i-th node.
[0206] In this embodiment, after performing the corresponding operations on the computing cluster, the second state data of the computing cluster can be obtained, denoted as matrix X' ∈ R N×D .
[0207] Further, divide the first state data, the second state data, and the same operation vector corresponding to the same operation vector into a group. At this time, a group of data obtained can be denoted as (X, A, X'), where X represents the first state data, A represents the operation vector, and X' represents the second state data.
[0208] The building unit 110 constructs an initial model based on a Variational Auto-Encoder (VAE).
[0209] Specifically, the building unit 110 constructing an initial model based on a Variational Auto-Encoder includes:
[0210] Obtain the output data of the encoder in the Variational Auto-Encoder, and obtain the mean signal and variance signal from the output data;
[0211] Obtain random noise;
[0212] Fuse the mean signal and the variance signal with the random noise to obtain a noise code;
[0213] Add an operation coding layer in the Variational Auto-Encoder, and deploy the operation vector on the operation coding layer;
[0214] Merge the noise code and the operation vector to obtain a hidden vector;
[0215] Input the hidden vector into the decoder in the Variational Auto-Encoder to obtain the initial model.
[0216] Specifically, in the initial model, the encoder and decoder in the Variational Auto-Encoder can use a graph convolutional neural network as the basic structure.
[0217] Specifically, in the encoder, taking a graph convolutional neural network with a single hidden layer as an example, its main structure can be expressed as:
[0218] Input layer: Input the first state data X of the computing cluster;
[0219] Hidden layer:
[0220] Mean output layer:
[0221] Variance output layer:
[0222] Among them, W1, W2, and W3 are weight coefficients, and σ(·) is an activation function.
[0223] Furthermore, superimpose the noise signal and the operation signal:
[0224] Sample Gaussian noise ε~N(0,1) from the Gaussian distribution as the random noise, and fuse it with the mean signal Z mean and variance signal Z log_var of the encoder output to obtain the noise code as:
[0225]
[0226] Merge the noise encoding with the operation vector A ∈ R N×K , to obtain a latent vector Z, representing the encoding of the system state - system operation:
[0227] Z = [Z noise |A],
[0228] Furthermore, in the decoder, taking the graph convolutional neural network with a single hidden layer as an example, its main structure can be expressed as:
[0229] Input layer: latent vector Z;
[0230] Hidden layer:
[0231] Output layer:
[0232] where W4 and W5 are weight coefficients.
[0233] Specifically, W1, W2, W3, W4, and W5 can be randomly initialized.
[0234] In the above embodiment, an operation encoding layer is added on the basis of the original variational auto - encoder. By changing the structure of the variational auto - encoder, the constructed initial model can be integrated with the actual operation.
[0235] The training unit 113 trains the initial model using the training data to obtain a data generation model.
[0236] In at least one embodiment of the present invention, the training unit 113 trains the initial model using the training data to obtain a data generation model, including:
[0237] Construct a target loss function;
[0238] Use the target loss function and perform gradient - descent training on the initial model based on the training data;
[0239] When the value of the target loss function is less than or equal to a preset threshold, stop training and determine the currently obtained model as the data generation model.
[0240] Among them, the preset threshold can be custom - configured.
[0241] Specifically, the constructing of the target loss function includes:
[0242] Calculate the output loss component using the following formula:
[0243]
[0244] Among them, L output represents the loss of the output component, X' represents the second state data obtained from the training data, and X' output represents the second state data output by the model during the training process;
[0245] The hidden layer loss component is calculated using the following formula:
[0246]
[0247] Among them, L latent represents the hidden layer loss component, I represents a constant matrix, and Z log_var represents the output data of the variance output layer of the model, and Z mean represents the output data of the mean output layer of the model;
[0248] Calculate the sum of the output loss component and the hidden layer loss component to obtain the target loss function.
[0249] For example: the target loss function can be expressed as: L = L output + L latent .
[0250] Among them, I is a constant matrix with all elements being 1 and having the same dimension as Z log_var and Z mean .
[0251] In the above formula, the operation SUM represents summing all elements of the matrix; square(·) represents squaring each element of the matrix respectively, and exp(·) represents calculating the exponential function for each element of the matrix.
[0252] Furthermore, when the training reaches that the loss function remains lower than the preset threshold, the model training ends, and the weight coefficients W1, W2, W3, W4, W5 of the optimized model are obtained.
[0253] Through the above embodiments, it is possible to comprehensively consider the output loss and the hidden layer loss during the model training, so that the accuracy of the trained model is higher.
[0254] The building unit 110 collects the data to be processed, inputs the data to be processed into the data generation model, and constructs a computing cluster data group according to the output of the data generation model.
[0255] In at least one embodiment of the present invention, the data to be processed can be randomly obtained from the training data, or the data can be collected from the production line environment as the data to be processed.
[0256] Specifically, the building unit 110 inputs the data to be processed into the data generation model, and constructs a computing cluster data group according to the output of the data generation model, including:
[0257] After the model training is completed, the generated model can be used to generate relevant data. The first state data can be randomly selected from the collected training data set, or the first state data can be collected in real time from the real production line environment as the initial state of the simulation system (i.e., the initial model), denoted as matrix X origin ;
[0258] Input the initial state X of the simulation system origin into the encoder, and forward propagate layer by layer to obtain the system state encoding, that is, the mean signal Z mean and the variance signal Z log_var .
[0259] Among them, the hidden layer output the mean output layer output the variance output layer output
[0260] Collect the Gaussian noise signal ε~N(0, 1) from the normal distribution, and construct a noisy system state signal and merge it with the operation vector A to be tested test to obtain the superimposed signal Z = [Z noise |A test .
[0261] Input the system state-system operation signal Z into the decoder to obtain the second state data X' generate . At the same time, a set of simulation data is obtained, denoted as (X origin , A test , X' generate ).
[0262]
[0263]
[0264] Repeat the above steps to generate a series of data sets, denoted as {(X origin , A test , X' generate )}, thereby realizing data simulation and data generation.
[0265] Through the above implementation methods, it is possible to simulate and simulate data through the model, and then automatically generate data in combination with artificial intelligence means, which is efficient and accurate. While generating a large amount of data, it effectively avoids adverse effects on the normal operation of the production line.
[0266] In at least one embodiment of the present invention, before inputting the data to be processed into the data generation model, the current task scenario is identified;
[0267] When the current task scenario requires generating abnormal data, backup processing devices that are mutually backed up are obtained from the processing devices to be processed;
[0268] The initial states of the mutually backed up processing devices to be processed are configured as abnormal.
[0269] For example: when devices A, B, and C are mutually backup devices, abnormal data will only be generated when all three are abnormal. Therefore, if the initial states of devices A, B, and C are all configured as abnormal (such as all being in a shutdown state), the simulation data generated by the data generation model is abnormal data.
[0270] In the above embodiment, a sufficient amount of abnormal data is generated through the model for subsequent use in fault detection, fault recovery problems, and the research and development of related systems. This not only does not pose any harm or risk to the real production line environment but also can obtain relatively rich and comprehensive abnormal data, further improving the fault repair efficiency of the computing cluster, shortening the fault location and repair time, and thus improving system availability.
[0271] It should be noted that, in order to further improve data security and prevent data from being maliciously tampered with, the data generation model can be stored in a blockchain node.
[0272] From the above technical solutions, it can be seen that the present invention can obtain processing devices to be processed, construct a computing cluster based on the processing devices to be processed, generate the topological structure of the computing cluster, and the generated topological structure can clearly reflect the relationship between each processing device in the computing cluster. The initial state of each processing device in the computing cluster, each operation data executed on the computing cluster, and the final state of each processing device after executing each operation data are collected. Training data is generated based on the topological structure, the initial state of each processing device, each operation data, and the final state of each processing device. An initial model is constructed based on a variational autoencoder, and the initial model is trained using the training data to obtain a data generation model. The data to be processed is collected, the data to be processed is input into the data generation model, and a computing cluster data group is constructed based on the output of the data generation model. It can simulate and emulate data through the model, and then automatically generate data in combination with artificial intelligence means, which is efficient and accurate. While generating a large amount of data, it effectively avoids having an adverse impact on the normal operation of the production line.
[0273] Such as Figure 3As shown, it is a schematic structural diagram of a computer device of a preferred embodiment of the data generation method for a computing cluster according to the present invention.
[0274] The computer device 1 may include a memory 12, a processor 13, and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a data generation program for a computing cluster.
[0275] Those skilled in the art can understand that the schematic diagram is only an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 can be either a bus structure or a star structure. The computer device 1 may also include more or fewer other hardware or software than shown, or different component arrangements. For example, the computer device 1 may also include input / output devices, network access devices, etc.
[0276] It should be noted that the computer device 1 is only an example. Other existing or future electronic products that can be adapted to the present invention should also be included within the protection scope of the present invention and are hereby incorporated by reference.
[0277] Among them, the memory 12 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 12 may be an internal storage unit of the computer device 1 in some embodiments, such as the mobile hard disk of the computer device 1. The memory 12 may also be an external storage device of the computer device 1 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 1. Further, the memory 12 may include both an internal storage unit and an external storage device of the computer device 1. The memory 12 can be used not only to store application software installed on the computer device 1 and various types of data, such as the code of the data generation program for a computing cluster, but also to temporarily store data that has been output or will be output.
[0278] In some embodiments, the processor 13 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control core of the computer device 1, connecting various components of the entire computer device 1 through various interfaces and circuits. By running or executing programs or modules stored in the memory 12 (such as executing a data generation program for a computing cluster), and by calling data stored in the memory 12, it executes various functions of the computer device 1 and processes data.
[0279] The processor 13 executes the operating system of the computer device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the data generation method embodiments of the above-mentioned various computing clusters, such as Figure 1 the steps shown.
[0280] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a construction unit 110, a generation unit 111, a collection unit 112, and a training unit 113.
[0281] The above-mentioned integrated units implemented in the form of software function modules may be stored in a computer-readable storage medium. The above-mentioned software function modules are stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the data generation method of the computing cluster in each embodiment of the present invention.
[0282] If the integrated module / unit of the computer device 1 is implemented in the form of a software function unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented.
[0283] Among them, the computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory, etc.
[0284] Furthermore, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area can store an operating system, application programs required for at least one function, etc.; the storage data area can store data created according to the use of the blockchain node, etc.
[0285] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, an application service layer, etc.
[0286] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, in Figure 3 it is only represented by a single straight line, but it does not mean that there is only one bus or one type of bus. The bus is set to realize the connection and communication between the memory 12 and at least one processor 13, etc.
[0287] Although not shown, the computer device 1 may further include a power supply (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, so as to realize functions such as charging management, discharging management, and power consumption management through the power management device. The power supply may further include any components such as one or more DC or AC power supplies, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The computer device 1 may further include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0288] Further, the computer device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the computer device 1 and other computer devices.
[0289] Optionally, the computer device 1 may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the computer device 1 and to display a visual user interface.
[0290] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0291] Figure 3 Only the computer device 1 with components 12 - 13 is shown. Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the computer device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0292] Combined with Figure 1 , the memory 12 in the computer device 1 stores multiple instructions to implement a method for generating data of a computing cluster, and the processor 13 can execute the multiple instructions to implement:
[0293] Obtain a device to be processed, and construct a computing cluster according to the device to be processed;
[0294] Generate a topological structure of the computing cluster;
[0295] Collect the initial state of each device to be processed, each operation data executed on the computing cluster, and the final state of each device to be processed after executing each operation data from the computing cluster;
[0296] Generate training data according to the topological structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed;
[0297] Construct an initial model based on a variational autoencoder;
[0298] Train the initial model using the training data to obtain a data generation model;
[0299] Collect the data to be processed, input the data to be processed into the data generation model, and construct a computing cluster data group according to the output of the data generation model.
[0300] Specifically, for the specific implementation method of the above instructions by the processor 13, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiments, which will not be elaborated here.
[0301] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0302] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0303] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0304] In addition, in each embodiment of the present invention, the various functional modules can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a hardware plus software functional module.
[0305] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0306] Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0307] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described in the present invention can also be implemented by one unit or device through software or hardware. The terms such as "first" and "second" are used to denote names and do not denote any particular order.
[0308] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A data generation method for a computing cluster, characterized in that, The data generation method for the computing cluster includes: Obtain the device to be processed, and construct a computing cluster according to the device to be processed; Generate the topological structure of the computing cluster; Collect the initial state of each device to be processed in the computing cluster, each operation data executed on the computing cluster, and the final state of each device to be processed after executing each operation data; Generate training data according to the topological structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed; Construct an initial model based on a variational autoencoder; Train the initial model using the training data to obtain a data generation model; Collect the data to be processed, input the data to be processed into the data generation model, and construct a computing cluster data group according to the output of the data generation model; Wherein, each operation data executed on the computing cluster refers to an operation on the computing cluster.
2. The data generation method of the computing cluster according to claim 1, characterized in that The generating the topological structure of the computing cluster includes: Obtain the call relationship between the devices to be processed; Determine each device to be processed as a node; Construct a directed edge between the nodes according to the call relationship between the devices to be processed; Number each node in the configured order to obtain the topological structure.
3. The data generation method of the computing cluster according to claim 1, wherein The generating training data according to the topological structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed includes: Encode each operation data to obtain an operation vector; For each node in the topological structure, perform vectorization processing on each node according to the initial state of each device to be processed to generate an initial state signal for each node; Construct a matrix based on the topological structure and the initial state signal of each node to obtain a set of first state data of the computing cluster; wherein, in the set of first state data, each first state data corresponds to each operation vector; Perform vectorization processing on each node according to the final state of each device to be processed to generate a final state signal for each node; Construct a matrix based on the topological structure and the final state signal of each node to obtain a set of second state data of the computing cluster; wherein, in the set of second state data, each second state data corresponds to each operation vector; Divide the first state data, the second state data, and the same operation vector corresponding to the same operation vector into a group to obtain at least one data group; Integrate the at least one data group to obtain the training data.
4. The data generation method of the computing cluster according to claim 3, wherein The constructing an initial model based on a variational autoencoder includes: Obtain the output data of the encoder in the variational autoencoder, and obtain a mean signal and a variance signal from the output data; Obtain random noise; Fuse the mean signal and the variance signal with the random noise to obtain a noise code; Add an operation coding layer in the variational autoencoder, and deploy the operation vector on the operation coding layer; Merge the noise code and the operation vector to obtain a latent vector; Input the latent vector into the decoder in the variational autoencoder to obtain the initial model.
5. The data generation method of the computing cluster according to claim 1, wherein Training the initial model using the training data to obtain a data generation model includes: Constructing an objective loss function; Using the objective loss function and performing gradient descent training on the initial model based on the training data; When the value of the objective loss function is less than or equal to a preset threshold, stop training and determine the currently obtained model as the data generation model.
6. The data generation method of the computing cluster according to claim 5, characterized in that, The constructing of the objective loss function includes: Calculating an output loss component using the following formula: ; Among them, represents the loss of the output component, represents the second state data obtained from the training data, represents the second state data output by the model during the training process; Calculating a hidden layer loss component using the following formula: ; Among them, represents the hidden layer loss component, represents a constant matrix, represents the output data of the variance output layer of the model, represents the output data of the mean output layer of the model; Calculating the sum of the output loss component and the hidden layer loss component to obtain the objective loss function.
7. The data generation method of the computing cluster according to claim 1, characterized in that, Before inputting the data to be processed into the data generation model, the method further includes: Identifying the current task scenario; When the current task scenario requires generating abnormal data, obtaining backup devices to be processed from the devices to be processed; Configuring the initial states of the backup devices to be processed as abnormal.
8. A data generation device for a computing cluster, characterized in that, The data generation device of the computing cluster includes: A construction unit, configured to obtain devices to be processed and construct a computing cluster according to the devices to be processed; A generation unit, configured to generate a topology structure of the computing cluster; An acquisition unit, configured to acquire the initial state of each device to be processed, each operation data executed on the computing cluster, and the final state of each device to be processed after executing each operation data from the computing cluster; The generation unit is further configured to generate training data according to the topology structure, the initial state of each device to be processed, each operation data, and the final state of each device to be processed; The construction unit is further configured to construct an initial model based on a variational autoencoder; A training unit, configured to train the initial model using the training data to obtain a data generation model; The construction unit is further configured to acquire data to be processed, input the data to be processed into the data generation model, and construct a computing cluster data group according to the output of the data generation model; Wherein, each operation data executed on the computing cluster refers to an operation on the computing cluster.
9. A computer device, characterized in that, The computer device includes: A memory, storing at least one instruction; and A processor, executing the instruction stored in the memory to implement the data generation method of the computing cluster according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: At least one instruction is stored in the computer-readable storage medium, and the at least one instruction is executed by a processor in the computer device to implement the data generation method of the computing cluster according to any one of claims 1 to 7.
Citation Information
Patent Citations
Cluster state monitoring method and device
CN112115031A