Machine fault prediction method and system based on large model, terminal and medium
Through the machine fault prediction method based on large models, the Time-LLM big model is trained using historical performance data and statistical features to predict the server failure risk, solving the problem of untimely fault handling in traditional operation and maintenance, and improving operation and maintenance efficiency and system stability.
Patent Information
- Application Number
- CN202510205323.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-30
AI Technical Summary
During the operation and maintenance of traditional server clusters, the workload is large and the task time is not fixed. In the event of a sudden failure of a key server, the failure of the operation and maintenance personnel to deal with it in a timely manner may affect business continuity.
The machine failure prediction method based on the big model is adopted, and the historical performance data of the failed machine is obtained, fault labels are set, statistical characteristics and lift and fall change trends are calculated, and the Time-LLM big model is trained to predict future failure risks.
It realizes the discovery of potential faults in advance, changes in traditional operation and maintenance methods, reduces the impact of faults, improves operation and maintenance efficiency and system stability, and reduces operation and maintenance complexity.
Smart Images

Figure CN120067760A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of server cluster operation and maintenance, and specifically relates to a machine fault prediction method, system, terminal and medium based on a large model. Background Art
[0002] A server cluster is a cluster system composed of multiple servers connected through a network, aiming to provide powerful processing capabilities, high availability and scalability by integrating computing, storage and network resources to meet large-scale business needs. Therefore, it is very important to perform daily management, maintenance and optimization on the server cluster to ensure its stable and efficient operation.
[0003] Currently, when operation and maintenance personnel perform operation and maintenance work on a server cluster, they often collect detailed fault information from various monitoring systems, server logs and alarm tools after a machine fails. According to the information collected, the operation and maintenance personnel perform fault detection on the faulty server, and conduct comprehensive analysis and root cause location based on various data collected during the fault detection process.
[0004] In the traditional server cluster operation and maintenance process, the workload is large, the task time is not fixed, and when some key servers suddenly fail, if the operation and maintenance personnel do not complete the fault handling in time, it will affect the continuity of the business. Summary of the Invention
[0005] In view of the above deficiencies of the prior art, the present invention provides a machine fault prediction method, system, terminal and medium based on a large model to solve the above technical problems.
[0006] In a first aspect, the present invention provides a machine fault prediction method based on a large model, including: S1, obtaining historical performance data of a faulty machine, where the historical performance data includes CPU performance data and RAM performance data; S2, setting a label indicating whether a fault occurs for the historical performance data at each moment, and calculating the statistical features and rising and falling change trends of the historical performance data; S3, training the Time-LLM large model based on the historical performance data, statistical features and rising and falling change trends after setting the label. The Time-LLM large model fuses the historical performance data, statistical features and rising and falling change trends after setting the label into a prompt text, and the model obtains predicted performance data based on the prompt text and outputs a fault prediction result within a preset future time period based on the predicted performance data; S4, obtaining the performance data of the physical machine to be judged currently, and inputting the performance data into the Time-LLM large model to obtain a fault prediction result within a preset future time period.
[0007] In an alternative embodiment, in step S2, when setting labels for historical performance data, it is determined whether there are missing values in the historical performance data based on the time intervals between the historical performance data. If there are missing values, the positions and quantities of the missing values are determined; For each missing value, based on the historical performance data points before and after the missing value, the cubic interpolation algorithm is used to calculate the estimated value of the missing value.
[0008] In an alternative embodiment, in step S3, the Time-LLM large model includes a prompt model, a large language model, and a post-classifier. The large language model includes a Transformer model, and the Transformer model has multiple encoders and decoders; The encoder includes an input layer, a position encoding layer, a residual connection layer, a normalization layer, and a feed-forward fully connected layer, and is used to extract the high-level features of the input feature vector; The decoding layer includes a multi-head attention layer, a feed-forward neural network layer, a residual link layer, a multi-head attention layer, and a normalization layer, and is used to generate a sequence output.
[0009] In an alternative embodiment, the prompt model includes: an input layer, a feature extraction and analysis module for calculating statistical features and rising and falling change trends, a prompt word generation module for fusing historical performance data, statistical features, and rising and falling change trends into a prompt text according to a predefined prompt word template, and a neural network conversion module for converting the prompt text into a numerical vector that is easy for the large language model to understand.
[0010] In an alternative embodiment, the post-classifier includes a binary classification fully connected layer, which receives the predicted performance data. The input performance data is linearly combined with a pre-learned weight matrix, and then a bias term is added to obtain the output result of the fully connected layer. Based on an activation function, the output result of the fully connected layer is converted into a probability value, and it is determined whether a fault will occur based on a preset classification threshold as the fault prediction result.
[0011] In an alternative embodiment, the specific training of the Time-LLM large model in step S3 includes: Constructing a data set based on the feature vectors, and dividing the data set into a training set, a test set, and a validation set; Training the model for multiple rounds based on the training set. The input training set is encoded, and the prompt model converts the input data into a vector form that can be processed by the large language model. The Transformer model trains the initial Time-LLM large model based on the encoded training set. After the encoder of the model processes the input content, the decoder generates the predicted performance data according to the extracted feature information, and the post-classifier analyzes and processes the predicted performance data to obtain the fault prediction result; Calculate the loss value between the performance data predicted by the computational model and the true performance data. When the loss value is greater than the threshold, calculate the gradient of the loss function with respect to the trainable parameters of the model based on the backpropagation algorithm, and update the parameters using the Adam optimizer based on the calculated gradient; After each round of training is completed, verify and adjust the model based on the validation set; After training is completed, test and adjust the model based on the test set.
[0012] In an alternative embodiment, the large language model is a pre-trained model with its weights frozen. When updating the parameters, only the parameters of the prompt model and the post-classifier are updated.
[0013] In a second aspect, the present invention provides a machine fault prediction system based on a large model. When the system is implemented, it executes the above-mentioned machine fault prediction method based on a large model. The system includes: A data acquisition module that acquires historical performance data of a faulty machine. The historical performance data includes CPU performance data and RAM performance data; A performance analysis module that sets a label indicating whether a fault has occurred for the historical performance data at each moment, and calculates the statistical features and rising and falling change trends of the historical performance data; A model training module that trains the Time-LLM large model based on the historical performance data, statistical features, and rising and falling change trends after setting the labels. The Time-LLM large model fuses the historical performance data, statistical features, and rising and falling change trends after setting the labels into a prompt text. The model obtains predicted performance data based on the prompt text and outputs a fault prediction result within a preset future time period based on the predicted performance data; A fault prediction module that acquires the performance data of the physical machine to be judged currently, and inputs the performance data into the Time-LLM large model to obtain a fault prediction result within a preset future time period.
[0014] In a third aspect, a terminal is provided, including: A processor and a memory, where The memory is used to store a computer program, The processor is used to call and run the computer program from the memory, so that the terminal executes the method of the above-mentioned terminal.
[0015] In a fourth aspect, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the methods described in the above aspects.
[0016] The beneficial effects of the present invention are as follows. The machine fault prediction method, system, terminal, and medium based on large models provided by the present invention obtain the historical performance data of the CPU and RAM of the faulty machine, set fault labels for it, calculate statistical features and rising and falling trends, train the Time-LLM large model with this to obtain prediction performance data, and then calculate the binary classification result, ultimately realizing fault prediction. This method can discover potential faults in advance, change the traditional operation and maintenance method, reduce the impact of faults, improve the operation and maintenance efficiency and system stability, and reduce the operation and maintenance complexity.
[0017] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a schematic flowchart of a machine fault prediction method based on a large model according to an embodiment of the present invention.
[0020] Figure 2 It is a schematic block diagram of a machine fault prediction system based on a large model according to an embodiment of the present invention.
[0021] Figure 3 It is a schematic structural diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0024] The following explains the key terms that appear in the present invention.
[0025] CPU utilization: The physical machine automatically collects the average CPU utilization within this period every hour.
[0026] RAM utilization: The physical machine automatically collects the average RAM utilization within this period every hour.
[0027] Fault alarm: When a physical machine fails, its built-in alarm module will automatically record the fault type and fault time. Many types of faults will be reflected in the changes of the machine's CPU and RAM utilization before they occur. Similarly, excessive load during machine operation may also be a cause of faults. For example, when the CPU and RAM of the machine are highly utilized for a long time, it may cause overheating and system crashes.
[0028] The machine fault prediction method based on a large model provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the machine fault prediction system based on a large model runs in the computer device.
[0029] Figure 1 is a schematic flowchart of the machine fault prediction method based on a large model according to an embodiment of the present invention. Among them, Figure 1 The execution subject can be a machine fault prediction system based on a large model. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.
[0030] As Figure 1 shown, the method includes: S1, obtaining the historical performance data of the faulty machine, where the historical performance data includes CPU performance data and RAM performance data; In the training stage, first query the machines with alarms above a certain level in the database, regard them as faulty machines, connect to the database according to the unique identification number of the faulty machine, and obtain the CPU and RAM performance data for more than 144 hours before the fault; in the prediction stage, directly obtain the CPU and RAM data for 128 hours before the current period of the physical machine to be detected.
[0031] To provide a data basis for subsequent model training and prediction, complete and accurate historical performance data can enable the model to learn the performance change rules of faulty machines, help improve the accuracy of fault prediction, and enable the model to better identify potential fault risks.
[0032] S2, setting labels for whether a fault occurs for the historical performance data at each moment, and calculating the statistical features and rising and falling change trends of the historical performance data; Based on the alarm records of the faulty machines, mark their historical performance data as having a fault (label set to 1), and mark the data of non-faulty machines as not having a fault (label set to 0). Calculate statistical features such as the mean, maximum, minimum, and standard deviation, and judge the rising and falling change trends by comparing the data at adjacent time points.
[0033] Provide sample data with fault identification for model training so that the model can learn the association between faults and performance data.
[0034] S3. Based on the historical performance data, statistical features, and rising and falling change trends after setting labels, train the Time-LLM large model. The Time-LLM large model integrates the historical performance data, statistical features, and rising and falling change trends after setting labels into a prompt text. The model obtains the predicted performance data based on the prompt text and outputs the fault prediction results within a preset future time period based on the predicted performance data. Organize the historical performance data, statistical features, and rising and falling change trends after setting labels into input samples in a specific format. Use these samples to fine-tune and train the Time-LLM large model. The model integrates this information to generate a prompt text, and then predicts the CPU and RAM performance data for a future period based on the prompt text through an internal mechanism. Finally, based on the preset fault judgment rules, output the fault prediction results based on the predicted performance data.
[0035] By training the large model by integrating multi-source information, it can fully learn the complex patterns and features related to faults, and improve the prediction ability of future performance data. The fault prediction results obtained based on the predicted performance data can early warn of potential faults, help operation and maintenance personnel take timely measures, and reduce the losses caused by faults.
[0036] S4. Obtain the performance data of the physical machine that needs to be judged currently, and input the performance data into the Time-LLM large model to obtain the fault prediction results within a preset future time period. Each of the above gives a simple implementation method and a beneficial effect in one paragraph.
[0037] Real-time collect the CPU and RAM performance data of the physical machine that needs to be judged currently, and perform operations such as normalization and missing value processing according to the preprocessing method during training. Input the processed data into the trained Time-LLM large model. The model analyzes and processes the input data based on the learned knowledge and patterns, and outputs the fault prediction results of this physical machine within a preset future time period. It can quickly evaluate and predict the running state of the current physical machine, and provide timely decision-making basis for operation and maintenance personnel.
[0038] Optionally, as an embodiment of the present invention, in step S1, due to the need for model training or a large amount of historical performance data and sample data of faulty machines to make a training set for fine-tuning the large model, in order to obtain historical performance data mixed with faulty machines and non-faulty machines, it is first necessary to define "machine failure". Generally speaking, when an alarm above a certain level occurs in this physical machine in the system, we consider that this physical machine "has failed". Therefore, when obtaining historical performance data, first query the corresponding alarmed machines in the database and regard these machines as faulty machines. Then, according to the unique identification numbers of these faulty machines, continue to connect to the database to find historical performance data. If we want the model to predict the future 16-hour machine performance trend based on the historical 128 decimal data (which can be adjusted according to the actual situation), then we need to take the data more than 144 hours before the machine fails.
[0039] Optionally, as an embodiment of the present invention, in step S2, when setting labels for historical performance data, since the data is collected at time intervals (such as every hour), theoretically the time series should be continuous. Based on the time interval between historical performance data and the normal collection time interval, determine whether there are missing values in the historical performance data. If there are missing values, determine the positions and quantities of the missing values; For each missing value, based on the historical performance data points before and after the missing value, use the cubic interpolation algorithm to calculate the estimated value of the missing value, so as to ensure the integrity and availability of the data and provide a reliable data basis for subsequent analysis and model training.
[0040] Optionally, as an embodiment of the present invention, in step S3, the Time-LLM large model includes a prompt model, a large language model, and a post-classifier. Generally, the large language model is good at processing conversations but is slightly lacking in predicting sequence data. Here, by inputting the data into the pre-trainable prompt model, extract information such as the maximum value, average value, and rising and falling changes of the sequence data, and convert this information into an abstract representation that is easier for the large language model to understand through prompt words and neural networks for subsequent processing by the large model.
[0041] The large language model uses the Llama2.0 - 7B version, which includes a Transformer model. The Transformer model has multiple encoders and decoders; The encoder includes an input layer, a position encoding layer, a residual connection layer, a normalization layer, and a feed-forward fully connected layer, which is used to extract the high-level features of the input feature vector; The input layer is the starting part of the encoder. Its main function is to receive the original input data and convert it into a format suitable for subsequent processing. Since the Transformer model itself does not have the ability to capture sequence order information, the position encoding layer is introduced to make up for this defect, using sine and cosine functions for position encoding. The main function of the residual connection is to solve the problems of gradient disappearance and gradient explosion in deep neural networks, and it also helps the model to learn and train better. The normalization layer is used to normalize the input data so that the mean of the data is 0 and the variance is 1. This helps to accelerate the convergence speed of the model and can improve the stability and generalization ability of the model. The feed-forward fully connected layer is used to perform further feature transformation and extraction on the feature vectors processed by the previous layers.
[0042] The decoding layer includes a multi-head attention layer, a feed-forward neural network layer, a residual connection layer, a multi-head attention layer and a normalization layer, which are used to generate sequence outputs.
[0043] The multi-head attention layer has two main functions. The first multi-head attention layer (self-attention layer) is used to perform attention calculation on the input sequence inside the decoder, so that the word vectors at each position can pay attention to other positions in the sequence, thereby capturing the dependencies inside the sequence. The second multi-head attention layer (encoder-decoder attention layer) is used to associate the output of the encoder with the input of the decoder, so that the decoder can use the feature information extracted by the encoder to generate the output. Similar to the feed-forward fully connected layer in the encoder, the feed-forward neural network layer in the decoder is used to perform further feature transformation and extraction on the feature vectors processed by the multi-head attention layer to generate higher-level feature representations for the final output. Just like in the decoder, the residual connection layer helps the model to learn and train better by directly adding the input to the output processed by some layers, and solves the problems of gradient disappearance and gradient explosion. The output after the residual connection is normalized so that the mean of the data is 0 and the variance is 1 to accelerate the convergence speed of the model and improve the stability of the model.
[0044] Optionally, as an embodiment of the present invention, the prompt model includes: an input layer, a feature extraction and analysis module for calculating statistical features and rising and falling change trends, a prompt word generation module for fusing historical performance data, statistical features and rising and falling change trends into prompt text according to a predefined prompt word template, and a neural network conversion module for converting the prompt text into a numerical vector that is easy for the large language model to understand.
[0045] Optionally, as an embodiment of the present invention, the post-classifier includes a binary classification fully connected layer that receives the predicted performance data. The input performance data is linearly combined with a pre-learned weight matrix and then added with a bias term to obtain the output result of the fully connected layer. Based on an activation function, the output result of the fully connected layer is converted into a probability value, and whether a failure will occur is judged based on a preset classification threshold as the failure prediction result.
[0046] Optionally, as an embodiment of the present invention, the specific training of the Time-LLM large model in step S3 includes: Construct a data set based on the feature vectors and divide the data set into a training set, a test set, and a validation set; Train the model for multiple rounds based on the training set. Encode the input training set, and the prompt model converts the input data into a vector form that can be processed by the large language model. The Transformer model trains the initial Time-LLM large model based on the encoded training set. After the encoder of the model processes the input content, the decoder generates the predicted performance data according to the extracted feature information, and the post-classifier analyzes and processes the predicted performance data to obtain the failure prediction result; Calculate the loss value between the performance data predicted by the model and the real performance data. When the loss value is greater than the threshold, calculate the gradient of the loss function with respect to the trainable parameters of the model based on the backpropagation algorithm, and update the parameters using the Adam optimizer based on the calculated gradient; After each round of training is completed, verify and adjust the model based on the validation set; After training is completed, test and adjust the model based on the test set.
[0047] Optionally, as an embodiment of the present invention, the large language model is a pre-trained model with its weights frozen. When updating the parameters, only the parameters of the prompt model and the post-classifier are updated.
[0048] In some embodiments, the large model-based machine fault prediction system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the large model-based machine fault prediction system can be stored in the memory of the computer device and executed by at least one processor to execute (see details Figure 1 description) the functions of large model-based machine fault prediction.
[0049] In this embodiment, the large model-based machine fault prediction system can be divided into multiple functional modules according to the functions it performs, such as Figure 2As shown in the figure. The functional modules of the system may include: a data acquisition module, a performance analysis module, a model training module, and a fault prediction module. The modules referred to in the present invention refer to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments. The system includes: A data acquisition module that acquires historical performance data of a faulty machine, and the historical performance data includes CPU performance data and RAM performance data; A performance analysis module that sets a label indicating whether a fault has occurred for the historical performance data at each moment, and calculates the statistical characteristics and rising and falling change trends of the historical performance data; A model training module that trains the Time-LLM large model based on the historical performance data, statistical characteristics, and rising and falling change trends after setting the labels. The Time-LLM large model fuses the historical performance data, statistical characteristics, and rising and falling change trends after setting the labels into a prompt text. The model obtains predicted performance data based on the prompt text and outputs a fault prediction result within a preset future time period based on the predicted performance data; A fault prediction module that acquires the performance data of the physical machine to be judged currently, and inputs the performance data into the Time-LLM large model to obtain a fault prediction result within a preset future time period.
[0050] The historical CPU and RAM performance data of the faulty machine is collected through the data acquisition module. The performance analysis module marks the fault label for it, calculates the characteristics and trends. The model training module trains the Time-LLM large model based on this information, enabling it to fuse the information to generate a prompt text and obtain predicted performance data, and then output a fault prediction result. The fault prediction module can also analyze the current physical machine performance data using the model, realizing the effective prediction of physical machine faults, which can help operation and maintenance personnel discover potential fault hazards in advance, take preventive measures, reduce the fault occurrence rate and losses, and improve system stability and operation and maintenance efficiency.
[0051] Figure 3 It is a schematic structural diagram of a terminal provided by an embodiment of the present invention, and this terminal can be used to execute the method for predicting machine faults based on a large model provided by an embodiment of the present invention.
[0052] Among them, the terminal may include: a processor, a memory, and a communication unit. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0053] Among them, the memory can be used to store the execution instructions of the processor. The memory can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. When the execution instructions in the memory are executed by the processor, the terminal can execute some or all of the steps in the above method embodiments.
[0054] The processor is the control center of the storage terminal, connecting various parts of the entire electronic terminal through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory, and calling the data stored in the memory, it executes various functions of the electronic terminal and / or processes data. The processor can be composed of an integrated circuit (IC). For example, it can be composed of a single packaged IC, or can be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor can include only a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single operation core or can include multiple operation cores.
[0055] The communication unit is used to establish a communication channel so that the storage terminal can communicate with other terminals. It receives user data sent by other terminals or sends user data to other terminals.
[0056] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the embodiments provided by the present invention. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.
[0057] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disc, etc., which are various media that can store program codes, including several instructions for causing a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0058] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the terminal embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the descriptions in the method embodiments.
[0059] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the systems or modules can be in an electrical, mechanical or other forms.
[0060] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place, or can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0061] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0062] Although the present invention has been described in detail by referring to the accompanying drawings and in conjunction with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and all such modifications or substitutions should be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of changes or substitutions, which should all be covered by the protection scope of the present invention.
Claims
1. A machine failure prediction method based on a large model, characterized in that: The following steps are involved: S1, obtaining historical performance data of the faulty machine, the historical performance data including CPU performance data and RAM performance data; S2: Set a label for whether a fault has occurred for each moment of historical performance data, and calculate the statistical characteristics and rising and falling trend of the historical performance data; S3: Train the Time-LLM large model based on the historical performance data, statistical features, and rising and falling change trends after labeling. The Time-LLM large model integrates the historical performance data, statistical features, and rising and falling change trends after labeling into prompt text. The model obtains predicted performance data based on the prompt text, and outputs fault prediction results within a preset future time period based on the predicted performance data. S4, obtain the performance data of the physical machine that needs to be judged currently, input the performance data into the Time-LLM large model, and obtain the fault prediction result within the preset future time period.
2. The machine failure prediction method based on a large model according to claim 1, characterized in that: In step S2, when setting labels for historical performance data, it is determined whether there are missing values in the historical performance data based on the time interval between the historical performance data, and if there are missing values, the location and number of the missing values are determined; For each missing value, the cubic interpolation algorithm is used to calculate the estimated value of the missing value based on the historical performance data points before and after the missing value.
3. The machine failure prediction method based on a large model according to claim 1, characterized in that: In step S3, the Time-LLM large model includes a prompt model, a large language model, and a post-classifier, the large language model includes a Transformer model, and the Transformer model has multiple encoders and decoders; The encoder includes an input layer, a position encoding layer, a residual connection layer, a normalization layer, and a feed-forward fully connected layer, which are used to extract high-order features of the input feature vector; The decoding layer includes a multi-head attention layer, a feedforward neural network layer, a residual link layer, a multi-head attention layer and a normalization layer to generate sequence output.
4. The machine failure prediction method based on a large model according to claim 3 is characterized in that: The prompt model includes: an input layer, a feature extraction and analysis module for calculating statistical features and rising and falling change trends, a prompt word generation module for fusing historical performance data, statistical features and rising and falling change trends into prompt text according to a pre-defined prompt word template, and a neural network conversion module for converting the prompt text into a numerical vector that is easy for the large language model to understand.
5. The machine failure prediction method based on a large model according to claim 4 is characterized in that: The post-classifier includes a binary fully connected layer, which receives the predicted performance data. The input performance data is linearly combined with the pre-learned weight matrix, and the bias term is added to obtain the output result of the fully connected layer. The output result of the fully connected layer is converted into a probability value based on the activation function, and whether a fault will occur is judged based on the preset classification threshold as the fault prediction result.
6. The machine failure prediction method based on a large model according to claim 5 is characterized in that: The specific training of the Time-LLM large model in step S3 includes: Construct a data set based on the feature vector and divide the data set into training set, test set and validation set; The model is trained for multiple rounds based on the training set. The input training set is encoded. The prompt model converts the input data into a vector form that can be processed by the large language model. The Transformer model trains the initial Time-LLM large model based on the encoded training set. After the encoder of the model processes the input content, the decoder generates predicted performance data based on the extracted feature information. The post-classifier analyzes and processes the predicted performance data to obtain the fault prediction result. Calculate the loss value between the performance data predicted by the model and the actual performance data. When the loss value is greater than the threshold, calculate the gradient of the loss function to the model's trainable parameters based on the back-propagation algorithm, and use the Adam optimizer to update the parameters based on the calculated gradient. After each round of training is completed, the model is verified and adjusted based on the validation set; After training is completed, the model is tested and adjusted based on the test set.
7. The machine fault prediction method based on a large model according to claim 6, characterized in that: The large language model is a pre-trained model, and its weights are frozen. When updating parameters, only the parameters of the prompt model and the post-classifier are updated.
8. A machine failure prediction system based on a large model, characterized in that: When the system is implemented, the machine fault prediction method based on the large model as described in any one of claims 1 to 7 is executed, and the system includes: A data acquisition module is used to acquire historical performance data of the faulty machine, including CPU performance data and RAM performance data; The performance analysis module sets a label for the historical performance data at each moment to indicate whether a fault has occurred, and calculates the statistical characteristics and rising and falling trend of the historical performance data; The model training module trains the Time-LLM large model based on the historical performance data, statistical features, and rising and falling change trends after labeling. The Time-LLM large model integrates the historical performance data, statistical features, and rising and falling change trends after labeling into prompt text. The model obtains the predicted performance data based on the prompt text, and outputs the fault prediction results within a preset future time period based on the predicted performance data. The fault prediction module obtains the performance data of the physical machine that needs to be judged, inputs the performance data into the Time-LLM large model, and obtains the fault prediction results within a preset future time period.
9. A terminal, characterized in that: include: A memory for storing a machine failure prediction program based on a large model; A processor is used to implement the steps of the large model-based machine fault prediction method as described in any one of claims 1 to 7 when executing the large model-based machine fault prediction program.
10. A computer-readable storage medium, characterized in that: The readable storage medium stores a large model-based machine fault prediction program, which, when executed by a processor, implements the steps of the large model-based machine fault prediction method as described in any one of claims 1 to 7.
Citation Information
Cited By
Comprehensive pipe gallery operation and maintenance system and method based on AI large model
CN120746554A