Equipment health state prediction question-answering system and method for pre-training large language model
Through pre-training large language models and template-prompted question-answering systems, the generalization and data missing problems in equipment health status prediction are solved, efficient and accurate equipment health status prediction and instant warning are achieved, and the stability and reliability of the production line are improved.
Patent Information
- Application Number
- CN202510935160.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-21
AI Technical Summary
Existing deep learning equipment health status prediction methods have limitations in generalization, data missing and reasoning ability, especially in the case of very few samples or zero samples, the learning ability is poor, and traditional models are difficult to fully consider the complexity and diversity of equipment operating status, resulting in inaccurate prediction results.
By adopting a pre-trained large language model, combined with reversible instance normalization, embedding and tokenization, prompt learning-based reprogramming and template prompt question-answering system, the prediction of equipment health status and question-answering communication are realized through collection, labeling, editing and interaction modules.
It improves the accuracy and user-friendliness of equipment health status prediction, can process structured and unstructured data, provide immediate warning information, reduce maintenance costs, and improve the stability and reliability of production lines.
Smart Images

Figure CN120821804A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and artificial intelligence technology, and in particular to a system and method for predicting equipment health status using a pre-trained large language model. Background Art
[0002] Equipment health prediction is a crucial component of industrial production and equipment management. It involves real-time monitoring of equipment operation, data analysis, and fault prediction. Predicting equipment health can help companies promptly identify potential faults and issues, implement preventative maintenance measures, and reduce production interruptions and repair costs. This ensures efficient equipment operation and smooth production processes, ultimately improving production efficiency and reliability.
[0003] In recent years, with the rapid development of deep learning and big data technologies, the application of neural network models in equipment health prediction has attracted considerable attention. However, existing deep learning equipment health prediction methods are typically strictly specialized within a specific domain, making them difficult to generalize across different scenarios and devices. Furthermore, these methods typically require large amounts of data to train the models, resulting in poor learning performance with very few or no samples. Furthermore, equipment health prediction requires models with good reasoning capabilities. Due to the complexity and diversity of equipment operating states, the reasoning process is often complex. Traditional prediction models may not fully account for the interactions and influences between various factors, resulting in insufficient reasoning capabilities and inaccurate prediction results.
[0004] Large language models demonstrate powerful pattern recognition and reasoning capabilities on complex labeled sequences, but their potential for predicting equipment health status remains untapped. Furthermore, the predictive paradigm of question-and-answer dialogue is relatively accessible and user-friendly for non-research users. Therefore, to address the limitations of current equipment health status prediction methods in terms of generalization, data loss, and reasoning capabilities, and to further explore the predictive capabilities of large language models for equipment time-series operational data, this paper proposes an equipment health status prediction question-and-answer system using a pre-trained large language model. Summary of the Invention
[0005] In response to a series of problems in current equipment health status prediction methods in terms of generalization, data missing, and reasoning ability, as well as the fact that the development of pre-trained large language models in the field of equipment time series data is still limited by data sparsity, the present invention proposes an equipment health status prediction question-answering system based on a pre-trained large language model.
[0006] To address the issue of data distribution drift in equipment time series forecasting, a reversible instance normalization method is introduced to explicitly restore non-stationary information after the model output. This allows the model to ignore data drift during learning while avoiding the loss of non-stationary information. To better preserve local semantic information and reduce computational burden, the input sequence is embedded and tokenized. To align the time series input features with the natural language text domain, the embedding sub-block is reprogrammed into the source data representation space to obtain word representations of the equipment data. To better guide the transformation of the reprogrammed time series sub-block, prompts are constructed as prefixes based on prompt learning to enrich the input time series with additional context and provide task instructions in natural language. Furthermore, to enable non-research users to complete the forecasting task in a relatively easy and user-friendly question-answering format, template-based prompt descriptions are used to achieve efficient data-to-text conversion. The model predicts time series in a sentence-by-sentence manner. The input prompts provide historical information and future input queries, and the output prompts answer the true answer to the question.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] An equipment health status prediction question-answering system based on a pre-trained large language model, characterized by comprising: an acquisition module, a tagging module, an editing module, a prediction module, and an interaction module;
[0009] The acquisition module is used to collect time series data of industrial production equipment and perform preprocessing to obtain processed data;
[0010] The marking module is used to divide and mark the processed data to obtain a marking sequence;
[0011] The editing module is used to reprogram the tag sequence memory to obtain word representation of time series data;
[0012] The prediction module is used to collect the health status of industrial production equipment based on the word representation and perform prediction to obtain a prediction result;
[0013] The interactive module is used to interact with the user and complete question-and-answer communication on the prediction results.
[0014] Preferably, the acquisition module includes: a sensor and a processing unit;
[0015] The sensors are placed at key locations of industrial production equipment to acquire time series data;
[0016] The processing unit is used to preprocess the time series data to obtain processed data.
[0017] Preferably, the workflow of the labeling module includes: dividing the processed data into a plurality of L-lengthp The total number of input sub-blocks is Where S is the horizontal sliding step size; T is the sequence length; given a sub-block in, Represents the real number field; a simple linear layer is used as a sub-block embedder to create dimension d m , embed it as Enter characteristics for rig timing.
[0018] Preferably, the editing module comprises: a reprogramming unit and a construction unit;
[0019] The reprogramming unit is used to reprogram the edit sequence to obtain a word representation of the time series data;
[0020] The construction unit is used to guide the reprogramming unit to convert the time series sub-blocks.
[0021] Preferably, the workflow of the reprogramming unit includes: reprogramming the time series using the source data pattern; the process includes:
[0022] Reprogram the backbone using pre-trained word vectors;
[0023] Perform linear testing on pre-trained word vectors and maintain a small set of text prototypes;
[0024] Let the text prototype learn to connect language clues to represent the local sub-block information of time series data;
[0025] The text description corresponding to the sub-block is obtained through the multi-head self-attention mechanism to complete the reprogramming.
[0026] Preferably, the workflow of the construction unit includes: constructing prompts based on prompt learning as prefixes to enrich the input context content, which is used to guide the reprogramming unit to perform the conversion of time series sub-blocks; the constructed prompts include: dataset context, task instructions and statistical descriptions.
[0027] The present invention also provides a method for predicting equipment health status using a pre-trained large language model. The method is applied to the above-mentioned system and comprises the following steps:
[0028] Collect time series data of industrial production equipment and preprocess it to obtain processed data;
[0029] dividing and marking the processed data to obtain a marked sequence;
[0030] reprogramming the tag sequence memory to obtain word representation of time series data;
[0031] Based on the word representation, the health status of industrial production equipment is collected and predicted to obtain a prediction result.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] This invention can process large amounts of structured and unstructured data, enabling more accurate identification of key indicators and patterns of equipment health, predicting equipment health and providing immediate early warning information. This helps companies take necessary maintenance measures before failures occur, avoiding production interruptions and losses. This not only improves the stability and reliability of production lines but also reduces maintenance costs for companies, contributing to the development of intelligent manufacturing and industrial sectors. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0035] Figure 1 Schematic diagram of the system structure of an embodiment of the present invention;
[0036] Figure 2 Schematic diagram of the reprogramming process according to an embodiment of the present invention;
[0037] Figure 3 This is a schematic diagram of the structure of a pre-trained large language model according to an embodiment of the present invention;
[0038] Figure 4 Schematic diagram of the overall process of equipment health status prediction according to an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0040] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] Example 1
[0042] This embodiment provides a pre-trained large language model for equipment health status prediction question answering system, such as Figure 1As shown, it includes: an acquisition module, a labeling module, an editing module, a prediction module and an interaction module; the acquisition module is used to collect the time series data of industrial production equipment and perform preprocessing to obtain processed data; the labeling module is used to divide and label the processed data to obtain a label sequence; the editing module is used to reprogram the label sequence memory to obtain a word representation of the time series data; the prediction module is used to collect the health status of industrial production equipment based on the word representation and predict the prediction result; the interaction module is used to interact with the user and complete the question-and-answer communication of the prediction result.
[0043] The following will describe in detail how the present invention solves technical problems in practical work in conjunction with this embodiment.
[0044] The sensor data received by the processing unit suffers from distribution drift in time series data, which can easily lead to inconsistent distributions between the training and test sets. Therefore, reversible instance normalization (RIN) is used to normalize each input channel individually to have zero mean and unit standard deviation to mitigate time series distribution drift. RIN consists of two symmetrical parts: normalization and denormalization. Before inputting the data into the model, the data is normalized. After the model learns, the model output is denormalized.
[0045] First, the original data is normalized, and for each input x (i) Perform instance normalization, that is, use its own mean and variance for normalization. The specific calculation process is as follows: The input of the time series prediction task is expressed as: The corresponding output is expressed as: Where N represents the number of sequences. Define K as the number of variables, T x is the input sequence length, T y is the output sequence length. Given input Prediction output: ;The normalization formula is as follows:
[0046]
[0047] Among them, E t represents the mean; Var represents the standard deviation; express Normalized data; represents the value of the kth variable at time step t in the i-th sample; j represents the index of the time step; ∈ represents a small positive number used to avoid the denominator being zero, usually set to 10 -6 or smaller; k represents the index of the variable in the time series; γ and β represent the learned affine transformation functions.
[0048] The denormalization method is as follows, using the same parameter values as the normalization stage:
[0049]
[0050] in, represents the predicted value of the kth variable in the i-th sample after denormalization at time step t; Represents the model prediction output.
[0051] Afterwards, the labeling module divides the processed data into several segments of length L. p Continuous overlapping or non-overlapping sub-blocks, so the total number of input sub-blocks is Where S is the horizontal sliding step size; T is the sequence length. in, Represents the real number field; a simple linear layer is used as a sub-block embedder to create dimension d m , embed it as That is the equipment timing input feature.
[0052] This operation better preserves local semantic information by aggregating local information into each sub-block and tokenizes the input to form a compact input token sequence, reducing the computational burden.
[0053] The labeled sequence is then reprogrammed using the reprogramming unit in the editing module. To align the time series and natural language patterns and activate the backbone's time series understanding and reasoning capabilities, the present invention reprograms the time series using the source data schema. Because time series data cannot be directly edited or described losslessly in natural language, traditional reprogramming methods struggle to meet these requirements.
[0054] To solve this problem, this embodiment proposes to use pre-trained word vectors in the backbone Reprogramming is performed, where V is the vocabulary size and D is the hidden dimension of the backbone model. Since there is no prior knowledge indicating which source tokens are directly related, improper use of E will lead to a huge and potentially dense reprogramming space. This embodiment chooses to perform linear detection on E to solve this problem, thereby maintaining a smaller set of text prototypes, denoted as Where V′<<V. Let the text prototype learn to connect language clues, for example, red lines represent short rises and blue lines represent slow declines, and combine them to represent local sub-block information of time series data without leaving the space of language model pre-training.
[0055] The text description corresponding to the sub-block is adaptively obtained through the multi-head self-attention mechanism, allowing adaptive selection of relevant source information. Specifically, for each detection head k = {1, ..., K}, the query matrix is defined Bond Matrix Value Matrix in D represents the hidden dimension of the backbone model, The time series sub-block in each attention head is reprogrammed as follows:
[0056]
[0057] By aggregating each get A linear projection is then performed to align the hidden dimensions with the backbone model, resulting in As word representation of temporal data. The reprogramming process is as follows Figure 2 shown.
[0058] In order to better guide the transformation of the reprogrammed time series sub-blocks, the construction unit of this embodiment constructs prompts as prefixes based on prompt learning to enrich the input context content. The constructed prompts have three key components: dataset context, task instructions, and statistical descriptions. The dataset context provides the large language model with basic operating information about the equipment input time series, which usually exhibit different characteristics in various fields. Task instructions are important guides for the large language model in the patch embedding transformation of a specific task. In addition, this embodiment uses key statistical data (such as trends and lags) to enrich the input time series to facilitate pattern recognition and reasoning.
[0059] Finally, the prediction module is used to obtain the prediction results. The prediction module of this embodiment builds a prediction model based on the large language model. In order to maintain the data-independent representation learning ability of the model, most of its parameters are frozen. After the prompts and embedding blocks are packaged and forwarded through the pre-trained large language model, the prefix part is discarded and the output representation is obtained. They are flattened and linearly projected to obtain the final prediction results. The structure of the pre-trained large language model is as follows: Figure 3 shown.
[0060] In this embodiment, the pre-trained large language model structure specifically includes:
[0061] This paper describes an architecture based on a pre-trained large language model. Its core structure uses the Transformer model, which has powerful feature extraction and generation capabilities. Its modular design enables adaptability and scalability to complex tasks. The model consists of an input embedding layer, a stacked encoder layer, a decoder layer, and an output layer. The functions of each module are as follows:
[0062] The input embedding layer is primarily responsible for mapping discrete input data into a high-dimensional continuous vector space. It extracts features from the input data using the embedding matrix and, combined with position embedding, encodes the temporal and sequential information in the sequence. This overcomes the Transformer architecture's limitations in modeling sequential order, enabling the model to handle complex time series data or other ordered data.
[0063] The encoder layer consists of multiple stacked layers, each of which incorporates a self-attention mechanism, a feedforward neural network, residual connections, and layer normalization. The self-attention mechanism captures global context by calculating the correlation between elements in the input sequence, enhancing the model's ability to process long sequences. The feedforward neural network further enhances feature representation through nonlinear transformations, while the introduction of residual connections ensures the stability of gradient flow, preventing the vanishing gradient problem in deep networks and improving training efficiency. Layer normalization further optimizes the training process, enhancing the stability and robustness of the model.
[0064] The decoder layer adds an encoder-decoder attention mechanism to the encoder layer, ensuring that the model effectively incorporates contextual information from the input sequence for decoding. Furthermore, the decoder's self-attention mechanism uses masking to restrict its access to only the preceding information required to generate the current output, thereby maintaining causal relationships and ensuring the rationality and consistency of the generated content. The decoder layer collaborates with the encoder layer to achieve a precise mapping from input data to output targets.
[0065] The output layer, consisting of a fully connected network and a softmax function, maps the features generated by the decoder into the target space and generates a probability distribution for classification or generation tasks. In specific applications of this invention, the output can be a predicted equipment health status label or a natural language description.
[0066] By optimizing the combination of these modules, the pre-trained large language model of this invention not only learns rich linguistic features from a large-scale corpus but also, through task fine-tuning, further adapts it to the scenario of equipment health status prediction. This model has significant advantages in feature extraction, global context capture, and complex causal relationship modeling. It can demonstrate excellent predictive performance even in situations with limited data or complex tasks, providing efficient and reliable technical support for equipment health status prediction.
[0067] Finally, users use the interactive module to interact with the system and complete question-and-answer communication about the prediction results.
[0068] This embodiment chooses to use descriptions based on template prompts to achieve efficient conversion of data to text. The model predicts time series in a sentence-by-sentence manner to complete question-and-answer communication with users.
[0069] The template consists of two main parts: input prompts and output prompts. The input prompts include the description of the historical observations and the indicators of the predicted target time step, which can be divided into the context part and the question part. The context provides historical information for the prediction, and the question part can be regarded as an input query about the future. The output prompts process the required predicted value, which is used as the true value label for training or evaluation and is the real answer to the question. Based on the equipment time series health data prediction task setting, the input is at t obs History of numerical data points collected at consecutive time steps: in Represents the target U observed at time step t m The predicted target (output) is the value of n time steps t in the future. obs+1 , t obs+2 ,...,t obs+n Data value
[0070] Table 1
[0071]
[0072] The template-based value-statement conversion is shown in Table 1. The overall process of system prediction is as follows Figure 4 As stated.
[0073] Example 2
[0074] This embodiment also provides a method for predicting equipment health status using a pre-trained large language model. The method includes the following steps:
[0075] S1. Collect time series data of industrial production equipment and preprocess it to obtain processed data.
[0076] Data is received from sensors deployed on industrial production equipment. Time series data suffers from distribution drift, which can easily lead to inconsistent distributions between the training and test sets. Therefore, reversible instance normalization (RIN) is used to normalize each input channel individually to have zero mean and unit standard deviation to mitigate time series distribution drift. RIN consists of two symmetrical parts: normalization and denormalization. Before data is input into the model, it is normalized. After the model learns, the model output is denormalized.
[0077] First, the original data is normalized, and for each input x (i) Perform instance normalization, that is, use its own mean and variance for normalization. The specific calculation process is as follows: The input of the time series prediction task is expressed as: The corresponding output is expressed as: Where N represents the number of sequences. Define K as the number of variables, T xis the input sequence length, T y is the output sequence length. Given input , predicted output: ;The normalization formula is as follows:
[0078]
[0079] Among them, E t represents the mean; Var represents the standard deviation; express Normalized data; represents the value of the kth variable at time step t in the i-th sample; j represents the index of the time step; ∈ represents a small positive number used to avoid the denominator being zero, usually set to 10 -6 or smaller; k represents the index of the variable in the time series; γ and β represent the learned affine transformation functions.
[0080] The denormalization method is as follows, using the same parameter values as the normalization stage:
[0081]
[0082] in, represents the predicted value of the kth variable in the i-th sample after denormalization at time step t; Represents the model prediction output.
[0083] S2. Divide and label the processed data to obtain a label sequence.
[0084] After processing, the input sequence is divided into several blocks of length L. p Continuous overlapping or non-overlapping sub-blocks, so the total number of input sub-blocks is Where S is the horizontal sliding step size; T is the sequence length. in, Represents the real number field; a simple linear layer is used as a sub-block embedder to create dimension d m , embed it as That is the equipment timing input feature.
[0085] This operation better preserves local semantic information by aggregating local information into each sub-block and tokenizes the input to form a compact input token sequence, reducing the computational burden.
[0086] S3. Reprogram the tag sequence memory to obtain word representation of temporal data.
[0087] To align the time series and natural language patterns and activate the backbone's time series understanding and reasoning capabilities, this paper reprograms the time series using the source data patterns. Because time series data cannot be directly edited or described losslessly in natural language, traditional reprogramming methods are difficult to meet these requirements.
[0088] To solve this problem, this embodiment proposes to use pre-trained word vectors in the backbone Reprogramming is performed, where V is the vocabulary size and D is the hidden dimension of the backbone model. Since there is no prior knowledge indicating which source tokens are directly related, improper use of E will lead to a huge and potentially dense reprogramming space. This embodiment chooses to perform linear detection on E to solve this problem, thereby maintaining a smaller set of text prototypes, denoted as Where V′<<V. Let the text prototype learn to connect language clues, for example, red lines represent short rises and blue lines represent slow declines, and combine them to represent local sub-block information of time series data without leaving the space of language model pre-training.
[0089] The text description corresponding to the sub-block is adaptively obtained through the multi-head self-attention mechanism, allowing adaptive selection of relevant source information. Specifically, for each detection head k = {1, ..., K}, the query matrix is defined Bond Matrix Value Matrix in D represents the hidden dimension of the backbone model, The time series sub-block in each attention head is reprogrammed as follows:
[0090]
[0091] By aggregating each get A linear projection is then performed to align the hidden dimensions with the backbone model, resulting in As word representation of temporal data. The reprogramming process is as follows Figure 2 shown.
[0092] In order to better guide the conversion of the reprogrammed time series sub-blocks, this embodiment constructs prompts as prefixes based on prompt learning to enrich the input context content. The constructed prompts have three key components: dataset context, task instructions, and statistical descriptions. The dataset context provides the large language model with basic operational information about the equipment input time series, which usually exhibit different characteristics in various fields. Task instructions are important guides for the large language model in the conversion of patch embeddings for specific tasks. In addition, this embodiment uses key statistical data (such as trends and lags) to enrich the input time series to facilitate pattern recognition and reasoning.
[0093] S4. Based on word representation, collect the health status of industrial production equipment for prediction and obtain the prediction results.
[0094] A prediction model is built based on the large language model. In order to maintain the model's data-independent representation learning ability, most of its parameters are frozen. After the prompts and embedding blocks are packaged and forwarded through the pre-trained large language model, the prefix part is discarded and the output representation is obtained. They are flattened and linearly projected to obtain the final prediction result. The pre-trained large language model structure is as follows: Figure 3 shown.
[0095] In this embodiment, the pre-trained large language model structure specifically includes:
[0096] This paper describes an architecture based on a pre-trained large language model. Its core structure uses the Transformer model, which has powerful feature extraction and generation capabilities. Its modular design enables adaptability and scalability to complex tasks. The model consists of an input embedding layer, a stacked encoder layer, a decoder layer, and an output layer. The functions of each module are as follows:
[0097] The input embedding layer is primarily responsible for mapping discrete input data into a high-dimensional continuous vector space. It extracts features from the input data using the embedding matrix and, combined with position embedding, encodes the temporal and sequential information in the sequence. This overcomes the Transformer architecture's limitations in modeling sequential order, enabling the model to handle complex time series data or other ordered data.
[0098] The encoder layer consists of multiple stacked layers, each of which incorporates a multi-head attention mechanism, a feedforward neural network, residual connections, and layer normalization. The multi-head attention mechanism captures global context by calculating the correlation between elements in the input sequence, enhancing the model's ability to process long sequences. The feedforward neural network further enhances feature representation through nonlinear transformations, while the introduction of residual connections ensures the stability of gradient flow, preventing the vanishing gradient problem in deep networks and improving training efficiency. Layer normalization further optimizes the training process, enhancing the stability and robustness of the model.
[0099] The decoder layer adds an encoder-decoder attention mechanism to the encoder layer, ensuring that the model effectively incorporates contextual information from the input sequence for decoding. Furthermore, the decoder's self-attention mechanism uses masking to restrict its access to only the preceding information required to generate the current output, thereby maintaining causal relationships and ensuring the rationality and consistency of the generated content. The decoder layer collaborates with the encoder layer to achieve a precise mapping from input data to output targets.
[0100] The output layer, consisting of a fully connected network and a softmax function, maps the features generated by the decoder into the target space and generates a probability distribution for classification or generation tasks. In specific applications of this invention, the output can be a predicted equipment health status label or a natural language description.
[0101] By optimizing the combination of these modules, the pre-trained large language model of this invention not only learns rich linguistic features from a large-scale corpus but also, through task fine-tuning, further adapts it to the scenario of equipment health status prediction. This model has significant advantages in feature extraction, global context capture, and complex causal relationship modeling. It can demonstrate excellent predictive performance even in situations with limited data or complex tasks, providing efficient and reliable technical support for equipment health status prediction.
[0102] Finally, this embodiment also provides a description based on template prompts to achieve efficient conversion of data into text, and uses the model to predict time series in a sentence-by-sentence manner to complete question-and-answer communication with users.
[0103] The template consists of two main parts: input prompts and output prompts. The input prompts include the description of the historical observations and the indicators of the predicted target time step, which can be divided into the context part and the question part. The context provides historical information for the prediction, and the question part can be regarded as an input query about the future. The output prompts process the required predicted value, which is used as the true value label for training or evaluation and is the real answer to the question. Based on the equipment time series health data prediction task setting, the input is at t obs History of numerical data points collected at consecutive time steps: in Represents the target U observed at time step t m The predicted target (output) is the value of n time steps t in the future. obs+1 , t obs+2 ,...,t obs+n Data value
[0104] The template-based value-statement conversion is shown in Table 1. The overall process of system prediction is as follows Figure 4 As stated.
[0105] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A pre-trained large language model for equipment health status prediction question-answering system, characterized by: include: Acquisition module, labeling module, editing module, prediction module and interaction module; The acquisition module is used to collect time series data of industrial production equipment and perform preprocessing to obtain processed data; The marking module is used to divide and mark the processed data to obtain a marking sequence; The editing module is used to reprogram the tag sequence memory to obtain word representation of time series data; The prediction module is used to collect the health status of industrial production equipment based on the word representation and perform prediction to obtain a prediction result; The interactive module is used to interact with the user and complete question-and-answer communication on the prediction results.
2. The equipment health status prediction question-answering system based on a pre-trained large language model according to claim 1 is characterized in that: The acquisition module includes: a sensor and a processing unit; The sensors are placed at key locations of industrial production equipment to acquire time series data; The processing unit is used to preprocess the time series data to obtain processed data.
3. The equipment health status prediction question-answering system based on a pre-trained large language model according to claim 1 is characterized in that: The workflow of the labeling module includes: dividing the processed data into a number of L-length p The total number of input sub-blocks is Where S is the horizontal sliding step size; T is the sequence length; given a sub-block in, Represents the real number field; a simple linear layer is used as a sub-block embedder to create dimension d m , embed it as Enter characteristics for rig timing.
4. The equipment health status prediction question-answering system based on a pre-trained large language model according to claim 1 is characterized in that: The editing module includes: a reprogramming unit and a construction unit; The reprogramming unit is used to reprogram the edit sequence to obtain a word representation of the time series data; The construction unit is used to guide the reprogramming unit to convert the time series sub-blocks.
5. The equipment health status prediction question-answering system based on a pre-trained large language model according to claim 4 is characterized in that: The workflow of the reprogramming unit includes: reprogramming the time series using the source data pattern; the process includes: Reprogram the backbone using pre-trained word vectors; Perform linear testing on pre-trained word vectors and maintain a small set of text prototypes; Let the text prototype learn to connect language clues to represent the local sub-block information of time series data; The text description corresponding to the sub-block is obtained through the multi-head self-attention mechanism to complete the reprogramming.
6. The equipment health status prediction question-answering system based on a pre-trained large language model according to claim 5 is characterized in that: The workflow of the construction unit includes: constructing prompts based on prompt learning as prefixes to enrich the input context content, which is used to guide the reprogramming unit to perform the conversion of time series sub-blocks; the constructed prompts include: dataset context, task instructions and statistical descriptions.
7. A method for predicting equipment health status using a pre-trained large language model, the method being applied to the system according to any one of claims 1 to 6, characterized in that the steps include: Collect time series data of industrial production equipment and preprocess it to obtain processed data; dividing and marking the processed data to obtain a marked sequence; reprogramming the tag sequence memory to obtain word representation of time series data; Based on the word representation, the health status of industrial production equipment is collected and predicted to obtain a prediction result.
Citation Information
Cited By
Hydrogen energy equipment health state prediction method and system based on CASP-LLM
CN121502611A
Hydrogen energy equipment health state prediction method and system based on casp-llm
CN121502611B