An intelligent cloud file storage method and medium for a power plant based on cloud computing
By applying the intelligent cloud archive storage method on the cloud computing platform of the power plant, and using the core text index model to analyze the power plant operation logs, the problem of insufficient log analysis efficiency and accuracy in the existing technology is solved, and more efficient log text processing and key information extraction is achieved.
Patent Information
- Application Number
- CN202411135954.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-08-19
AI Technical Summary
The prior art is difficult to effectively analyze and process complex and variable power plant operation logs, especially when faced with unstructured or semi-structured data, with insufficient accuracy and efficiency.
The intelligent cloud archive storage method based on cloud computing is adopted. By obtaining the running log text of the power plant's operating status, the implicit representation of the target text is extracted, and these implicit representations are iteratively optimized based on the core text index model to extract the implicit representation of the core text item of the content of interest.
It improves the detection accuracy and efficiency of the content of interest in the log text, and can more accurately capture the deep semantic information and contextual relationships in the log, meeting the high requirements of power plants for real-time and accuracy.
Smart Images

Figure CN119127823B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and more particularly, to an intelligent cloud file storage method and medium for a power plant based on cloud computing. Background Art
[0002] With the rapid development of information technology and industrial automation technology, the operation and maintenance management of power plants increasingly relies on large-scale data processing and analysis capabilities. As a key infrastructure for energy supply, the massive operation logs generated during the daily operation of power plants are crucial for fault troubleshooting, performance optimization, and safety management. However, traditional log analysis methods often rely on manual reading and analysis, which are not only inefficient but also error-prone, making it difficult to meet the high requirements for real-time and accuracy in modern power plants. Currently, although some automated log analysis tools have been developed, most of these tools are based on simple keyword matching or rule engines and are difficult to effectively process complex and ever-changing operation log texts. Especially when faced with unstructured or semi-structured log data, the accuracy and efficiency of these tools are often greatly reduced. In addition, since the operation logs of power plants usually contain a large number of professional terms and complex scenario descriptions, traditional methods based on statistics or shallow machine learning are difficult to capture the deep semantic information and context relationships in the logs. To overcome the above technical problems, the industry has begun to explore the use of advanced artificial intelligence technologies such as deep learning to improve log analysis methods. However, when existing deep learning models are applied to power plant operation log analysis, a complex debugging process is often required in the model training stage, resulting in insufficient efficiency. Summary of the Invention
[0003] The object of the present invention is to provide an intelligent cloud file storage method and medium for a power plant based on cloud computing. The embodiments of the present application are implemented as follows:
[0004] First aspect, an intelligent cloud file storage method for a power plant based on cloud computing, which is applied to a cloud computing server. The method includes: obtaining operation log texts of the operation status of a target power plant uploaded by at least one intelligent terminal device of the power plant to form a to-be-analyzed operation log text; extracting a target text implicit representation from the to-be-analyzed operation log text, where the to-be-analyzed operation log text includes content of interest; iteratively optimizing the target text implicit representation and the implicit representation of a control core text item obtained by arbitrarily assigning values based on a core text index model to obtain an implicit representation of a target core text item of the content of interest, where the implicit representation of the target core text item represents the implicit representation of the core text item detected in the content of interest included in the to-be-analyzed operation log text; wherein, the core text index model is obtained by performing p rounds of iterative optimization on each operation log text template according to an initialized core text index model. In the a-th core text index model after completing the a-th round of iterative optimization, iterative optimization is completed on the b-th template implicit representation binary group obtained in the b-th round of iterative optimization to obtain the a-th template implicit representation binary group, and the weights and biases in the a-th core text index model are updated according to the implicit representation error between the b-th template implicit representation binary group and the a-th template implicit representation binary group, and stop when meeting a preset stop requirement, where 1≤a≤p, b = a - 1, the a-th template implicit representation binary group includes the a-th template text implicit representation and the a-th control template core text item implicit representation, the a-th template text implicit representation is obtained by performing a rounds of iterative optimization on an initial template text implicit representation extracted from the operation log text template, and the a-th control template core text item implicit representation is obtained by performing a rounds of iterative optimization on an initial control template core text item implicit representation obtained by arbitrarily assigning values; determining, based on the implicit representation of the target core text item, the text paragraphs in which the core text item in the content of interest exists in the to-be-analyzed operation log text, and marking the core text item to obtain a marked operation log text, and storing the marked operation log text.
[0005] Second aspect, the present application provides an intelligent cloud file storage device for a power plant based on cloud computing. The method includes: a log text acquisition module, configured to acquire operation log texts of the operation status of a target power plant uploaded by at least one intelligent terminal device of the target power plant, and form a to-be-analyzed operation log text; an implicit representation extraction module, configured to extract a target text implicit representation from the to-be-analyzed operation log text, where the to-be-analyzed operation log text includes content of interest; a core feature determination module, configured to perform iterative optimization on the target text implicit representation and the implicit representation of a control core text item obtained by arbitrary assignment start based on a core text index model, to obtain an implicit representation of a target core text item of the content of interest, where the implicit representation of the target core text item represents an implicit representation of a core text item detected in the content of interest included in the to-be-analyzed operation log text; where the core text index model is obtained by performing p rounds of iterative optimization on each operation log text template according to an initialized core text index model. In the a-th core text index model after completing the a-th round of iterative optimization, iterative optimization is performed on the b-th template implicit representation binary group obtained in the b-th round of iterative optimization to obtain the a-th template implicit representation binary group, and the weights and biases in the a-th core text index model are updated according to the implicit representation error between the b-th template implicit representation binary group and the a-th template implicit representation binary group, and stop when meeting a preset stop requirement, where 1 ≤ a ≤ p, b = a - 1, the a-th template implicit representation binary group includes an a-th template text implicit representation and an a-th control template core text item implicit representation, the a-th template text implicit representation is obtained by performing a rounds of iterative optimization on an initial template text implicit representation extracted from the operation log text template, and the a-th control template core text item implicit representation is obtained by performing a rounds of iterative optimization on an initial control template core text item implicit representation obtained by arbitrary assignment start; a log text storage module, configured to determine, based on the implicit representation of the target core text item, a text paragraph in which the core text item in the content of interest exists in the to-be-analyzed operation log text, mark the core text item, obtain a marked operation log text, and store the marked operation log text.
[0006] Third aspect, the present application provides a cloud computing server, including: one or more processors; a memory; one or more computer programs; where the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method described in the first aspect above is implemented.
[0007] Fourthly, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a processor, the processor can execute the method described in the first aspect above.
[0008] The beneficial effects of the present application at least include: The present application extracts the implicit representation of the target text from the running log text to be analyzed. The running log text to be analyzed includes the content of interest. Based on the core text index model, the implicit representation of the target text and the implicit representation of the control core text item obtained by arbitrary assignment start are iteratively optimized to obtain the implicit representation of the target core text item of the content of interest. The implicit representation of the target core text item represents the implicit representation of the core text item detected in the content of interest included in the running log text to be analyzed. The core text index model is obtained by performing p rounds of iterative optimization on each running log text template according to the initialized core text index model. In the a-th core text index model after completing the a-th round of iterative optimization, the iterative optimization of the b-th template implicit representation binary group obtained in the b-th round of iterative optimization is completed to obtain the a-th template implicit representation binary group, and the weights and biases in the a-th core text index model are updated according to the implicit representation error between the b-th template implicit representation binary group and the a-th template implicit representation binary group, and stop when meeting the preset stop requirements, 1≤a≤p, b = a - 1. The a-th template implicit representation binary group includes the a-th template text implicit representation and the a-th control template core text item implicit representation. The a-th template text implicit representation is obtained by performing a rounds of iterative optimization on the initial template text implicit representation extracted from the running log text template. The a-th control template core text item implicit representation is obtained by performing a rounds of iterative optimization on the initial control template core text item implicit representation obtained by arbitrary assignment start. Then, based on the implicit representation of the target core text item, the text paragraph where the core text item in the content of interest exists in the running log text to be analyzed is determined, and the core text item is marked to obtain the marked running log text, and the marked running log text is stored. That is to say, the process of determining the text paragraph where the core text item in the content of interest included in the running log text to be analyzed exists by using the core text index model obtained by iterative self-guided learning in the present application can not only improve the detection accuracy but also improve the detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.
[0010] Figure 1 It is a flowchart of an intelligent cloud file storage method for a power plant based on cloud computing provided by an embodiment of the present application.
[0011] Figure 2 It is a schematic diagram of the functional module architecture of an intelligent cloud file storage device for a power plant based on cloud computing provided by an embodiment of the present application.
[0012] Figure 3 It is a schematic diagram of the composition of a cloud computing server provided by an embodiment of the present application. Detailed implementation manners
[0013] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the implementation manners part of the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application.
[0014] In the embodiments of the present application, the execution subject of the intelligent cloud file storage method for a power plant based on cloud computing is a cloud computing server, including but not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers in cloud computing. Among them, cloud computing is a type of distributed computing, which is a super virtual computer composed of a group of loosely coupled computer sets. Among them, the cloud computing server can run alone to implement the present application, or can be connected to the network and implement the present application through interaction with other cloud computing servers in the network. Among them, the network where the cloud computing server is located includes but not limited to the Internet, wide area network, metropolitan area network, local area network, VPN network, etc.
[0015] The embodiments of the present application provide an intelligent cloud file storage method for a power plant based on cloud computing. This method is applied to a cloud computing server, as Figure 1 shown, this method includes:
[0016] Step S100: Obtain the operation log text of the operation status of a target power plant uploaded by at least one intelligent terminal device, and form a to-be-analyzed operation log text.
[0017] As a real-time scenario, assume that in a large power group, there are multiple power plants distributed in different regions. In order to monitor and ensure the stable operation of these power plants, each power plant is equipped with a variety of intelligent terminal devices, such as sensors, controllers, and data recorders. These devices collect the operation status data of the power plant in real time, including but not limited to information such as the temperature, pressure, speed, power generation, and fuel consumption of the generator set, and convert these data into operation log text.
[0018] Whenever intelligent terminal devices (such as temperature sensors, pressure gauges, etc.) complete a data sampling, they will automatically upload the collected data to the data center of the power plant or the cloud computing server via wired or wireless means (such as 4G / 5G network, Wi-Fi, or dedicated industrial Ethernet). The format of data upload is usually in the standard log format, which contains key information such as timestamp, device ID, parameter name, parameter value, etc. For example, a typical log record might look like this: "2023-04-01 12:00:00, Turbine1_TempSensor, 500℃", indicating that at 12:00 noon on April 1, 2023, the temperature sensor of Turbine 1 read 500 degrees Celsius. The cloud computing server listens for the data streams from intelligent terminal devices of each power plant through predefined API interfaces or message queues (such as Kafka, RabbitMQ).
[0019] When receiving a new operation log text, the cloud computing server first validates and cleans the data to ensure the integrity and accuracy of the data. For example, checking whether the timestamp is valid, whether the parameter value is within a reasonable range, etc.
[0020] The preliminarily processed data is integrated into a unified format to form the operation log text to be analyzed. These texts may contain the operation records of multiple power plants and various devices at different time points.
[0021] For the convenience of subsequent processing, the cloud computing server can further organize these texts, such as grouping them by power plant, device type, or time period.
[0022] The operation log text to be analyzed is securely stored in the database on the cloud computing platform to ensure the reliability and accessibility of the data. These databases may adopt a distributed architecture, such as Hadoop HDFS or Amazon S3, to support the storage and efficient query of large-scale data. At the same time, the cloud computing server will also regularly back up these data to prevent data loss or damage.
[0023] Through such a process, step S100 ensures the timely and accurate collection of the operation status data of the power plant, providing a solid foundation for subsequent data analysis and intelligent decision-making. In subsequent steps, the cloud computing server will utilize these data, combined with machine learning models (such as LSTM recurrent neural network for time series prediction, or natural language processing models such as BERT for text mining), to further explore the data value and improve the operation and maintenance efficiency and safety of the power plant.
[0024] Step S200: Extract the implicit representation of the target text from the operation log text to be analyzed, where the operation log text to be analyzed includes the content of interest.
[0025] The operation log text to be analyzed refers to the set of power plant operation status log texts that are ready for in-depth analysis after preliminary screening and processing. These texts usually contain various contents such as the operation status of power plant equipment, parameter changes, warning messages, fault records, etc. For example, "2023-04-01 12:00, Turbine 1, RPM: 3000, Temp: 500℃, Warning: High Temperature" is a record in the operation log text to be analyzed, which records the rotational speed, temperature, and high-temperature warning information of Turbine 1 at a specific time point.
[0026] The implicit representation of the target text refers to converting the key information or semantic content in the text into a form that can be directly processed and understood by a computer, usually a high-dimensional vector or tensor. This representation method captures the deep semantic features of the text, rather than just the surface words.
[0027] The content of interest refers to the information in the operation log text that is of great value for power plant operation and maintenance management, fault diagnosis, performance optimization, etc. Such information includes, for example, the abnormal status of equipment, parameters exceeding the normal range, specific warning or error messages, etc. For example, in the above log text, "Warning: High Temperature" is one of the contents of interest because it indicates that the turbine may have an overheating problem, which requires the attention and handling of operation and maintenance personnel.
[0028] Specifically, the cloud computing server performs the following operations to complete step S200: First, the cloud computing server receives and processes the operation log text to be analyzed from step S100. These texts contain various contents such as the operation data, fault records, and maintenance information of various power plant equipment, with a large amount of data and diverse formats. To extract valuable information from these unstructured texts, the cloud computing server uses specific NLP techniques. Then, the cloud computing server uses a pre-trained word embedding model (such as Word2Vec, GloVe, or BERT's Token Embeddings) to convert each word or phrase in the text into a high-dimensional vector, which is the implicit representation of the word or phrase. This conversion enables the cloud computing server to understand the semantic relationships between words and thus capture the context information in the text.
[0029] For example, in a log record "2023-04-01 14:00, Turbine 3 overheated, automatically shut down", the cloud computing server will first convert each word (such as "Turbine", "overheated", "automatically", "shut down") into a corresponding word vector. Subsequently, using deep learning models such as the Attention Mechanism or Convolutional Neural Network (CNN), the cloud computing server can identify the key phrase "Turbine 3 overheated" because it is closely related to the equipment operation status of the power plant and is the content of interest. In order to more accurately represent this key information, the cloud computing server can further use a recurrent neural network (RNN) or its variants (such as LSTM, GRU) to process the entire sentence or paragraph, thereby generating a more advanced implicit representation vector that not only contains the information of the keywords, but also incorporates the semantic features of the entire context. In addition, for certain specific tasks (such as fault prediction and anomaly detection), the cloud computing server may also combine domain knowledge bases or expert cloud computing servers to enhance the accuracy and relevance of the implicit representation of the text through rule matching or conditional reasoning. Finally, after processing in step S200, the cloud computing server obtains a series of implicit representations of the target text, which accurately capture the interesting content in the run log text to be analyzed, providing strong data support for subsequent analysis and storage. These implicit representations will be passed as input to the core text index model in step S300 to further explore the data value and optimize storage efficiency.
[0030] Step S300: Iteratively optimize the implicit representation of the target text and the implicit representation of the control core text item obtained by any assignment based on the core text index model to obtain the target core text item implicit representation of the content of interest, wherein the target core text item implicit representation represents the implicit representation of the core text item detected in the content of interest contained in the run log text to be analyzed.
[0031] In step S300, the cloud computing server uses a pre-trained core text index model to deeply process the implicit representation of the target text obtained from step S200 to accurately identify and characterize the core text items of the content of interest that are critical to the operation and maintenance management of the power plant.
[0032] Specifically, the cloud computing server first initializes an implicit representation of the reference core text item, which is usually done by random assignment (i.e., starting with arbitrary assignment) to ensure that the model does not get stuck in a local optimum during the iterative optimization process. This reference representation is used as a benchmark and is input into the core text indexing model together with the implicit representation of the target text. The core text indexing model itself is, for example, a complex neural network architecture that combines advanced components such as recurrent neural networks (e.g., LSTM), convolutional neural networks (CNN), or self-attention mechanisms (e.g., Transformer). These models are designed to process sequence data (such as text) and can learn the internal structure and patterns of the data from it. During the iterative optimization process, the model repeatedly adjusts its internal parameters (such as weights and biases) to minimize the objective function (such as the cross-entropy loss function), thereby improving the model's recognition accuracy for the implicit representation of the target text. This optimization process typically involves the backpropagation algorithm, which adjusts the model parameters based on the difference between the model's prediction results and the true labels.
[0033] The optimized implicit representation of the target core text item will more accurately reflect the key information in the original text, providing strong support for subsequent tasks such as log analysis, fault warning, and performance evaluation. For example, this implicit representation can be used to build a real-time monitoring cloud computing server for a power plant. Once a text pattern like "Turbine overheating" is detected, the cloud computing server can immediately trigger an alarm and take corresponding countermeasures.
[0034] It should be noted that due to the complexity of the core text indexing model and the diversity of power plant operation logs, the specific process and effect of iterative optimization will be affected by various factors, including the initial parameter settings of the model, the quality and quantity of training data, and the choice of optimization algorithm. Therefore, in practical applications, the model needs to be carefully adjusted and optimized according to specific circumstances.
[0035] Step S400: Based on the implicit representation of the target core text item, determine the text paragraphs in the text of the operation log to be analyzed where the core text item in the content of interest exists, mark the core text item, obtain the marked operation log text, and store the marked operation log text.
[0036] In step S400, the cloud computing server uses the implicit representation of the target core text item obtained from step S300 to accurately locate and mark the text paragraphs where the core text item of the content of interest in the text of the operation log to be analyzed is located, and then generates the marked operation log text and stores it securely on the cloud platform.
[0037] Specifically, the cloud computing server first searches and matches in the original operation log text according to the semantic features and location information contained in the implicit representation of the target core text item. This process may involve string matching algorithms, regular expression matching, or more advanced text parsing techniques, depending on the form and complexity of the implicit representation of the target core text item. Take a typical power plant operation log as an example: "2023-04-15 14:30, Turbine 2, Warning: Oil pressure low, please check filters and replenish oil as necessary. System operating normally otherwise." In this log, "Oil pressure low" is identified as the core text item of the content of interest, and its implicit representation has been optimized in step S300 to the extent that it can accurately characterize this key information. In step S400, the cloud computing server successfully locates the text paragraph containing "Oil pressure low": "Warning: Oil pressure low, please check filters and replenish oil as necessary." by comparing the implicit representation of the target core text item with each paragraph in the log text. Subsequently, the cloud computing server will mark this paragraph, and the marking method can be highlighting, adding specific tags or annotations, etc., for subsequent reference and analysis.
[0038] After the marking is completed, the cloud computing server will generate a marked operation log text. This text not only retains all the information of the original log but also intuitively highlights the location of the key information through marking, greatly improving the readability and usability of the log. Finally, this marked operation log text is securely stored on the cloud platform, leveraging the high availability and scalability of cloud computing to ensure the persistent preservation and fast access of data.
[0039] By executing step S400, the power plant operation and maintenance personnel can more conveniently obtain the key information in the operation log, take necessary measures in a timely manner to address potential problems or risks, thereby ensuring the safe and stable operation of the power plant. At the same time, the storage of the marked operation log text also provides a valuable data source for subsequent data analysis and mining, helping the power plant continuously optimize the operation and maintenance strategy and improve the operation efficiency.
[0040] In the embodiment of the present application, the core text index model is obtained by performing p rounds of iterative optimization on each operation log text template based on the initialized core text index model. In the a-th core text index model after completing the a-th round of iterative optimization, iterative optimization is performed on the b-th template implicit representation binary tuple obtained in the b-th round of iterative optimization to obtain the a-th template implicit representation binary tuple, and the weights and biases in the a-th core text index model are updated according to the implicit representation error between the b-th template implicit representation binary tuple and the a-th template implicit representation binary tuple, and stop when meeting the preset stop requirements, where 1 ≤ a ≤ p and b = a - 1. The a-th template implicit representation binary tuple includes the a-th template text implicit representation and the a-th control template core text item implicit representation. The a-th template text implicit representation is obtained by performing a rounds of iterative optimization on the initial template text implicit representation extracted from the operation log text template, and the a-th control template core text item implicit representation is obtained by performing a rounds of iterative optimization on the initial control template core text item implicit representation obtained by starting with arbitrary assignment.
[0041] The training process of the core text index model is a crucial step in ensuring that the model can accurately extract and index key information from power plant operation logs. This process involves iterative learning of a large number of operation log text templates, continuously optimizing the model's parameters to improve its ability to identify core text items.
[0042] First, the cloud computing server initializes a core text indexing model. This model is, for example, a complex neural network structure, such as a Transformer-based encoder-decoder architecture, which can learn from text data and generate effective implicit representations. During initialization, the weights and biases of the model are usually set to random values or predefined default values. Next, the cloud computing server starts iterative training of the model using the running log text templates. These templates are carefully selected from historical running logs and contain various typical scenarios and key information in the operation of the power plant. The training process is divided into multiple rounds (set as p rounds), and each round aims to further optimize the performance of the model. In the a-th round of iterative optimization (where 1 ≤ a ≤ p), the cloud computing server first loads the b-th template implicit representation binary tuple obtained after the (a - 1)-th (i.e., the b-th, since b = a - 1) round of iterative optimization. This binary tuple consists of two parts: the b-th template text implicit representation and the b-th control template core text item implicit representation. The former is the result of optimizing the initial text representation extracted from the running log text template after b rounds, and the latter is the control core text item implicit representation obtained by arbitrary assignment startup (i.e., random initialization) and optimized for the same number of rounds. Then, the cloud computing server uses the current core text indexing model to further iteratively optimize these two implicit representations. The optimization process may include steps such as forward propagation, loss calculation, backpropagation, and parameter update. Specifically, the model will attempt to predict the core text items in the current text and compare them with the known control core text items to calculate a loss value (such as cross-entropy loss), which measures the difference between the model prediction and the actual situation.
[0043] The loss value is then used to adjust the weights and biases of the model through the backpropagation algorithm to reduce future prediction errors. This adjustment process is based on gradient descent, that is, updating the parameter values along the negative gradient direction of the loss function with respect to the model parameters. The specific update formula is, for example:
[0044]
[0045] Among them, θ represents the model parameters (weights and biases), L represents the loss function, and η is the learning rate, which controls the step size of parameter updates. After a rounds of iterative optimization, the cloud computing server obtains the a - th template implicit representation tuple, including a more accurate implicit representation of the a - th template text and an implicit representation of the core text items of the a - th control template. At the same time, according to the implicit representation error between the b - th template implicit representation tuple and the a - th template implicit representation tuple (measured by calculating the Euclidean distance or cosine similarity between the two), the cloud computing server further adjusts the parameters in the core text index model to minimize this error. When the model performance reaches the preset stop requirements (such as the loss value is lower than a certain threshold, the accuracy on the validation set no longer improves significantly, etc.), the training process ends. At this time, the cloud computing server obtains a well - trained core text index model, which can accurately extract and index core text items from the power plant operation logs. Through this process, the cloud computing server not only improves the model's understanding ability of specific - domain texts, but also enhances its generalization and robustness in practical applications.
[0046] In a possible implementation, step S300, iteratively optimizing the implicit representation of the target text and the implicitly represented control core text items obtained by arbitrary assignment based on the core text index model to obtain the implicitly represented target core text items of the content of interest, may include:
[0047] Complete the following process in the core text index model:
[0048] Step S310: Load the x - th target text implicit representation and the x - th implicitly represented control core text items output by the x - th iterative optimization component in the obtained core text index model into the (x + 1)-th iterative optimization component in the core text index model, where the core text index model includes t iterative optimization components, and 1 ≤ x ≤ t.
[0049] When the cloud computing server executes step S310, the input implicit representation of the target text has been preliminarily processed by the first x iterative optimization components of the core text index model. Here, the implicit representation of the target text is obtained by converting the power plant operation log text through natural language processing (NLP) technology. It contains the semantic information of the log text but has not been precisely focused on the core text items of interest.
[0050] After the x-th iterative optimization component completes its task, it outputs two important results: the x-th implicit representation of the target text and the x-th implicit representation of the control core text item. The x-th implicit representation of the target text is the latest understanding of the target text content in the current iteration round, while the x-th implicit representation of the control core text item is a reference representation corresponding to the true core text item, which is generated, for example, through random initialization or other means, to assist the model in learning during the training process. Subsequently, step S310 loads these outputs as inputs into the (x + 1)-th iterative optimization component in the core text indexing model. This component will continue to process the inputs more deeply to further refine and strengthen the key information in the text. For example, if the x-th component is a Transformer encoder layer based on the attention mechanism, it may have given higher weights to the key parts of the input text; the (x + 1)-th component may, on this basis, use more complex neural network structures (such as stacked Transformer layers, LSTM layers, or convolutional layers) to capture more subtle semantic features and context relationships. During the loading process, the cloud computing server ensures the integrity and consistency of the data, and at the same time, according to the requirements of the model architecture, it may be necessary to adjust the format or dimension of the data to meet the input needs of the (x + 1)-th iterative optimization component. This transfer and processing of data are completed automatically without manual intervention. It should be noted that the core text indexing model usually contains multiple (assumed to be t) such iterative optimization components, which process the input data in sequence until all components have been executed. Therefore, step S310 is not only a key step in a single iteration process but also a guarantee for the continuous progress of the entire iterative optimization process. Through such a series of iterative optimization processes, the core text indexing model can gradually approach an accurate understanding of the content of interest in the power plant operation log and finally output the implicit representation of the target core text item, providing strong support for subsequent data storage, analysis, and utilization.
[0051] Step S320: In the (x + 1)-th iterative optimization component, iteratively optimize the x-th implicit representation of the target text and the x-th implicit representation of the control core text item to obtain the (x + 1)-th implicit representation of the target text and the (x + 1)-th implicit representation of the control core text item.
[0052] In step S320, the cloud computing server transfers the x-th implicit representation of the target text and the x-th implicit representation of the control core text item output by the x-th iterative optimization component as inputs to the (x + 1)-th iterative optimization component. This component is, for example, a complex neural network layer, such as a Transformer-based encoder layer, an LSTM layer, or a convolutional neural network layer combined with the attention mechanism. These neural network layers deeply analyze and process the input data through their specific algorithms and parameters.
[0053] Taking the Transformer encoder layer as an example, this layer first uses the self-attention mechanism to process the input implicit representation. The self-attention mechanism allows the model to refer to the information at all positions in the input sequence when processing the information at each position, thereby capturing the long-range dependencies in the text. In the (x + 1)-th iterative optimization component, the self-attention mechanism calculates the attention weights between the feature vectors in the x-th target text implicit representation and the x-th control core text item implicit representation, and performs a weighted sum of the feature vectors according to these weights to generate a new set of feature vectors. Next, this layer can apply a positional encoding to preserve the order information of the elements in the input sequence because the Transformer model itself does not understand the order of the input. The positional encoding is added to the feature vectors generated in the previous step to obtain a set of feature vectors containing positional information.
[0054] Subsequently, these feature vectors pass through a series of feed-forward neural networks, which usually include two linear transformation layers and a ReLU activation function. The feed-forward neural networks perform further non-linear transformations on the feature vectors to extract higher-level features.
[0055] After the above processing, the (x + 1)-th iterative optimization component outputs the (x + 1)-th target text implicit representation and the (x + 1)-th control core text item implicit representation. These two implicit representations are more refined and accurate than the input because they have been deeply optimized through complex neural network layers.
[0056] Specifically, the (x + 1)-th target text implicit representation may contain a series of feature vectors that are more focused on the content of interest in the power plant operation log, such as features related to equipment failures, performance anomalies, etc. Although the (x + 1)-th control core text item implicit representation is still randomly initialized, after this round of iterative optimization, the gap between it and the true core text item implicit representation has been narrowed, providing a better starting point for subsequent iterations.
[0057] It should be noted that the entire iterative optimization process is a cyclic process. The output of the (x + 1)-th iterative optimization component will become the input of the next iterative optimization component until all t iterative optimization components have been executed. Through such continuous iteration and optimization, the core text indexing model can gradually approach an accurate understanding of the content of interest in the power plant operation log.
[0058] Step S330: When x + 1 = t, use the (x + 1)-th control core text item implicit representation as the target core text item implicit representation.
[0059] Step S330 is an important judgment and execution point in the iterative optimization process of the core text index model. It marks the end of the entire iterative process and determines the final optimized result as the implicit representation of the target core text item. In the iterative optimization process of the core text index model, the cloud computing server sequentially processes the implicit representation of the input target text and the implicit representation of the reference core text item through t iterative optimization components. Each component further optimizes based on the output of the previous component to generate a new implicit representation. This process is carried out cyclically until all components have been traversed once. When the iteration variable x increases such that x + 1 is equal to the total number of components t, it means that all the predetermined iterative optimization components have completed the processing of the input data. At this time, Step S330 is triggered, and the cloud computing server regards the implicit representation of the reference core text item at the current (i.e., the output of the t-th component) as the final implicit representation of the target core text item. Although the implicit representation of the reference core text item is randomly initialized at the beginning of the iteration, through the layer-by-layer processing and optimization of multiple iterative optimization components, it gradually approaches the core implicit representation of the truly interesting content in the target text. In the scenario of intelligent cloud archive storage in a power plant, this interesting content may include, but is not limited to, equipment fault alarms, performance anomaly indicators, changes in key operating parameters, etc.
[0060] Specifically, if the target text is about the operation log of a certain generator set, which contains the description "Generator temperature is too high, automatically start the cooling cloud computing server", then through the iterative optimization of the core text index model, the final implicit representation of the target core text item obtained will be able to highly represent the key information that the generator temperature is too high. This implicit representation is, for example, a set of a series of feature vectors, and each vector corresponds to the semantic features of a certain keyword or phrase in the text, jointly constituting an accurate description of the core text item.
[0061] Therefore, in Step S330, the cloud computing server does not have a further calculation or optimization process, but directly takes the implicit representation of the reference core text item output by the last iterative optimization component as the final result. This result can then be used in subsequent data analysis, storage, or retrieval tasks to provide strong support for the operation and maintenance management of the power plant.
[0062] Step S340: When x + 1 < t, load the obtained (x + 1)-th target text implicit representation and the (x + 1)-th reference core text item implicit representation into the (x + 2)-th iterative optimization component to perform the (x + 2)-th iterative optimization.
[0063] Step S340 ensures the continuity of the entire iterative process until the preset number of iterations t is reached. In the scenario of intelligent cloud file storage in a power plant based on cloud computing, the execution of this step is crucial for accurately extracting key information from a large amount of operation logs.
[0064] When executing step S340, the cloud computing server first checks whether the current iteration round x + 1 is less than the total number of iterations t. If the condition is met, that is, there are more iterative optimization components to process the input data, then the cloud computing server will continue to execute the iterative optimization process.
[0065] In this step, the cloud computing server takes the (x + 1)-th target text implicit representation and the (x + 1)-th control core text item implicit representation output by the (x + 1)-th iterative optimization component as inputs and loads them into the (x + 2)-th iterative optimization component. These two implicit representations are the results of the previous round of iterative optimization. They have passed through the complex processing of the (x + 1)-th component and contain richer and more accurate text information.
[0066] The (x + 2)-th iterative optimization component is, for example, a complex model with a multi-layer neural network structure. It will further analyze and optimize the input implicit representations. This component can use convolutional layers to extract local features in the text, use recurrent layers to capture temporal dependencies in sequential data, or use attention mechanisms to focus on the most critical parts of the text.
[0067] Taking the Transformer encoder layer as an example, the (x + 2)-th component can use the self-attention mechanism to enable the model to refer to the information at all positions in the input sequence when processing the information at each position, thereby capturing the global dependencies in the text. Through the multi-head attention mechanism, the model can process the information from different representation subspaces in parallel, further enhancing its ability to understand the text content.
[0068] Inside the (x + 2)-th iterative optimization component, the computing cloud server will perform a series of complex calculations, including linear transformation, activation function application, normalization operations, etc., to generate new implicit representations. These new implicit representations will more accurately reflect the key information in the text and provide strong support for subsequent data processing and analysis. After completing the processing of the (x + 2)-th iterative optimization component, the cloud computing server will obtain the (x + 2)-th target text implicit representation and the (x + 2)-th control core text item implicit representation. These two new implicit representations will then be passed as inputs to the next iterative optimization component (if any) to continue the iterative optimization process.
[0069] By continuously repeating step S340 until x + 1 equals t, the cloud computing server can ensure that the core text index model has fully iteratively optimized the input text. The finally obtained target implicit representation of the core text item will very precisely reflect the key information in the text, providing strong data support for the operation and maintenance management of the power plant.
[0070] In a possible implementation, in step S320, in the (x + 1)-th iterative optimization component, iteratively optimizing the x-th target text implicit representation and the x-th control core text item implicit representation to obtain the (x + 1)-th target text implicit representation and the (x + 1)-th control core text item implicit representation may include:
[0071] Complete the following process in the (x + 1)-th iterative optimization component:
[0072] Step S321: In the first iterative optimization module of the (x + 1)-th iterative optimization component, perform a linear mapping on the log text sub-implicit representation output by the x-th iterative optimization component to obtain the first log text sub-implicit representation; perform a linear mapping on the control core text item sub-implicit representation output by the x-th iterative optimization component to obtain the first control core text item sub-implicit representation.
[0073] In step S321, the cloud computing server preliminarily processes the log text sub-implicit representation and the control core text item sub-implicit representation output by the x-th iterative optimization component through linear mapping to generate new implicit representations, laying a foundation for subsequent more complex optimizations. Specifically, when the cloud computing server executes step S321, it first focuses on the output results of the x-th iterative optimization component. These outputs include two parts: the log text sub-implicit representation and the control core text item sub-implicit representation. The log text sub-implicit representation is usually a high-dimensional vector that captures the semantic information of a sub-paragraph or keyword in the power plant operation log; while the control core text item sub-implicit representation is a randomly initialized vector corresponding to the true implicit representation of the core text item, used to guide the model to learn during the iteration process.
[0074] To further refine this information in the (x + 1)-th component, the cloud computing server will apply linear mapping within the first iterative optimization module. Linear mapping is a basic mathematical transformation that converts an input vector into a new output vector through a weight matrix and a bias vector. In this scenario, linear mapping is used to perform a preliminary spatial transformation on the input sub-implicit representation to make it more suitable for subsequent optimization steps.
[0075] Suppose the log text sub-implicit representation output by the x-th iterative optimization component is where d x is the dimension of this vector; the control core text item sub-implicit representation is The dimensions are the same. In the first iterative optimization module, the cloud computing server defines two sets of weight matrices W text ∈R dx+1×dx and W core ∈R dx+1×dx , and two sets of bias vectors b text ∈R dx+1 and b core ∈R dx+1 , where d x+1 is the dimension of the new implicit representation, which may be different from the dimension of the input vector.
[0076] Then, the cloud computing server calculates the first log text sub-implicit representation and the first control core text item sub-implicit representation
[0077]
[0078] These two newly generated implicit representations are then used as the output of the first iterative optimization module and may be further passed to other iterative optimization modules within this component for further optimization.
[0079] Through such linear mapping processing, the cloud computing server can preliminarily transform the input sub-implicit representation in a simple and effective way, laying a foundation for subsequent non-linear transformation and feature extraction. This processing method not only retains the main features of the input information but also provides more optimization space for the model by introducing new dimensions and linear combinations.
[0080] Step S322: In the e-th iterative optimization module of the (x + 1)-th iterative optimization component, load the f-th log text sub-implicit representation and the f-th control core text item sub-implicit representation output by the f-th iterative optimization module of the (x + 1)-th iterative optimization component into the e-th iterative optimization module in the (x + 1)-th iterative optimization component, where 2 ≤ e and f = e - 1.
[0081] In step S322, the cloud computing server is responsible for using the output of the previous iterative optimization module (the f-th module, where f = e - 1) as the input of the current iterative optimization module (the e-th module) to achieve continuous data processing and optimization. Specifically, within the (x + 1)-th iterative optimization component, there are multiple iterative optimization modules arranged in sequence. Each module performs specific processing on the input data and passes the processing result to the next module.
[0082] Taking the scenario of intelligent cloud file storage in a power plant as an example, assume that the (x + 1)-th iterative optimization component is processing an operation log text about abnormal generator temperature. In this log, the key information (i.e., the core text items) may include words or phrases such as "generator temperature" and "abnormal increase". During the iterative optimization process, each module tries to extract more representative features from these words or phrases to more accurately represent the key information in the log.
[0083] In step S322, the cloud computing server first obtains the output of the f-th iterative optimization module (i.e., the previous module), which includes the f-th log text sub-implicit representation and the f-th control core text item sub-implicit representation. These two implicit representations respectively capture the semantic information of a part of the log text and the current state of the randomly initialized information corresponding to the core text item.
[0084] Then, the cloud computing server loads these two implicit representations as inputs into the e-th iterative optimization module. Here, e is an integer greater than f (since f = e - 1, so e is at least 2), indicating that the data is flowing forward along the module chain. Inside the e-th module, the cloud computing server further processes these inputs, which may include various operations such as linear transformation, non-linear activation, feature extraction, and attention mechanism application, depending on the design and purpose of the module. These operations are aimed at extracting more valuable features from the input data to more accurately represent the core text item.
[0085] It should be noted that step S322 not only realizes the transfer of data between modules but also implies the dependency relationship and optimization order between modules. Since the output of each module is calculated based on its input, the arrangement order and internal processing logic of the modules have an important impact on the final optimization result. By continuously repeating step S322 and subsequent optimization steps (such as steps S323 and S324), the cloud computing server can gradually refine the most critical and representative information in the log text and store it in the core text index model in the form of implicit representation. These implicit representations not only facilitate subsequent log retrieval and analysis tasks but also provide strong data support for the operation and maintenance management of the power plant.
[0086] Step S323: In the e-th iterative optimization module, perform a linear mapping on the f-th log text sub-implicit representation to obtain a first mapped feature vector, and use the first mapped feature vector as the e-th log text sub-implicit representation output by the e-th iterative optimization module; and perform a linear mapping on the f-th control core text item sub-implicit representation to obtain a second mapped feature vector; based on the first mapped feature vector and the second mapped feature vector, obtain the e-th control core text item sub-implicit representation output by the e-th iterative optimization module.
[0087] Step S323 describes how to process the input sub - implicit representation of log text and the sub - implicit representation of the control core text item within the e - th iterative optimization module, and generate a new implicit representation for use in subsequent steps. In the cloud computing server, when executing the e - th iterative optimization module of the (x + 1) - th iterative optimization component, it first receives two inputs from the f - th module (where f = e - 1): the f - th sub - implicit representation of log text and the f - th sub - implicit representation of the control core text item. These two inputs respectively represent the intermediate representations of the log text and the control core text item within the model in the current iteration round.
[0088] The processing process of step S323 can be broken down into the following key steps:
[0089] The cloud computing server first performs a linear mapping on the f - th sub - implicit representation of log text. Linear mapping usually involves a weight matrix and a bias vector The first mapped feature vector is calculated by the following formula
[0090]
[0091] where, is the sub - implicit representation of log text output by the f - th module, and is the first mapped feature vector obtained after linear mapping. This vector captures higher - level feature information in the log text.
[0092] In some cases, the first mapped feature vector itself can directly serve as the sub - implicit representation of log text output by the e - th iterative optimization module However, in more complex models, it may be necessary to further process it through a non - linear activation function (such as ReLU, Sigmoid, etc.) or other transformation steps However, for the sake of simplicity in illustration, it is assumed here that is directly adopted as the output. as the output.
[0093] Similarly, the cloud computing server performs a linear mapping on the f - th sub - implicit representation of the control core text item to obtain the second mapped feature vector The weight matrix and bias vector used in this process may be different from those used when processing log text, and are denoted as and
[0094]
[0095] where, is the sub - implicit representation of the control core text item output by the f - th module.
[0096] Finally, the cloud computing server generates the sub - implicit representation of the control core text item output by the e - th iterative optimization module based on the first mapping feature vector and the second mapping feature vector This step may involve various operations such as concatenation, summation, dot product, or more complex non - linear transformations. For illustration, a simple fusion method such as weighted summation can be assumed. This step may involve various operations, such as splicing, summing, dot - product, or more complex non - linear transformations. To illustrate, a simple fusion method, such as weighted summation, can be assumed.
[0097] Through the above steps, the cloud computing server has completed the processing within the e - th iterative optimization module and generated a new sub - implicit representation of the log text and a sub - implicit representation of the control core text item. These representations will be used as the input for the next round of iteration and continue to participate in the model optimization process.
[0098] Step S324: When the e - th iterative optimization module is the last iterative optimization module in the (x + 1) - th iterative optimization component, use the e - th sub - implicit representation of the log text as the (x + 1) - th target text implicit representation obtained by executing the (x + 1) - th iterative optimization component, and use the e - th sub - implicit representation of the control core text item as the (x + 1) - th control core text item implicit representation obtained by executing the (x + 1) - th iterative optimization component.
[0099] In step S324, according to the position of the currently processed iterative optimization module, it is decided when to use the intermediate processing result as the final output of the component. The core text index model processes the input log text through multiple iterative optimization components to extract key information. Each iterative optimization component contains multiple iterative optimization modules inside, and these modules process the input data in sequence, gradually refining and optimizing the implicit representation. When the data flows through the last iterative optimization module, the output of this module is regarded as the final result of the entire component.
[0100] In the cloud computing server, when processing the e - th iterative optimization module of the (x + 1) - th iterative optimization component, the cloud computing server will first check whether this module is the last module within the component. This judgment is usually achieved based on the design parameters of the component or module counting. If it is the last module (that is, there is no other subsequent module that can receive its output), then the following operations are performed:
[0101] Determine the (x + 1) - th target text implicit representation: directly use the e - th sub - implicit representation of the log text output by the e - th iterative optimization module as one of the final outputs of the (x + 1) - th iterative optimization component, that is, the (x + 1) - th target text implicit representation. This implicit representation synthesizes the processing results of all modules within the component on the input log text and is the latest understanding of the key information of the current log text by the model.
[0102] Determine the implicit representation of the (x + 1)-th control core text item: Similarly, the implicit sub-representation of the e-th control core text item output by the e-th iterative optimization module is used as another final output of the (x + 1)-th iterative optimization component, that is, the implicit representation of the (x + 1)-th control core text item. Although this representation is introduced as a randomly initialized vector corresponding to the true core text item during the iteration process, through the iterative optimization of multiple modules, it has gradually approximated the implicit representation of the true core text item.
[0103] For example, assume that the (x + 1)-th iterative optimization component includes three iterative optimization modules (e = 1, 2, 3), and the third module (i.e., e = 3) is currently being processed. When the third module completes its iterative optimization task, the implicit sub-representation of the 3rd log text and the implicit sub-representation of the 3rd control core text item output by it will be regarded as the implicit representation of the (x + 1)-th target text and the implicit representation of the (x + 1)-th control core text item respectively, because these representations are the final results obtained after all modules within the component have been processed.
[0104] Through step S324, the cloud computing server ensures that each iterative optimization component can output meaningful implicit representations, which will be used as inputs for subsequent processing steps (such as steps S330 and S340) and continue to participate in the overall optimization process of the core text indexing model. Finally, when all iterative optimization components have completed processing, the model will be able to generate an accurate implicit representation of the target core text item to support the efficient storage and retrieval of the power plant intelligent cloud archive.
[0105] In a possible implementation, before extracting the implicit representation of the target text from the log text to be analyzed, the method further includes an initialization step of the core text indexing model, which specifically may include:
[0106] Step S200A: Extract the initial template text implicit representation from the obtained u-th log text template, and obtain the initial implicit representation of the control template core text item obtained by arbitrarily assigning values, where 1 ≤ u ≤ q.
[0107] In a cloud computing server for intelligent cloud archives storage of a power plant based on cloud computing, the cloud computing server first obtains a series of operation log text templates from the database of the power plant. These templates represent various typical scenarios in the daily operation of the power plant, such as equipment startup, normal operation, fault alarm, etc., and are of great significance for the model to learn the characteristics of the power plant operation logs. For each obtained operation log text template (numbered the u-th, where 1 ≤ u ≤ q, and q is the total number of templates), the cloud computing server first applies natural language processing technology (NLP) to extract the initial implicit representation of the template text. This step usually involves text vectorization, that is, converting the text into a high-dimensional vector form that can be directly processed by a computer. These vectors capture the key semantic information in the text, enabling the cloud computing server to understand and analyze the text content.
[0108] Specifically, the cloud computing server can use a pre-trained word embedding model (such as Word2Vec, GloVe, or BERT, etc.) to convert each word in the text into a corresponding word vector. Then, through a certain aggregation strategy (such as average pooling, weighted average, or a more complex attention mechanism), these word vectors are combined into a single vector as the initial implicit representation of the template text. Although this vector cannot fully restore all the information of the original text, it is sufficient to capture the main features and meanings of the text.
[0109] At the same time, in order to provide a benchmark corresponding to the implicit representation of the real core text item, the cloud computing server also needs to generate the initial implicit representation of the control template core text item. Since there is no real core text item for reference at this stage, this representation is obtained by arbitrary assignment startup (i.e., random initialization). Specifically, the cloud computing server can randomly generate a vector with the same dimension as the implicit representation of the initial template text but with random content as the control representation. This vector has no practical meaning in the initialization stage, but in the subsequent iterative optimization process, it will serve as a target to guide the model to gradually learn the implicit representation of the real core text item.
[0110] Through the above steps, the cloud computing server generates a pair of initial implicit representations for each operation log text template: the initial implicit representation of the template text and the initial implicit representation of the control template core text item. These two sets of representations will be used as the input data for the initialization of the core text index model and participate in the subsequent iterative optimization process. In the iterative process, the model will gradually learn how to extract an implicit representation closer to the real core text item from the initial implicit representation of the template text, thereby improving the model's ability to understand and analyze the power plant operation logs.
[0111] The following process is cycled for the initial implicit representation of the template text and the initial implicit representation of the control template core text item, and stops when p rounds of iterative optimization are reached:
[0112] Step S200B: In the a-th core text index model that has completed the a-th round of iterative optimization, perform iterative optimization on the b-th template implicit representation binary tuple obtained from the b-th round of iterative optimization to obtain the a-th template implicit representation binary tuple. When a = 1, the b-th template implicit representation binary tuple obtained from the b-th round of iterative optimization represents the initial template implicit representation binary tuple, and the initial template implicit representation binary tuple includes the initial template text implicit representation and the initial control template core text item implicit representation.
[0113] Step S200B involves further optimizing the template implicit representation based on the completed iterative rounds. In each iteration, the cloud computing server processes the template implicit representation using the current core text index model. These template implicit representations were initially extracted from the running log text template through Step S200A and include the initial template text implicit representation and the initial control template core text item implicit representation, which together form the initial template implicit representation binary tuple. As the iteration progresses, these representations are continuously optimized to more accurately reflect the key information in the text.
[0114] Specifically for Step S200B, after completing the a-th round of iterative optimization, the cloud computing server obtains the a-th core text index model. This model has accumulated a large amount of learning experience through the previous a - 1 rounds of iteration and can process text data relatively accurately. At this time, the cloud computing server needs to use this optimized model to perform further iterative optimization on the b-th template implicit representation binary tuple obtained from the b-th round of iterative optimization.
[0115] It should be noted that the b-th round and the a-th round here are not completely independent. In fact, at the beginning of the a-th round of iteration, the cloud computing server needs to use the template implicit representation binary tuple obtained from the b-th round (actually the a - 1-th round because b = a - 1) of iterative optimization as input. When a = 1, this binary tuple is the initial template implicit representation binary tuple, that is, the set of data containing the initial template text implicit representation and the initial control template core text item implicit representation.
[0116] During the iterative optimization process, the cloud computing server inputs the b-th template implicit representation binary tuple into the a-th core text index model. This model is, for example, a complex neural network structure that includes multiple hidden layers, activation functions, and loss functions, etc. Through mechanisms such as forward propagation and backward propagation, the model processes the input data and gradually adjusts its internal parameters (such as weights and biases) to minimize the loss function value and improve the model's fitting ability to text data.
[0117] After a new round of iterative optimization, the cloud computing server obtains the implicit representation binary tuple of the a-th template. This new binary tuple contains the implicit representation of the template text after the a-th round of iterative optimization and the implicit representation of the core text items of the control template. Compared with the implicit representation binary tuple of the b-th template, the representations in the implicit representation binary tuple of the a-th template are more accurate and effective because they have been optimized through further processing by the a-th core text index model. This process will be repeated continuously until the preset number of iterative rounds p is reached. In each round of iteration, the cloud computing server uses the model and optimization results obtained in the previous round of iteration to guide the optimization process of the current round. In this way, the core text index model can gradually learn the internal laws and feature representation methods of the text data, providing strong support for subsequent analysis of the operation logs of the power plant.
[0118] Step S200C: Determine the difference in implicit representation of templates between the implicit representation binary tuple of the b-th template and the implicit representation binary tuple of the a-th template according to the implicit representation error between the two.
[0119] In the cloud computing server, when the a-th round of iterative optimization is completed, a new set of implicit representation binary tuples of templates, namely the implicit representation binary tuple of the a-th template, is obtained. This binary tuple contains the implicit representation of the template text after optimization and the implicit representation of the core text items of the control template. To evaluate the effect of this round of iteration, the cloud computing server needs to compare the binary tuple of the current round (the a-th round) with the binary tuple of the previous round (the b-th round, where b = a - 1). This is the core task of step S200C.
[0120] Specifically, step S200C first calculates the implicit representation error between the implicit representation binary tuple of the b-th template and the implicit representation binary tuple of the a-th template. This error is usually measured by comparing the similarity or distance between the two sets of representations. In the vector space, distance is an intuitive and effective metric. Commonly used distance calculation methods include Euclidean distance, Manhattan distance, cosine similarity, etc.
[0121] Taking Euclidean distance as an example, assume that the implicit representation of the template text in the implicit representation binary tuple of the b-th template is vector and the implicit representation of the core text items of the control template is vector The corresponding vectors in the implicit representation binary tuple of the a-th template are respectively and Then the difference in implicit representation of templates between the implicit representation binary tuple of the b-th template and the implicit representation binary tuple of the a-th template (taking Euclidean distance as an example) can be calculated by the following formula:
[0122]
[0123] Among them, ∥·∥ 2 represents the calculation of the Euclidean distance, where d is the dimension of the vector, and respectively represent the value of the i-th element in the implicit representation vectors of the template text in the a-th round and the b-th round. Similarly, and represent the values of the corresponding elements in the implicit representation vector of the core text item of the control template.
[0124] In practical applications, the cloud computing server can calculate the difference between the implicit representation of the template text and the implicit representation of the core text item of the control template respectively, or merge them into a comprehensive difference measure (e.g., by weighted summation). This difference measure reflects the change in the model performance during the iterative optimization process. If the difference measure is small, it indicates that the output of the model changes little in two consecutive iterations, which may mean that the model is approaching the convergence state; if the difference measure is large, it indicates that there is still a large room for optimization of the model.
[0125] Through the calculation in step S200C, the cloud computing server can quantitatively evaluate the effect of each round of iterative optimization, providing important feedback information for the subsequent training process. This helps the cloud computing server to more intelligently adjust the training strategy and improve the accuracy and efficiency of the core text indexing model.
[0126] Step S200D: Fuse the difference measures of the implicit representations of the first a templates to obtain the debugging cost value of the a-th implicit representation;
[0127] The purpose of step S200D is to meaningfully integrate the difference measures of the implicit representations of the templates generated in each round of iteration (i.e., the output of step S200C), so as to obtain an index that can reflect the current performance of the model - the debugging cost value of the a-th implicit representation.
[0128] When the cloud computing server executes step S200D, it has already calculated the difference measures of the implicit representations of the templates generated in each round of the first a rounds of iteration through step S200C. These difference measures may include the difference measure of the implicit representation of the template text and the difference measure of the implicit representation of the core text item of the control template, which respectively measure the degree of change in the implicit representations of the same template text and the control core text item by the model in two consecutive iterations.
[0129] To obtain a comprehensive evaluation index, the cloud computing server fuses these difference measures. The fusion method can be simple arithmetic mean, weighted mean or other more complex aggregation methods. Here, weighted summation is used as an example for illustration.
[0130] First, assume that for the a-th round of iteration, the cloud computing server has calculated the set of difference measures of the implicit representations of the template text in the first a rounds of iteration:
[0131] {Differencetext,1,Differencetext,2,…,Differencetext,a};
[0132] and the set of difference amounts implicitly represented by the core text items of the control template:
[0133] {Differencecore,1,Differencecore,2,…,Differencecore,a}.
[0134] Then, the cloud computing server can assign a set of weights {α1, α2, …, αa} and {β1, β2, …, βa} to these two sets of difference amounts respectively. These weights can be set according to the actual situation. For example, they can be assigned based on the importance of the impact of the difference amount on the model performance. Next, the cloud computing server performs weighted summation on the difference amounts implicitly represented by the template text and the difference amounts implicitly represented by the core text items of the control template respectively to obtain two sets of weighted sums.
[0135] Finally, in order to obtain a unified implicit representation debugging cost value, the cloud computing server can further fuse the above two sets of weighted sums. The fusion method can be a simple arithmetic average or other aggregation functions designed according to actual needs. In this way, the cloud computing server obtains the a-th implicit representation debugging cost value, which reflects the overall performance change trend of the model in the first a rounds of iterations. By comparing the debugging cost values of different iteration rounds, the cloud computing server can evaluate the convergence situation of the model and then decide whether to continue iterative optimization or adjust the optimization strategy. In the intelligent cloud file storage scenario of a power plant based on cloud computing, the significance of this step is to provide a quantifiable monitoring index for model training. By real-time monitoring the change of the debugging cost value, the cloud computing server can timely discover and solve the problems that may occur in the training process, ensure that the model can converge to the optimal state stably and efficiently, and thus provide more accurate and reliable support for the operation log analysis of the power plant.
[0136] Step S200E: If the implicit representation debugging cost value determined according to all the difference amounts implicitly represented by the templates obtained after the first a rounds of iterative optimization does not meet the cost threshold, then update the weights and biases in the a-th core text index model to obtain the (a + 1)-th core text index model; complete the (a + 1)-th round of iterative optimization in the (a + 1)-th core text index model.
[0137] In a cloud computing server, as iterative optimization progresses, after each round of iteration is completed, the cloud computing server evaluates the current performance of the model based on the amount of difference in the implicit representation of the template calculated during the previous a rounds of iteration. These amounts of difference are calculated through step S200C and are fused into the implicit representation debugging cost value in step S200D. This cost value reflects the degree of optimization of the implicit representation of the template text and the core text items of the control template during the iterative process of the model. Next, in step S200E, the cloud computing server compares this implicit representation debugging cost value with a preset cost threshold. The cost threshold is a standard set according to actual requirements and is used to determine whether the model has reached a sufficient degree of optimization. If the implicit representation debugging cost value is greater than the cost threshold, it indicates that the current model performance has not reached the satisfactory standard and iterative optimization needs to continue.
[0138] At this time, the cloud computing server updates the weights and biases in the a-th core text index model according to the results of the previous a rounds of iteration. Weights and biases are key parameters in a neural network model, and they determine the way the model responds to input data and the accuracy of the output results. The process of updating these parameters usually relies on the backpropagation algorithm, which guides the adjustment direction of the parameters based on the difference (i.e., loss or error) between the predicted result and the actual result of the model. After updating the weights and biases, the cloud computing server obtains a new model version, namely the (a + 1)-th core text index model. This new model inherits the optimization results of the previous a rounds of iteration and is further adjusted and optimized on this basis. Subsequently, the cloud computing server continues to execute the iterative optimization process on the (a + 1)-th core text index model, that is, enters the (a + 1)-th round of iteration.
[0139] In the (a + 1)-th round of iteration, the cloud computing server repeats the previous steps, including processing the template text and the core text items of the control template using the new model, calculating the amount of difference in the implicit representation, evaluating the model performance, etc. As the iteration progresses, the performance of the model will gradually improve, and the implicit representation debugging cost value will gradually decrease. When the implicit representation debugging cost value finally meets the cost threshold, it indicates that the model has reached the expected degree of optimization, and at this time the iterative process will stop.
[0140] Step S200F: If the implicit representation debugging cost value determined based on all the amounts of difference in the implicit representation of the template obtained after the previous a rounds of iterative optimization meets the cost threshold, then determine that the a-th round of iterative optimization is the p-th round of iterative optimization.
[0141] Step S200F ensures that after a sufficient number of iterations, the model can meet the expected performance standards, thus avoiding unnecessary waste of computing resources. When the execution reaches step S200F in the cloud computing server, the first a rounds of iterative optimization have been completed, and based on the difference in the implicitly represented template calculated during these iterations, an implicitly represented debugging cost value has been fused through step S200E. This cost value reflects the performance of the model in the current iteration round.
[0142] Next, the cloud computing server will compare this implicitly represented debugging cost value with a preset cost threshold. The cost threshold is a standard set according to actual requirements and is used to evaluate whether the model has reached a satisfactory level of optimization. If the implicitly represented debugging cost value meets (i.e., is less than or equal to) the cost threshold, then it can be considered that the current model performance is good enough and no more iterative optimization is required.
[0143] At this time, step S200F is triggered, and the cloud computing server determines that the a - th round of iterative optimization is the last round of iterative optimization in the entire initialization process, that is, the p - th round of iterative optimization. Here, p is a preset maximum number of iterative rounds limit. However, in actual execution, if the model can meet the performance requirements before reaching p rounds, then the iteration can be ended in advance to save computing resources.
[0144] After determining the number of rounds of iterative optimization, the cloud computing server can save and use the currently optimized core text index model as the initialized model. This model has been optimized through multiple rounds of iteration and has the ability to extract key information from the power plant operation logs, providing strong support for subsequent log analysis and storage.
[0145] It should be noted that although step S200F determines the end of iterative optimization, the entire model optimization process does not completely depend on this step. In practical applications, the cloud computing server may also comprehensively judge whether to continue the iteration based on other factors (such as time limit, resource usage, etc.). In addition, with the continuous addition of new data and the change of the model application scenario, the initialized model may need to be further online - learned or updated regularly to adapt to new requirements.
[0146] In a possible implementation solution, step S200C, determining the difference in the implicitly represented template between the b - th implicitly represented template binary tuple and the a - th implicitly represented template binary tuple according to the implicit representation error between the b - th implicitly represented template binary tuple and the a - th implicitly represented template binary tuple, may include:
[0147] Step S200C1: Determine the difference quantity between the implicit representation of the b-th template text in the implicit representation binary tuple of the b-th template and the implicit representation of the a-th template text in the implicit representation binary tuple of the a-th template to obtain the implicit representation difference quantity of the b-th template;
[0148] Step S200C2: Determine the difference quantity between the implicit representation of the core text item of the b-th control template in the implicit representation binary tuple of the b-th template and the implicit representation of the core text item of the a-th control template in the implicit representation binary tuple of the a-th template to obtain the implicit representation difference quantity of the b-th control template.
[0149] During the iterative optimization process of the core text index model, the implicit representation of the template text is an internal encoding method of the model for the key information in the power plant operation log. As the iteration progresses, these implicit representations will gradually approach the true semantic features, thus more accurately reflecting the content of the log text. The purpose of Step S200C1 is to quantify this approximation process, that is, to calculate the difference quantity between the implicit representations of the template text in the b-th iteration and the a-th iteration.
[0150] Suppose the implicit representation of the template text obtained after the b-th iteration optimization is a vector where d is the dimension of the vector, representing the position of the text in the implicit space. Similarly, the implicit representation of the template text obtained after the a-th iteration optimization is a vector To quantify the difference between these two vectors, various distance measurement methods can be used, and the most commonly used is the Euclidean distance.
[0151] In some cases, one may be more concerned about the direction difference between vectors rather than the absolute distance. In this case, the cosine similarity can be used as the measurement criterion. In the power plant intelligent cloud archive storage cloud computing server, the implicit representation of the template text may represent a certain specific device state or operating parameter. For example, a log about the generator temperature may be encoded as an implicit representation vector, which contains information about temperature values, change trends, etc. in multiple dimensions. As the iteration progresses, this vector will be gradually optimized to more accurately reflect the key information in the log.
[0152] Suppose in the b-th iteration, the model encodes a log about the generator temperature being too high and obtains the implicit representation vector Then, in the a-th iteration, this vector is further optimized to By calculating the Euclidean distance between the two, it is possible to evaluate whether the model's understanding of this log has deepened between two iterations, that is, whether it has captured the key information in the log more accurately. If the calculated distance is small, it indicates that the implicit representations obtained in the two iterations are very close, and the model may be approaching the convergence state; conversely, if the distance is large, it means that there is still a large room for optimization of the model.
[0153] Similar to the implicit representation of the template text, the implicit representation of the control template core text item is a randomly initialized vector introduced by the model during the iterative optimization process to guide learning. As the iteration progresses, this vector will gradually approach the true implicit representation of the core text item, thus playing a key role in the training process. The purpose of step S200D2 is to quantify this approximation process, that is, to calculate the difference between the implicit representation of the control template core text item in the b-th iteration and that in the a-th iteration.
[0154] Similar to step S200D1, the same distance metric method can be used to calculate the difference between the implicit representations of the control template core text items.
[0155] In the power plant intelligent cloud archive storage cloud computing server, although the implicit representation of the control template core text item is initially randomly initialized, it will gradually approach the true core text item representation during the iterative optimization process. This approximation process helps the model to more accurately identify the key information segments (i.e., core text items) in the log. By calculating the Euclidean distance or other distance metric values between the two vectors, the convergence of the implicit representation of the control template core text item by the model during the iteration can be evaluated. If the distance gradually decreases and stabilizes, it indicates that the model has successfully guided the implicit representation of the control template core text item to approach the true core text item representation; conversely, if the distance fluctuates greatly or continues to increase, it may mean that the model encounters difficulties in the optimization process or needs to adjust the optimization strategy.
[0156] Step S200C monitors the learning progress of the model and evaluates the effect of iterative optimization by calculating the difference between the implicit representation of the template text and the implicit representation of the control template core text item during the iteration.
[0157] In a possible implementation, step S200D, which fuses the difference amounts of the first a template implicit representations to obtain the a-th implicit representation debugging cost value, may include:
[0158] Step S200D1: Fuse the difference amounts of the first a template implicit representations to obtain the template implicit representation debugging cost value, and fuse the difference amounts of the first a control template implicit representations to obtain the control implicit representation debugging cost value;
[0159] Step S200D2: Fuse the debugging cost value of the template implicit representation and the debugging cost value of the control implicit representation to obtain the a-th implicit representation debugging cost value.
[0160] During the iterative optimization process, each iteration generates a set of template implicit representation difference amounts and a set of control template implicit representation difference amounts (as described in step S200C). These difference amounts respectively reflect the approximation degree of the model to the template text implicit representation and the control template core text item implicit representation during the iteration. The task of step S200D1 is to fuse these difference amounts to obtain two independent debugging cost values: the debugging cost value of the template implicit representation and the debugging cost value of the control implicit representation.
[0161] Suppose that in the first a iterations, the template implicit representation difference amounts generated in each iteration are Dtemplate1, Dtemplate2,..., Dtemplatea respectively. These difference amounts can be obtained by calculating the distance (such as the Euclidean distance) between the template text implicit representation vectors. To fuse these difference amounts into a single debugging cost value, the cloud computing server can adopt various methods, such as arithmetic mean, weighted mean, geometric mean, etc.
[0162] Similar to the template implicit representation, the control template implicit representation also generates difference amounts in each iteration, denoted as D control1 , D control2 ,..., D controla . These difference amounts reflect the changes in the control template core text item implicit representation during the iteration. Similarly, the cloud computing server can adopt methods such as arithmetic mean to fuse these difference amounts into the debugging cost value of the control implicit representation.
[0163] In the initial stage of the iteration, due to the model not having fully learned, both the template implicit representation difference amount and the control implicit representation difference amount may be relatively large. As the iteration progresses, these difference amounts will gradually decrease, indicating that the model is gradually converging. By calculating the debugging cost value C templatea of the template implicit representation and the debugging cost value C controla of the control implicit representation, the cloud computing server can quantitatively evaluate the learning progress and effect of the model during the iteration.
[0164] After calculating the debugging cost value of the template implicit representation and the debugging cost value of the control implicit representation respectively, the task of step S200D2 is to further fuse these two cost values to obtain the a-th implicit representation debugging cost value Ca. This cost value comprehensively reflects the approximation degree of the model to the template text and the control template core text item implicit representation during the iteration, and is an important indicator for evaluating the overall performance of the model.
[0165] There are various methods for fusing the debugging cost value of the template implicit representation and the debugging cost value of the control implicit representation. Common ones include arithmetic mean, weighted mean, maximum value, minimum value, etc. Which fusion method to choose depends on the specific application scenario and evaluation requirements. Here, it is assumed that the template implicit representation and the control implicit representation are equally important when evaluating the model performance, so a simple arithmetic mean is used for fusion. However, in actual applications, if a certain type of implicit representation is considered more important in the evaluation, a higher weight can be assigned to it.
[0166] In the power plant intelligent cloud archive storage cloud computing server, the template implicit representation and the control implicit representation may have different importance when evaluating the model performance. For example, if the main task of the model is to accurately identify the operating state of the generator set (i.e., the template text implicit representation), then the debugging cost value of the template implicit representation can be given a higher weight during fusion.
[0167] Assume that the weighted average method is used for fusion, and the weight α is assigned to the debugging cost value of the template implicit representation, and the weight 1 - α is assigned to the debugging cost value of the control implicit representation. Then the debugging cost value C a of the a-th implicit representation becomes:
[0168] C a = αC templatea + (1 - α)C controla ;
[0169] Step S200E calculates the debugging cost value of the a-th implicit representation by fusing the difference amount of the template implicit representation and the difference amount of the control template implicit representation generated in the previous a rounds of iterations. This cost value comprehensively reflects the learning progress and effect of the model during the iteration process, providing an important basis for subsequent training decisions. In the actual application of the power plant intelligent cloud archive storage cloud computing server, by accurately calculating the debugging cost value of the implicit representation and timely adjusting the training strategy, the accuracy and efficiency of the model can be further improved, thus better supporting the operation and maintenance management work of the power plant.
[0170] In a possible implementation scheme, in step S200B, after completing the iteration optimization of the a-th core text index model in the a-th round of iteration, and after completing the iteration optimization of the b-th template implicit representation binary group obtained in the b-th round of iteration to obtain the a-th template implicit representation binary group, the method further includes:
[0171] Step S200B1: Input the implicit representation of the a-th control template core text item in the a-th template implicit representation binary group into the a-th dense network component to obtain the a-th core text item inference distribution paragraph.
[0172] Step S200B1 further processes the template implicit representation using a dense network component to predict the possible positions of the core text items in the text. In the cloud computing server, after the a-th round of iterative optimization is completed, a set of optimized template implicit representation pairs will be obtained, which includes the a-th template text implicit representation and the a-th control template core text item implicit representation. Step S200C1 focuses on processing the a-th control template core text item implicit representation in this pair. This implicit representation is a high-dimensional vector that captures the characteristics of the control template core text item in the implicit space. However, this vector itself does not directly provide the specific position information of the core text item in the original text. To obtain this information, the cloud computing server inputs this implicit representation into a dense network component. A dense network component (dense layer) is a fully connected neural network layer, where each neuron is connected to all neurons in the previous layer. In this scenario, the dense network component receives the a-th control template core text item implicit representation as input and extracts deeper features through a series of non-linear transformations (such as activation functions) and linear combinations.
[0173] After being processed by the dense network component, the output is no longer a simple vector, but an inference distribution paragraph representing the possible positions of the core text item in the text. This distribution paragraph may be given in the form of a probability distribution, where each position (such as a word or phrase in the text) corresponds to a probability value indicating the likelihood that this position is part of the core text item.
[0174] For example, assume that a running log of a power plant describes a fault situation of a generating unit, and "Generator temperature is too high, and it has automatically shut down" is the core text item. In Step S200C1, the cloud computing server will input the implicit representation corresponding to this core text item into the dense network component. After processing, the output is, for example, a probability distribution indicating that the phrases "Generator temperature is too high" and "has automatically shut down" have a high probability of being recognized as part of the core text item in the text.
[0175] The key to this process is that the dense network component can automatically extract the position features of key information from complex text data by learning the mapping relationship between the input implicit representation and the positions of the core text items. This ability is crucial for the cloud computing server of the intelligent cloud archive storage of the power plant because it allows the cloud computing server to quickly locate and extract the key content in the running log, providing strong support for subsequent tasks such as fault analysis and performance evaluation.
[0176] Step S200B2: Calculate the inference difference amount of the a-th core text item between the inference distribution paragraph of the a-th core text item and the prior distribution paragraph of the core text item in the u-th running log text template.
[0177] In the cloud computing server, after the processing in step S200B1, the inference distribution paragraph of the a-th core text item has been obtained, which is a probability distribution representing the possible positions of the core text item in the text. At the same time, for the u-th running log text template, the prior distribution paragraph of its core text item has been determined in advance based on domain knowledge or manual annotation, that is, the true or expected position of the core text item in the actual text.
[0178] The core task of step S200B2 is to compare the differences between these two distribution paragraphs to quantitatively evaluate the accuracy of the model prediction. To achieve this goal, the cloud computing server adopts a suitable metric to calculate the difference amount between the two distributions, that is, the inference difference amount of the a-th core text item. In practical applications, there are various metrics available for selection, such as Kullback-Leibler divergence (KL divergence), Jensen-Shannon divergence, cross-entropy, etc. Here, KL divergence is taken as an example for illustration because KL divergence is a commonly used method to measure the difference between two probability distributions P and Q, and is particularly suitable for evaluating the difference between the model prediction distribution and the true distribution.
[0179] The calculation formula of KL divergence is:
[0180]
[0181] where X is the set of all possible events or positions, P(x) is the probability of position x in the inference distribution paragraph of the a-th core text item, and Q(x) is the probability of position x in the prior distribution paragraph of the core text item in the u-th running log text template.
[0182] Suppose there is a running log text template describing a certain fault process of a generator set, and "generator bearing temperature is too high" is the core text item. Through step S200C1, the cloud computing server has given the inference distribution of the possible positions of this core text item in the text. At the same time, based on domain knowledge or expert annotation, it is known that in the actual text, the expression "generator bearing temperature is too high" does exist and is located at a specific position in the text.
[0183] In step S200B2, the cloud computing server will use KL divergence (or other suitable metrics) to calculate the difference amount between the inference distribution and the prior distribution. The smaller this difference amount is, the more accurate the model prediction is; conversely, the larger the difference amount is, the greater the deviation between the model prediction and the actual situation. By calculating this difference amount, the cloud computing server can monitor the prediction performance of the model for the positions of core text items in real time and make adjustments as needed during the iterative optimization process. This is of great significance for improving the accuracy and efficiency of power plant running log analysis.
[0184] Step S200B3: Fuse the inference difference amounts of the first a core text items to obtain the inference debugging cost value of the a-th core text item.
[0185] In the cloud computing server, through the calculation in Step S200B2, the inference difference amounts of each round in the first a rounds of iteration have been obtained. These difference amounts respectively quantify the differences between the position distributions of the core text items predicted by the model and the true or prior position distributions in each round of iteration. The task of Step S200C3 is to fuse these difference amounts to obtain a comprehensive evaluation index - the inference debugging cost value of the a-th core text item. The fusion process usually involves some form of aggregation or weighted average of the difference amounts. The specific way of aggregation depends on the nature of the difference amounts and the evaluation objectives. In practical applications, common aggregation methods include arithmetic mean, weighted average, geometric mean, etc. In the scenario of intelligent cloud file storage in a power plant, the significance of this step is to provide a quantifiable monitoring index for the iterative optimization process of the model. By comparing the inference debugging cost values of the core text items in different iteration rounds, the cloud computing server can evaluate the convergence trend and stability of the model's prediction of the core text item positions. If the debugging cost value gradually decreases and tends to be stable as the number of iteration rounds increases, it indicates that the prediction performance of the model is gradually improving and tending to be mature; on the contrary, if the debugging cost value fluctuates greatly or continues to increase, it may be necessary to adjust the model's training strategy or optimization algorithm. For example, assume that an operation log of a power plant describes the fault handling process of a certain generator set, which contains multiple key core text items (such as fault phenomena, handling measures, etc.). During the iterative optimization process of the model, the cloud computing server will gradually learn the position characteristics of these core text items in the text and calculate the inference difference amounts in each round of iteration through Step S200B2. Then, in Step S200B3, these difference amounts are fused into a debugging cost value to evaluate the prediction performance of the current model. If the debugging cost value is small, it indicates that the model can already accurately predict the positions of the core text items; if the debugging cost value is large, it indicates that the model still needs further learning and optimization.
[0186] Step S200B4: Fuse the a-th implicit representation debugging cost value and the a-th core text item inference debugging cost value to obtain the a-th target cost value.
[0187] In the cloud computing server, through the calculations in the foregoing steps, the a-th implicit representation debugging cost value and the a-th core text item inference debugging cost value have been obtained respectively. The implicit representation debugging cost value reflects the optimization degree of the model for the implicit representation of the text during the iteration process, while the core text item inference debugging cost value measures the accuracy of the model's prediction of the core text item position. These two cost values evaluate the performance of the model from different dimensions. The task of step S200B4 is to fuse these two cost values to obtain a more comprehensive evaluation metric - the a-th target cost value. The fusion process usually involves some form of aggregation or weighted average of the two cost values, and the specific aggregation method depends on the evaluation objective and actual requirements.
[0188] Taking weighted average as an example, assume that the a-th implicit representation debugging cost value is C implicita , and the a-th core text item inference debugging cost value is C positiona , and the weights assigned to these two cost values are α and 1 - α (the value range of α is between 0 and 1), then the a-th target cost value C targeta can be calculated through the following formula:
[0189]
[0190] In this formula, the setting of the weight α depends on the relative importance attached to the optimization of the implicit representation and the accuracy of the core text item position prediction in the evaluation objective. If it is considered that the optimization of the implicit representation is more important, α can be set to a larger value; conversely, if more importance is attached to the prediction accuracy of the core text item position, α can be set to a smaller value. In the scenario of intelligent cloud file storage in a power plant, the target cost value provides a comprehensive evaluation metric for the performance of the cloud computing server. By comparing the target cost values in different iteration rounds, the cloud computing server can intuitively understand the overall optimization trend of the model during the iteration process. If the target cost value gradually decreases as the iteration rounds increase, it indicates that the model is continuously improving in both the optimization of the implicit representation and the prediction of the core text item position; conversely, if the target cost value fluctuates greatly or does not decrease continuously, it may be necessary to analyze the reasons in depth and adjust the optimization strategy. For example, assume that an operation log of a power plant details the discovery, diagnosis, and handling process of a certain equipment failure. During the iterative optimization process of the core text index model, the cloud computing server gradually improves the optimization degree of the implicit representation of this log and the prediction accuracy of the core text item position. Through the calculation of step S200B4, the cloud computing server obtains the target cost value for each iteration round. If the target cost value shows an obvious downward trend as the iteration progresses, then it can be judged that the model is gradually approaching the optimal solution, providing more reliable technical support for the analysis of the operation logs of the power plant.
[0191] Step S200B5: If the a-th target cost value cannot meet the cost threshold, update the weights and biases in the a-th core text index model and the a-th dense network component to obtain the (a + 1)-th core text index model and the (a + 1)-th dense network component; complete the (a + 1)-th round of iterative optimization in the (a + 1)-th core text index model, and obtain the (a + 1)-th core text item inference distribution paragraph in the (a + 1)-th dense network component.
[0192] In the cloud computing server, after the calculation in step S200B4, the a-th target cost value has been obtained. This value comprehensively evaluates the performance of the model in two aspects: implicit representation optimization and core text item position prediction. Step S200B5 first compares this target cost value with a preset cost threshold. The cost threshold is a standard set according to actual requirements and is used to determine whether the performance of the current model has reached a satisfactory level. If the a-th target cost value is greater than the cost threshold (i.e., the performance of the model has not reached the preset standard), then the cloud computing server will decide to continue optimizing the model.
[0193] The specific process of optimization includes updating the weights and biases in the a-th core text index model and the a-th dense network component. These two components are key parts of the model, and their parameters determine the prediction ability and generalization performance of the model. Through the backpropagation algorithm (a commonly used neural network training method), the cloud computing server will adjust these parameters according to the difference between the prediction results and the actual results of the current model, expecting to obtain better performance in the next round of iteration.
[0194] After the parameter update is completed, the cloud computing server obtains the (a + 1)-th core text index model and the (a + 1)-th dense network component. These two new models inherit the optimization results of the previous round of iteration and are further adjusted on this basis. Subsequently, the cloud computing server will continue to execute the iterative optimization process in the (a + 1)-th core text index model, that is, enter the (a + 1)-th round of iteration. In this round of iteration, the (a + 1)-th dense network component will receive the new template implicit representation as input and generate the (a + 1)-th core text item inference distribution paragraph, that is, the possible positions of the predicted core text items in the text.
[0195] For example, assume that a certain abnormal condition of a generator set and its handling process are detailed in an operation log of a power plant. During the iterative optimization process of the core text index model, if the target cost value after the a-th iteration is still higher than the cost threshold, it indicates that the current model still has deficiencies in extracting key information and predicting the positions of core text items. At this time, the cloud computing server will initiate a new round of iterative optimization to improve the performance of the model by adjusting the parameters of the core text index model and the dense network component. In the (a + 1)-th iteration, the updated model will again attempt to extract key information from the operation log and predict the positions of core text items in order to obtain more accurate results.
[0196] This process will be continuously repeated until the target cost value is lower than the cost threshold. At this time, it can be considered that the performance of the model has reached the preset standard, and the iterative optimization can be stopped and the model can be applied to the actual cloud computing server for intelligent cloud file storage in the power plant.
[0197] Step S200B6: If the a-th target cost value meets the cost threshold, determine that the a-th iterative optimization is the p-th iterative optimization.
[0198] In the cloud computing server, after the decision-making process of step S200C5, if the a-th target cost value meets the preset cost threshold, it indicates that the performance of the current model in both implicit representation optimization and core text item position prediction has reached a satisfactory level, and no more iterative optimization is required. At this time, step S200C6 is triggered, and the cloud computing server will determine that the a-th iterative optimization is the last round of the entire iterative process, that is, the p-th iterative optimization. Here, p is a preset upper limit of the number of iterative rounds, but in the actual execution process, the model may not need to reach this upper limit to meet the performance requirements. By setting the cost threshold, the cloud computing server can terminate the iteration in advance when the model performance meets the standard, thereby saving computing resources and accelerating the training process.
[0199] For example, assume that a power plant is using a cloud computing-based intelligent cloud file storage cloud computing server to manage its operation logs. These logs contain a large amount of key data such as equipment status information, operation records, and possible fault reports. In order to quickly extract useful information from this massive amount of data, the power plant uses a core text index model to automatically identify and index key text items.
[0200] During the iterative optimization process of the model, the cloud computing server improves its performance by continuously adjusting the model parameters. As the iteration progresses, the degree of optimization of the implicit representation of the text by the model and the prediction accuracy of the core text item positions gradually increase, resulting in a continuous decrease in the target cost value. When the target cost value after the a-th iteration is first lower than the preset cost threshold, the cloud computing server determines that the model has reached the standard according to the decision logic of step S200C6, and thus determines that the a-th iteration is the final p-th iteration optimization.
[0201] At this time, the cloud computing server saves the parameters and status of the current model and uses them for subsequent text analysis and indexing tasks. Since the model has been fully trained and optimized, it can accurately and efficiently process the operation log data of the power plant, providing timely and accurate information support for the operation and maintenance personnel. At the same time, due to the timely termination of the iteration process, the cloud computing server also avoids unnecessary waste of computing resources, improving the overall efficiency and economic benefits.
[0202] Based on the same principle as the method shown in Figure 1 In this application embodiment, an intelligent cloud file storage device 10 for a power plant based on cloud computing is also provided, as shown in Figure 2 The device 10 includes:
[0203] A log text acquisition module 11, configured to acquire the operation log text of the operation status of a target power plant uploaded by at least one intelligent terminal device of the power plant to form a to-be-analyzed operation log text;
[0204] An implicit representation extraction module 12, configured to extract a target text implicit representation from the to-be-analyzed operation log text, where the to-be-analyzed operation log text includes content of interest;
[0205] A core feature determination module 13, configured to iteratively optimize the target text implicit representation and the implicit representation of a control core text item obtained by arbitrarily assigning values based on a core text index model, to obtain an implicit representation of a target core text item of the content of interest, where the implicit representation of the target core text item represents the implicit representation of a core text item detected in the content of interest included in the to-be-analyzed operation log text;
[0206] Among them, the core text index model is obtained by performing p rounds of iterative optimization on each running log text template based on the initialized core text index model. In the a-th core text index model after completing the a-th round of iterative optimization, the implicit representation binary tuple of the b-th template obtained in the b-th round of iterative optimization is iteratively optimized to obtain the implicit representation binary tuple of the a-th template, and the weights and biases in the a-th core text index model are updated according to the implicit representation error between the implicit representation binary tuple of the b-th template and the implicit representation binary tuple of the a-th template, and stop when meeting the preset stop requirements, where 1 ≤ a ≤ p, b = a - 1. The implicit representation binary tuple of the a-th template includes the implicit representation of the a-th template text and the implicit representation of the core text item of the a-th control template. The implicit representation of the a-th template text is obtained by performing a rounds of iterative optimization on the initial template text implicit representation extracted from the running log text template, and the implicit representation of the core text item of the a-th control template is obtained by performing a rounds of iterative optimization on the initial implicit representation of the core text item of the control template obtained by arbitrary assignment;
[0207] The log text storage module 14 is configured to determine, based on the implicit representation of the target core text item, the text paragraphs in the to-be-analyzed running log text where the core text item in the content of interest exists, mark the core text item to obtain a marked running log text, and store the marked running log text.
[0208] The above embodiments introduce the intelligent cloud file storage device 10 of a power plant based on cloud computing from the perspective of virtual modules. The following introduces a cloud computing server from the perspective of physical modules, which is specifically as follows: The embodiment of the present application provides a cloud computing server, as Figure 3 shown, the cloud computing server 100 includes: a processor 101 and a memory 103. Among them, the processor 101 and the memory 103 are connected, such as connected through a bus 102. Optionally, the cloud computing server 100 may further include a transceiver 104. It should be noted that in actual applications, the transceiver 104 is not limited to one, and the structure of the cloud computing server 100 does not constitute a limitation to the embodiment of the present application.
[0209] The cloud computing server provided by the embodiment of the present application. The cloud computing server in the embodiment of the present application includes: one or more processors; a memory; one or more computer programs, where one or more computer programs are stored in the memory and configured to be executed by one or more processors. When the one or more programs are executed by the processor, the intelligent cloud file storage method of the power plant based on cloud computing described above is implemented.
[0210] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a processor, the processor can execute the corresponding content in the foregoing method embodiment.
[0211] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially as indicated by the arrows. The foregoing is only part of the implementation manners of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. An intelligent cloud archive storage method for a power plant based on cloud computing, characterized in that: Applied to a cloud computing server, the method comprises: Acquire a power plant operation status operation log text uploaded by at least one intelligent terminal device of the target power plant to form an operation log text to be analyzed; Extracting an implicit representation of a target text from the to-be-analyzed operation log text, wherein the to-be-analyzed operation log text includes content of interest; Iteratively optimizing the target text implicit representation and the control core text item implicit representation obtained by arbitrary assignment based on the core text index model to obtain the target core text item implicit representation of the content of interest, wherein the target core text item implicit representation represents the implicit representation of the core text item detected and obtained in the content of interest contained in the operation log text to be analyzed; The core text index model is obtained by performing p rounds of iterative optimization on each running log text template according to the initialized core text index model, and in the ath core text index model that completes the ath round of iterative optimization, the implicit representation bigram of the bth template obtained by the bth round of iterative optimization is iteratively optimized to obtain the ath template implicit representation bigram, and the weights and biases in the ath core text index model are updated according to the implicit representation error between the bth template implicit representation bigram and the ath template implicit representation bigram, and the process stops when the preset stopping requirements are met, wherein 1≤a≤p, b=a-1, when a is equal to 1, the b-th template implicit representation tuple obtained by the b-th round of iterative optimization represents the initial template implicit representation tuple; the a-th template implicit representation tuple includes the a-th template text implicit representation and the a-th reference template core text item implicit representation, the a-th template text implicit representation is obtained by performing a round of iterative optimization based on the initial template text implicit representation extracted from the running log text template, and the a-th reference template core text item implicit representation is obtained by performing a round of iterative optimization based on the initial reference template core text item implicit representation obtained by arbitrary value assignment startup; Determine, based on the implicit representation of the target core text item, a text paragraph in which the core text item in the content of interest exists in the to-be-analyzed operation log text, mark the core text item, obtain a marked operation log text, and store the marked operation log text; The step of iteratively optimizing the implicit representation of the target text and the implicit representation of the control core text item obtained by any assignment based on the core text index model to obtain the implicit representation of the target core text item of the content of interest includes: The following process is completed in the core text index model: Loading the obtained x-th target text implicit representation and x-th comparison core text item implicit representation output by the x-th iterative optimization component in the core text index model into the x+1-th iterative optimization component in the core text index model, wherein the core text index model includes t iterative optimization components, 1≤x≤t-1; In the x+1th iterative optimization component, the xth target text implicit representation and the xth comparison core text item implicit representation are iteratively optimized to obtain the x+1th target text implicit representation and the x+1th comparison core text item implicit representation; When x+1=t, the x+1th reference core text item implicit representation is used as the target core text item implicit representation; When x+1<t, the obtained x+1th target text implicit representation and the x+1th control core text item implicit representation are loaded into the x+2th iterative optimization component to perform the x+2th iterative optimization.
2. The method according to claim 1, characterized in that In the x+1th iterative optimization component, the xth target text implicit representation and the xth comparison core text item implicit representation are iteratively optimized to obtain the x+1th target text implicit representation and the x+1th comparison core text item implicit representation, including: The following process is completed in the x+1th iterative optimization component: In the first iterative optimization module of the x+1th iterative optimization component, linear mapping is performed on the log text sub-implicit representation output by the xth iterative optimization component to obtain the first log text sub-implicit representation; linear mapping is performed on the control core text item sub-implicit representation output by the xth iterative optimization component to obtain the first control core text item sub-implicit representation; In the e-th iterative optimization module of the x+1-th iterative optimization component, the f-th log text sub-implicit representation and the f-th comparison core text item sub-implicit representation output by the f-th iterative optimization module of the x+1-th iterative optimization component are loaded into the e-th iterative optimization module in the x+1-th iterative optimization component, where 2≤e, f=e-1; In the e-th iterative optimization module, linear mapping is performed on the f-th log text sub-implicit representation to obtain a first mapping feature vector, and the first mapping feature vector is used as the e-th log text sub-implicit representation output by the e-th iterative optimization module; and linear mapping is performed on the f-th control core text item sub-implicit representation to obtain a second mapping feature vector; based on the first mapping feature vector and the second mapping feature vector, the e-th control core text item sub-implicit representation output by the e-th iterative optimization module is obtained; In the case that the e-th iterative optimization module is the last iterative optimization module in the x+1-th iterative optimization component, the e-th log text sub-implicit representation is used as the x+1-th target text implicit representation obtained by executing the x+1-th iterative optimization component, and the e-th control core text item sub-implicit representation is used as the x+1-th control core text item implicit representation obtained by executing the x+1-th iterative optimization component.
3. The method according to claim 1, characterized in that Before extracting the implicit representation of the target text from the to-be-analyzed operation log text, the method further includes an initialization step of a core text index model, including: Extract the initial template text implicit representation from the obtained u-th running log text template, and obtain the initial comparison template core text item implicit representation obtained by any assignment start, 1≤u≤q, q is the total number of templates; The following process is completed cyclically for the implicit representation of the initial template text and the implicit representation of the initial comparison template core text item, and the optimization stops when p rounds of iteration are reached: In the ath core text index model that has completed the ath round of iterative optimization, iterative optimization is completed on the bth template implicit representation bigram obtained by the bth round of iterative optimization to obtain the ath template implicit representation bigram, when a is equal to 1, the bth template implicit representation bigram obtained by the bth round of iterative optimization represents the initial template implicit representation bigram, and the initial template implicit representation bigram includes the initial template text implicit representation and the initial comparison template core text item implicit representation; Determining a template implicit representation difference amount between the b-th template implicit representation bigram and the a-th template implicit representation bigram according to an implicit representation error between the b-th template implicit representation bigram and the a-th template implicit representation bigram; The implicit representation differences of the first a templates are fused to obtain the ath implicit representation debugging cost value; If the implicit representation debugging cost value determined according to the implicit representation difference amounts of all templates obtained after the previous a rounds of iterative optimization cannot meet the cost threshold, then the weights and biases in the a-th core text index model are updated to obtain the a+1-th core text index model; Completing the a+1th round of iterative optimization in the a+1th core text index model; If the implicit representation debugging cost value determined based on all template implicit representation differences obtained after the previous a rounds of iterative optimization meets the cost threshold, the a-th round of iterative optimization is determined to be the p-th round of iterative optimization.
4. The method according to claim 3, characterized in that The step of determining the template implicit representation difference amount between the b-th template implicit representation binary and the a-th template implicit representation binary according to the implicit representation error between the b-th template implicit representation binary and the a-th template implicit representation binary comprises: Determine the difference between the implicit representation of the b-th template text in the b-th template implicit representation binary and the implicit representation of the a-th template text in the a-th template implicit representation binary to obtain the difference of the b-th template implicit representation; The difference between the implicit representation of the bth control template core text item in the bth template implicit representation tuple and the implicit representation of the ath control template core text item in the ath template implicit representation tuple is determined to obtain the difference of the bth control template implicit representation.
5. The method according to claim 4, characterized in that The step of fusing the first a template implicit representation differences to obtain the ath implicit representation debugging cost value includes: The previous a template implicit representation differences are merged to obtain the template implicit representation debugging cost value, and the previous a reference template implicit representation differences are merged to obtain the reference implicit representation debugging cost value; The template implicit representation debugging cost value and the reference implicit representation debugging cost value are merged to obtain the ath implicit representation debugging cost value.
6. The method according to claim 3, characterized in that In the ath core text index model that has completed the ath round of iterative optimization, after completing iterative optimization on the bth template implicit representation bigram obtained by the bth round of iterative optimization and obtaining the ath template implicit representation bigram, the method further includes: Input the implicit representation of the ath control template core text item in the ath template implicit representation tuple into the ath dense network component to obtain the ath core text item inference distribution paragraph; Calculating the inference distribution paragraph of the a-th core text item and the core text item prior distribution paragraph in the u-th running log text template to obtain the inference difference amount of the a-th core text item; The inference differences of the first a core text items are integrated to obtain the inference debugging cost value of the a-th core text item; The ath implicit representation debugging cost value and the ath core text item reasoning debugging cost value are merged to obtain the ath target cost value; If the a-th target cost value cannot meet the cost threshold, the weights and biases in the a-th core text index model and the a-th dense network component are updated to obtain the a+1-th core text index model and the a+1-th dense network component; the a+1-th round of iterative optimization is completed in the a+1-th core text index model, and the a+1-th core text item reasoning distribution paragraph is obtained in the a+1-th dense network component; If the a-th target cost value satisfies the cost threshold, the a-th round of iterative optimization is determined to be the p-th round of iterative optimization.
7. An intelligent cloud archive storage device for a power plant based on cloud computing, characterized in that: The intelligent cloud archive storage device of the power plant based on cloud computing includes: A log text acquisition module, used to acquire the operation log text of the power plant operation status uploaded by at least one intelligent terminal device of the target power plant to form the operation log text to be analyzed; An implicit representation extraction module, used to extract an implicit representation of a target text from the operation log text to be analyzed, wherein the operation log text to be analyzed includes content of interest; A core feature determination module, for iteratively optimizing the implicit representation of the target text and the implicit representation of the control core text item obtained by any assignment start based on the core text index model, to obtain the implicit representation of the target core text item of the content of interest, wherein the implicit representation of the target core text item represents the implicit representation of the core text item detected and obtained in the content of interest contained in the operation log text to be analyzed; The core text index model is obtained by performing p rounds of iterative optimization on each running log text template according to the initialized core text index model, and in the ath core text index model that completes the ath round of iterative optimization, the implicit representation bigram of the bth template obtained by the bth round of iterative optimization is iteratively optimized to obtain the ath template implicit representation bigram, and the weights and biases in the ath core text index model are updated according to the implicit representation error between the bth template implicit representation bigram and the ath template implicit representation bigram, and the stop is stopped when the preset stop requirements are met, wherein 1≤a≤p , b=a-1, when a is equal to 1, the b-th template implicit representation tuple obtained by the b-th round of iterative optimization represents the initial template implicit representation tuple; the a-th template implicit representation tuple includes the a-th template text implicit representation and the a-th reference template core text item implicit representation, the a-th template text implicit representation is obtained by performing a round of iterative optimization based on the initial template text implicit representation extracted from the running log text template, and the a-th reference template core text item implicit representation is obtained by performing a round of iterative optimization based on the initial reference template core text item implicit representation obtained by arbitrary assignment startup The method comprises the following steps: performing iterative optimization on the implicit representation of the target text and the implicit representation of the control core text item obtained by any assignment based on the core text index model to obtain the implicit representation of the target core text item of the content of interest, including: completing the following process in the core text index model: loading the xth implicit representation of the target text and the xth implicit representation of the control core text item output by the xth iterative optimization component in the core text index model into the x+1th iterative optimization component in the core text index model, wherein the core text index model includes t iterative optimization components. , 1≤x≤t-1; in the x+1th iterative optimization component, the xth target text implicit representation and the xth comparison core text item implicit representation are iteratively optimized to obtain the x+1th target text implicit representation and the x+1th comparison core text item implicit representation; when x+1=t, the x+1th comparison core text item implicit representation is used as the target core text item implicit representation; when x+1<t, the obtained x+1th target text implicit representation and the x+1th comparison core text item implicit representation are loaded into the x+2th iterative optimization component to perform the x+2th iterative optimization; The log text storage module is used to determine the text paragraph where the core text item in the content of interest exists in the run log text to be analyzed based on the implicit representation of the target core text item, mark the core text item, obtain the marked run log text, and store the marked run log text.
8. A cloud computing server, characterized in that: include: one or more processors; Memory; one or more computer programs; The one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program runs on a processor, the processor executes the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text generation method and device, model training method and device, electronic equipment and medium
CN114492456A
Log anomaly detection method and device, equipment and storage medium
CN116107834A