Power plant equipment fault intelligent diagnosis method
Through timestamp alignment and device identification association processing, the source data regularity score is calculated, high-quality data entries are screened, a structured diagnostic data set is established, and a distributed diagnostic data index is generated. This solves the problems of fault diagnosis efficiency and storage resource waste in power plant equipment in the existing technology, and realizes efficient diagnosis of power plant equipment faults.
Patent Information
- Application Number
- CN202511127221.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies ignore the inherent correlation between data age and equipment importance, resulting in wasted storage resources and slow response to diagnostic queries. It is difficult to reasonably set the data retrieval order, affecting the efficiency and accuracy of power plant equipment fault diagnosis.
Through timestamp alignment and device identification association processing, the source data regularity score is calculated, high-quality data entries are screened, and a structured diagnostic data set is established. A distributed diagnostic data index is generated by combining data age and device importance. The data fragment list is dynamically sorted using active alarms and associated device lists to generate a diagnostic data retrieval path.
It improves the accuracy and real-time performance of power plant equipment fault diagnosis, optimizes storage resource allocation, improves data retrieval efficiency and query accuracy, and ensures the efficient and safe operation of power plant equipment.
Smart Images

Figure CN120632175A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and in particular to an intelligent diagnosis method for power plant equipment faults. Background Art
[0002] The intelligent diagnosis method for power plant equipment failures is used to help power plant operation and maintenance personnel efficiently discover, locate and diagnose abnormal conditions and potential failures of power plant equipment. By intelligently analyzing power plant equipment sensor monitoring data, operation logs, alarm records and other data, a diagnostic view of equipment failures is quickly generated to reduce equipment failure rates and improve equipment operation reliability and safety.
[0003] Existing technologies often ignore the inherent relationship between data age and device importance. Most device data fragments are stored using a simple balanced distribution method, resulting in wasted storage resources or excessive access pressure on hotspot storage nodes, affecting system stability. When processing diagnostic query requests, existing technologies often rely solely on basic matching of device identifiers and time ranges, lacking comprehensive analysis of active alarms and associated devices. This makes it difficult to rationally set the retrieval order for data fragments, resulting in slow diagnostic data response and insufficient accuracy, further delaying fault diagnosis and repair. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose an intelligent diagnosis method for power plant equipment faults.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for intelligent diagnosis of power plant equipment faults, comprising the following steps: Collect power plant operating status log records and alarm records, perform timestamp alignment and device identification association processing, calculate the source data regularity score based on the data source integrity and time series continuity, and based on the source data regularity score, screen data entries with source data regularity scores that meet the standards and perform unified formatting and integration processing to establish a standardized structured diagnostic data set; Based on the structured diagnostic data set, analyzing the time attributes of the data records, the associated power plant equipment identifiers, and the access frequency statistics, combining the data age and the equipment importance to derive a data segment distribution fitness value, assigning each data segment to a target storage node or logical partition based on the data segment distribution fitness value, creating an index entry including storage location information, and generating a distributed diagnostic data index; receiving a diagnostic query request for a target power plant equipment fault, parsing the equipment identifier and time range, searching for a location data segment using the distributed diagnostic data index, calculating the query relevance of the segment in combination with the currently active alarm and the associated equipment list to obtain a contextual retrieval priority, sorting the list of data segments to be retrieved based on the contextual retrieval priority, and obtaining a diagnostic data retrieval path; According to the order and location information determined in the diagnostic data retrieval path, data fragments are retrieved from each specified storage node or logical partition, the returned data is verified against the expected data items defined by the path, and a retrieval task completion metric is calculated. Based on the retrieval task completion metric, the verified data is aggregated, the information is assembled according to the query request context, and an associated fault diagnosis view is generated.
[0006] Preferably, the steps for obtaining the source data regularity score are: Collect power plant operation status log records and alarm records, unify data source labels, unify sampling time bases, and linearly merge record time axes for power plant operation status log records and alarm records corresponding to the same equipment identifier to generate an equipment record fusion sequence; Based on the device record fusion sequence, the total number of fields, the number of null value fields, the sampling time interval, the frequency of field value changes, the number of abnormal alarm triggers and the cross-sequence data synchronization delay time are extracted one by one to obtain a six-dimensional parameter set of regularity; The source data regularity score is calculated based on the six-dimensional regularity parameter set.
[0007] Preferably, the steps of obtaining the standardized structured diagnostic data set are: Collect the source data regularity score, use the device type and the source data regularity score as screening parameters, determine and retain data entries whose source data regularity scores meet the preset standard, mark the status of each record entry that meets the standard, and generate a set of data entries whose scores meet the standard; Based on the set of data entries that meet the scoring criteria, the data field structures of the original data entries are called one by one, the field data types and field lengths are standardized, and the field positions and contents are rearranged according to the unified field sorting rules to form a set of data entries with a unified field format; According to the data entry set with a unified field format, field mapping of structured data is performed with device identification, timestamp and record category as associated primary keys to form a standardized structured diagnostic data set.
[0008] Preferably, the steps for obtaining the data segment distribution suitability value are: Based on the structured diagnostic data set, extract the timestamp information and power plant equipment identification number corresponding to each data record, form an equipment record time series matrix through continuous time window aggregation and equipment record tracking, and generate a time equipment index matrix; According to the time device index matrix, the number of record accesses, update frequency, and active period span of each device number within a fixed statistical period are counted, and the time difference between the data generation time corresponding to each record and the current time is calculated to obtain the access behavior feature set; Based on the access behavior feature set, a data segment distribution suitability value is calculated.
[0009] Preferably, the steps of obtaining the distributed diagnostic data index are: Based on the data segment distribution fitness value, setting a fitness threshold for each data segment, judging whether the data segment distribution fitness value of each data segment meets the fitness threshold standard, marking the data segments that meet the standard and pre-sorting them one by one by position allocation, and generating a set of data segments with pre-sorted position allocation; Allocating a pre-sorted set of data segments according to the position, calling the power plant equipment identifier and the corresponding data segment distribution suitability value of each data segment one by one, and assigning the data segments one by one to target storage nodes or logical partitions that meet capacity and access performance requirements in descending order of the data segment distribution suitability values, thereby forming a set of data segments that have completed storage node or logical partition assignment; Based on the set of data fragments that have completed the assignment of storage nodes or logical partitions, the location information of the target storage node or logical partition where each data fragment is located and the associated power plant equipment identification are recorded one by one, index entries are created in chronological order, and a distributed diagnostic data index is generated.
[0010] Preferably, the steps of obtaining the diagnostic data retrieval path are: Receive a target power plant equipment fault diagnosis query request, extract a device identifier field and a query time range field, use the device identifier field to match the device index primary key in the distributed diagnostic data index, and use the query time range field to perform interval matching in the time period index key to form a set of positioning data segments; Calculating a context retrieval priority of each data segment based on the set of positioning data segments; According to the context retrieval priority, all positioning data segments are sorted from high to low according to the context retrieval priority, and each data segment is arranged in turn into the list of data segments to be retrieved, and is dispatched to the retrieval task execution plan queue in order to generate a diagnostic data retrieval path.
[0011] Preferably, the steps for obtaining the retrieval task completion metric are: According to the search order and target logical partition location information set in the diagnostic data search path, the path address and storage node number of each data segment in the task are retrieved one by one, and cross-partition parallel data is pulled to generate a list of returned data segments; According to the returned data fragment list, the number of record entries, the number of accurate field comparison items, the number of missing field bits, the number of field bits that failed verification, the number of structural position misalignments, and the number of abnormal label fields of each returned data fragment are called to form a fragment field verification parameter set; A retrieval task completion metric is calculated based on the segment field validation parameter set.
[0012] Preferably, the steps of obtaining the associated fault diagnosis view are: Based on the retrieval task completion metric, the data segments whose retrieval task completion metric meets the preset qualification standard are compared one by one from the returned data segment list, and all data segments that pass the comparison are retained to form a set of data segments that pass the verification; Based on the verified data segment set, the power plant equipment identifier, query time range, current active alarms, and associated equipment list in the target power plant equipment fault diagnosis query request are called, and the aggregated data segment content is matched with the query request context one by one to obtain a data segment set with complete context association; According to the data segment set with context association completed, all associated fields in the data segment are called one by one, and information is assembled and formatted and output according to the device identifier and time range specified in the query request context to generate an associated fault diagnosis view.
[0013] Compared with the prior art, the advantages and positive effects of the present invention are: The present invention enhances the accuracy of the association between power plant operation logs and alarm records by introducing timestamp alignment and device identification association processing, ensuring the comprehensiveness and temporal continuity of data sources, and completing data quality quantitative control through source data regularity scoring. High-quality data entries are used as the basis for diagnostic data standardization, thereby improving the reliability of structured diagnostic data sets. In addition, the present invention dynamically generates data segment distribution fitness values based on data age and device importance factors, so that data segments can be efficiently distributed on storage nodes or logical partitions based on actual access conditions and data lifecycles, thereby achieving resource optimization configuration. Furthermore, the present invention further utilizes information from the current active alarm and associated device list to dynamically sort the data segment list by calculating the context retrieval priority, making the diagnostic data retrieval path more in line with real-time query requirements and improving data retrieval efficiency and query accuracy. At the same time, the integrity of the acquired data segments is measured and quantified through retrieval tasks, and the verified data is aggregated to construct a fault diagnosis view for the query request context, thereby improving the accuracy and real-time performance of abnormal diagnosis of power plant equipment, thereby ensuring the efficient and safe operation of power plant equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Schematic diagram of the steps of the present invention. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0016] See also Figure 1 The present invention provides a technical solution, a method for intelligent diagnosis of power plant equipment faults, comprising the following steps: Collect power plant operating status log records and alarm records, perform timestamp alignment and device identification association processing, calculate the source data regularity score based on the data source integrity and time series continuity, and based on the source data regularity score, screen data entries that meet the source data regularity score standards and perform unified formatting and integration processing to establish a standardized structured diagnostic data set; Based on the structured diagnostic data set, the time attributes of the data records, the associated power plant equipment identification and access frequency statistics are analyzed. The data age and equipment importance are combined to derive the data segment distribution fitness value. Based on the data segment distribution fitness value, each data segment is assigned to the target storage node or logical partition, and an index entry including storage location information is created to generate a distributed diagnostic data index. Receive diagnostic query requests for target power plant equipment faults, parse the equipment identifier and time range, use the distributed diagnostic data index to find the location data segment, calculate the query relevance of the segment based on the current active alarm and the list of associated devices to obtain the context retrieval priority, sort the list of data segments to be retrieved based on the context retrieval priority, and obtain the diagnostic data retrieval path; According to the order and location information determined in the diagnostic data retrieval path, data fragments are retrieved from the specified storage nodes or logical partitions, the returned data is verified against the expected data items defined by the path, and the retrieval task completion metric is calculated. Based on the retrieval task completion metric, the verified data is aggregated, the information is assembled according to the query request context, and an associated fault diagnosis view is generated.
[0017] The steps for obtaining the source data regularity score are: Collect power plant operation status log records and alarm records, unify data source labels, unify sampling time bases, and linearly merge record time axes for power plant operation status log records and alarm records corresponding to the same equipment identifier to generate an equipment record fusion sequence; Based on the fusion sequence of device records, the total number of fields, the number of empty value fields, the sampling time interval, the frequency of field value changes, the number of abnormal alarm triggers, and the cross-sequence data synchronization delay time are extracted one by one to obtain a six-dimensional parameter set of regularity; According to the six-dimensional parameter set of regularity, the regularity score of the source data is calculated using the following formula: ; in, represents the source data regularity score, Indicates the average value of the total number of fields, Indicates the average frequency of field value changes in the last 30 days. Indicates the maximum number of empty value fields. Indicates the sum of the number of abnormal alarm triggering times. represents the standard deviation of the sampling time interval, Indicates the maximum difference in sampling time intervals, Indicates the maximum value of the cross-sequence data synchronization delay time in the running status log. Indicates the minimum value of the cross-sequence data synchronization delay time in the alarm record.
[0018] Specifically, the operation status log records and alarm records of the power plant are collected, and the label consistency processing is performed on the operation status logs and alarm data corresponding to the same equipment identification. The obtained log records are matched according to the equipment identification number and the equipment number and recording time of each record are checked for duplication or omission. If multiple alarm entries are found under the same timestamp, the subsequent alarm entries are appended to the back of the same timestamp record in a fixed order. For entries with time span differences, the precise time difference is calibrated using the timing comparison method. Then, in the processing program, the sampling time of each record is offset by microseconds according to the same reference time. The entries with excessive offset in the record content are deeply checked to check whether the entry has any difference in the upstream data collection link. If it is valid, it will be updated according to the offset value after precise alignment. If it is invalid, it will be directly deleted. The time coordinate alignment operation of all records is completed by traversing. After the alignment is completed, the field contents from different sources are merged to form a joint structure under the same time series. The field names of repeated fields under the same timestamp are checked. If the field names are found to be consistent, their values are judged to be conflicting. If there is no conflict, they are retained and written at the established position. If there is a conflict, the item with a more complete field value range and higher quality is selected to cover the conflicting item. After all records are merged, an index list corresponding to the device identification and timestamp is established for subsequent query. In this way, information from the operation status log and the alarm record can be aggregated on the same timeline and field confusion can be avoided, and finally a fusion sequence of device records is obtained.
[0019] Based on the fusion sequence of device records, the total number of fields, the number of empty value fields, the sampling time interval, the frequency of field value changes, the number of abnormal alarm triggering and the cross-sequence data synchronization delay time are extracted one by one. The statistical function is called in the processing program to calculate the number of fields contained in each fusion record and obtain the number of corresponding empty value fields in all fields. Then, the sampling time interval is counted according to the timestamp difference between the previous and next records. The difference between adjacent records of the same field is recorded as a field value change. The frequency of field value changes is obtained by superimposing the number of changes of all fields in time series. The number of abnormal alarm triggering is directly obtained from the alarm type field in the fusion record. The number of segment statistics, and the acquisition of cross-sequence data synchronization delay time requires comparing the arrival order of the running status log and alarm data of the record at the same timestamp, extracting the time difference between the two and selecting the maximum and minimum delay values actually corresponding to each record in the subsequent statistical process. After the above extraction operation is completed, the extracted data of all records can be searched according to the device identification and timestamp index, and the six-dimensional parameter set of regularity is generated one by one. The parameter set includes the total number of fields, the number of empty value fields, the sampling time interval, the frequency of field value changes, the number of abnormal alarm triggering and the cross-sequence data synchronization delay time, which can provide unified statistical parameters for subsequent steps.
[0020] formula: The benefit of the formula is that by incorporating factors such as the number of fields, field value changes, the number of alarm triggers, and the synchronization delay between multi-source logs, it can fully reflect the regularity of the data in multiple dimensions. This comprehensive consideration is more comprehensive in actual power plant equipment fault diagnosis scenarios, because relying solely on field completeness or a single sampling interval cannot accurately measure data quality. This formula uses multiple key factors and uses a nonlinear structure for superimposed calculations to make a more balanced comprehensive score for data in terms of completeness, sparsity, change activity, and synchronization. In subsequent fault diagnosis or analysis, a higher score indicates more reliable source data, which helps to make subsequent diagnostic results more accurate.
[0021] The parameter is obtained by checking the field list of each record in the fused sequence and recording the number of fields. The field number of all records is then summed up and divided by the total number of fused records to obtain the average number of fields. During the acquisition of this parameter, the number of fields in each fused record is strictly counted one by one. For example, when 500 fused records are collected, the total number of fields is accumulated to 7500. This value is then ratioed to the number of fused records. The specific formula is: , after calculation, we get is 15.0.
[0022] The parameter acquisition steps are as follows: retrieve the null value fields of all records in the fusion sequence and count the number of null values in each record, then find the maximum value among the null value numbers corresponding to all records. If all fields in some records are null, it is counted as the upper limit of the number of null value fields. The number of null value fields in each record is strictly read from the field list and each field is judged to be null. Finally, the maximum value among all records is selected as In an actual statistics, it was found that the number of fields with the highest null value reached 3, so The value is 3.
[0023] The steps to obtain the parameters are to check the timestamp differences of adjacent records in the fused sequence, regard these timestamp differences as a sampling interval list, and then calculate the standard deviation of this set of sampling interval lists according to the mathematical formula. The specific calculation formula is: ,in Indicates the The time interval between adjacent records, Represents the mean of all time intervals, n is the number of adjacent records, and there are 600 adjacent record intervals in this statistics. is 5.2 seconds, and the standard deviation is 1.1 obtained by variance calculation, so The value is 1.1.
[0024] The parameter acquisition step is to further search the sampling interval list from the same batch of records, find the interval value with the largest value and the interval value with the smallest value in the list and calculate the difference between the two. For example, in the interval statistics of adjacent records, the largest interval value is 8.7 seconds and the smallest interval value is 2.2 seconds, then calculate , considering this difference as .
[0025] The steps to obtain the parameters are to count the changes in field values, sum the number of times each field in the fused record has content differences between adjacent records, and then divide the total value by the number of days in the statistical period. In an actual application, the statistical period is 30 days, and the total number of changes reaches 150. Then, the formula , calculated is 5.
[0026] The parameter acquisition step is to directly count the number of all abnormal alarm entries in the fusion sequence and add up the number of these alarms. For each record with an abnormal alarm identifier, the tag value can be read in the alarm type field. For example, if 10 abnormal alarms accumulate in the same statistical period, then The value is 10.
[0027] The parameter acquisition step is to obtain the recording time corresponding to the running status log and the arrival time of the log in the fusion process in each record, take the difference as the cross-sequence data synchronization delay, and then select the maximum value of the delay from all records as For example, in the process of detecting 1200 logs, it was found that the arrival time differed from the recorded time by 3 seconds once, and the rest were mostly within 2 seconds. Therefore, the maximum value was determined to be 3, that is, =3.
[0028] The parameter acquisition step is to compare the timestamp of the alarm record with the timestamp of the log synchronization completion after fusion processing, calculate the cross-sequence delay time list, and select the minimum value among all the results to be defined as In an actual test, the minimum value reached 0.5 seconds, and after repeated inspection, it was confirmed that the value was valid. The value is 0.5.
[0029] Calculation process: The first step is to calculate the absolute difference , according to the above parameters =15.0, =3, so ,but .
[0030] The second step is to calculate ,Depend on =1.1, then .
[0031] The third step is to multiply the above two terms, that is, , and then take the cube, we get .
[0032] Step 4: Calculate ,Depend on =5, =10, =6.5, first find , then the molecule , the denominator , and then the fractional result , after taking the absolute value, it is still 4.975.
[0033] The fifth step is to add the results of the third and fourth steps, that is .
[0034] Step 6: Calculate the denominator ,Depend on =3, =0.5, difference , add 1 to get 3.5, and then take the square root .
[0035] Finally, divide the numerator by the denominator and take the square root: ; This result shows that when The larger the value, the better the regularity of the source data. The value is approximately 62.0. This value is closely related to the comprehensive performance of the source data in terms of field completeness, field change activity, and multi-source synchronization. A value above 50 indicates good data regularity. If the value is below 20, it indicates that the data is seriously inconsistent in multiple dimensions. Therefore, the calculation result can be used to determine whether the quality of the current fused data meets the requirements of subsequent diagnostic steps.
[0036] The steps to obtain a standardized structured diagnostic data set are: Collect source data regularity scores, use device type and source data regularity scores as screening parameters, determine and retain data entries whose source data regularity scores meet the preset standards, mark the status of record entries that meet the standards one by one, and generate a set of data entries whose scores meet the standards; Based on the set of data entries that meet the scoring standards, the data field structure of the original data entries is called one by one, the field data type and field length standards are standardized, and the field positions and contents are rearranged according to the unified field sorting rules to form a set of data entries with a unified field format; According to the data entry set with unified field format, the field mapping of structured data is performed with the device identification, timestamp and record category as the associated primary key to form a standardized structured diagnostic data set.
[0037] Specifically, the source data regularity scores are collected, and the score distribution range obtained by batch aggregation of source data regularity scores in the same power plant scenario in the early stage is referred to. 500 score records are read and the maximum, minimum and average values of these records are counted. A more stringent screening threshold is set downward for the average value and the threshold is recorded as , by traversing the specific value of each scoring record and Compare, if a score is higher than The record is marked as an available data entry and an additional identification tag is generated for the corresponding device type. It will be recorded as temporarily unqualified data and subsequent investigation will be carried out. According to actual monitoring, the average score is 68, the highest score is 85 and the lowest score is 42. The threshold is set to 60 and the calculation process confirms that the threshold can screen out about 70% of the data entries, complete the automatic comparison and marking of the scoring records, and after the comparison is completed, all those that meet the threshold above 60 will be marked. The records are grouped into a unified set, and then the device identification and data source information of each item that meets the conditions are recorded. Finally, a unique identification number is assigned to the set and stored in a data structure called "scoring result index" to generate a set of data items whose scores meet the standards.
[0038] Based on the set of data entries that meet the scoring criteria, a pre-sorted field specification list is called and matched with the corresponding types according to the field names listed therein. This list defines, for example, that the temperature field corresponds to the float type, ranging from 0 to 200, retaining two decimal places, and a maximum character length of 6; the pressure field corresponds to the float type, ranging from 0 to 10, retaining three decimal places, and a maximum character length of 6; the device number field corresponds to the string type, and the maximum character length is 20; the timestamp field corresponds to the string type, and the format is fixed to YYYY-MM-DDHH:MM:SS, and the length is fixed to 19. Then, the data entries that meet the scoring criteria are retrieved one by one. The field names are checked against this list to confirm whether the data types are consistent. If the field types do not match, force conversion or discard the field to keep the format consistent. Then, refer to the field sorting rules to put the basic information field in the front position, and then arrange the measurement data field, alarm data field, and auxiliary description field in sequence. During this period, if there are large differences in field lengths, the excess characters are truncated within the acceptable range and an exception count value is appended to the record to indicate the number of times the field length exception occurs in this entry. Finally, all fields are sorted and merged into a unified structure and checked with the device number and timestamp as indexes to form a set of data entries with a unified field format.
[0039] According to a set of data entries with a unified field format, the device identification string of each record is read and combined with the timestamp field and the record category field as a composite primary key. The actual device number corresponding to each device identification is found by comparing the existing device dictionary mapping table and the valid range of the timestamp and category fields is confirmed. If the corresponding device cannot be retrieved in the existing dictionary mapping table, it is added to a newly created temporary device list for subsequent updates. Subsequently, within the scope of a single record entry, all fields are filled into the data container bound to the primary key according to the field name sequence under the unified structure, and their source, data type and length information are marked respectively. If it is found that some fields conflict with the device identification or timestamp, the conflict position is additionally recorded and tracked using an additional index record. Finally, after all records are mapped, the merged structured data is sequentially written into the database index of the diagnostic data set according to the primary key to form a standardized structured diagnostic data set.
[0040] The steps to obtain the data segment distribution fitness value are as follows: Based on the structured diagnostic data set, the timestamp information and power plant equipment identification number corresponding to each data record are extracted. Through continuous time window aggregation and equipment record tracking, the equipment record time series matrix is formed, and the time equipment index matrix is generated. According to the time device index matrix, the number of record accesses, update frequency, and active period span of each device number within a fixed statistical period are counted. The time difference between the data generation time corresponding to each record and the current time is calculated to obtain the access behavior feature set; Based on the access behavior feature set, the data segment distribution suitability value is calculated using the following formula: ; in, represents the data segment distribution fitness value, is the number of visits to the current power plant equipment number within the statistical period, The average time difference generated for the data record entries of the device within the statistical period. The update frequency of all records under the current device number. The time difference between the earliest timestamp recorded under the current device number and the current time. is the active time span corresponding to the device number in the time device index matrix, The maximum record change frequency for the current device ID in the past 30 days.
[0041] Specifically, based on the structured diagnostic data set obtained previously, the timestamp and power plant equipment identification number attached to each record are read, the time zone and accuracy of the timestamp field are proofread in a unified format, and millisecond-level numerical conversion is performed. Then, a fixed-length time sliding window is determined for each equipment identification number. The window length can be set according to the operating rules of the equipment in the past few hours. For example, referring to the operating cycle information recorded in the equipment maintenance document, the length of each window is set to 15 minutes and continues to scroll backward. When sliding, the records are classified into the corresponding time slices according to the boundaries of the current time period and the next time period, and at the same time, it is ensured that the records crossing the window boundary will be split into the two time slices before and after. Within each time slice, adjacent records are tracked one by one according to the equipment identification number and sorted in order Statistics, if it is found that the identification numbers of two records are the same, the timestamps of the records are arranged in sequence and the time series fragments of the device in this time slice are constructed. The time series fragments of all devices in all sliding windows are integrated and grouped with the device number as the index primary key. These groups are then merged to cover the records of multiple different devices in the same time range. After all groups are merged, a matrix index structure is generated. The horizontal axis corresponds to the sequential number of the time sliding window, and the vertical axis corresponds to the identification number of different devices. The elements in the matrix are the record entries of the device in the time period. If it is detected that no records are collected in some time periods, the corresponding positions are filled with null values and the gap information is additionally recorded in the monitoring log. Finally, the device record time series matrix is formed, and the time device index matrix is generated.
[0042] According to the time device index matrix obtained above, the total number of records appearing in a fixed statistical period for each device number is counted and this value is defined as the number of accesses. The difference between the current record and the previous record in the same window period is compared to calculate the update frequency. The update frequency can be determined by comparing whether the field value has changed. For records that have multiple consecutive differences in a time series, they are accumulated under the same device number. In addition, the earliest and latest appearance times of the device in the same time window set are extracted, and the active period span is obtained by subtracting them and the span is controlled within a reasonable range. For example, According to the daily operation cycle of a unit, its active period span is checked within the range of 2 to 6 hours. If the span is longer than 6 hours, this difference is marked in the monitoring log. Then, the time difference between the data generation time and the current time of each record is calculated. This time difference is obtained by directly subtracting the current system time. For example, if the time in the record is 08:00 on the same day and the current system time is 10:30 on the same day, the difference is 2.5 hours. Statistics are performed on each item and the corresponding time difference is summed up or averaged under the same equipment number to obtain the corresponding time difference summary. These values are integrated into the access behavior feature set to obtain the access behavior feature set.
[0043] formula: The benefit of the formula is that it measures the access frequency characteristics of the device during the record generation process by combining the square of the number of accesses and the logarithm of the log generation interval, and combines indicators such as update frequency and time age difference to reflect the performance of the data in terms of time series dimension and update activity. At the same time, the maximum change across cycles is taken into account in the denominator. In this way, the device data with frequent access and prominent update patterns can be prioritized in distributed storage, thereby providing more targeted sorting results for subsequent query and retrieval arrangements.
[0044] The parameter acquisition step is to use the access behavior feature set obtained above to count the total number of times each device number appears in a fixed statistical period and assign the value. For example, if device number A001 appears 280 times in a statistical period of 30 days, the value will be Set to 280.
[0045] The parameter acquisition step is to extract the time difference of all records of each device number within the statistical period, subtract the timestamps of adjacent records and obtain a series of interval values, and then sum up these interval values to get the average value of the time difference generated by the data record entries. For example, in the 280 records of device number A001, there are 279 adjacent intervals. If the total value of these adjacent intervals is 558 hours, the calculation formula is , thus we get =2.0 hours.
[0046] The steps to obtain the parameters are as follows: referring to the update frequency statistics method obtained above, accumulate all field changes of device number A001 within the statistical period, first record the field comparison results of each log with the previous log and add up the number of times the values of different fields have changed. For example, if a total of 200 field changes are observed in 280 records, and the total number of days for recording the device number is 30 days, the update frequency can be further calculated as .
[0047] The steps for obtaining the parameters are to locate the earliest timestamp in the record corresponding to each device number, subtract it from the current time, and obtain the time difference. For example, the timestamp of the first record of device number A001 is 08:00 on a certain day, and the current time is 18:00 on the same day, with an interval of 10 hours. If the statistical period spans multiple days, the overall conversion can be performed across the days. For example, if the earliest record is 08:00 on a certain day and the current time is 14:00 on the third day, the time difference can be calculated by adding 48 hours and 6 hours to get 54 hours.
[0048] The steps to obtain the parameters are as follows: in the time device index matrix, check the earliest active time point and the latest active time point covered by device number A001. The difference between the two timestamps is the active period span of the device. If the device is in an active recording state from 08:00 on the first day to 20:00 on the fifth day during the 30-day observation period, the active span of 4 days and 12 hours, totaling 108 hours, can be obtained through timing operations. Then Set to 108.
[0049] The parameter acquisition step is to read the maximum change frequency of the device records in the past 30 days and digitize it. The specific method is to average the number of changes of all fields on a daily basis and find the peak value. For example, if the device number A001 reaches 50 field changes on a certain day and the number of records on that day is 20, then the change frequency on that day is .
[0050] Calculation process: The first step is to calculate the numerator , take the above parameters of device number A001: =280, =2.0, =6.67, then , , , and then multiply to get 78401.585*7.67 601254.156.
[0051] Step 2: Calculate the denominator ,Pick =54, =108, =2.5, calculate first , and then calculate , multiplying the two together gives 1.571*2.6926 4.233, taking the absolute value unchanged, and adding 1 to it we get 5.233.
[0052] The third step is to divide the result of the first step by the result of the second step, that is, .
[0053] The fourth step is to take the absolute value of the result of the previous step (unchanged) and then square it. , record this value as .
[0054] The results show that the data fragment distribution suitability value of device number A001 is approximately 339. If the value is higher than 200, it means that it has a high distribution suitability in terms of access frequency and update characteristics. If it is lower than 50, it means that the distribution suitability is limited and the record distribution may need to be re-examined. This value can provide a reference for subsequent data sharding storage strategy or retrieval sorting decisions.
[0055] The steps for obtaining the distributed diagnostic data index are: Based on the data segment distribution fitness value, a fitness threshold is set for each data segment, and whether the data segment distribution fitness value of each data segment meets the fitness threshold standard is determined. The data segments that meet the standard are marked and pre-sorted by position allocation one by one to generate a set of data segments with pre-sorted position allocation; Allocate a pre-sorted set of data segments based on the position, call the power plant equipment identifier and the corresponding data segment distribution suitability value of each data segment one by one, and assign the data segments one by one to target storage nodes or logical partitions that meet the capacity and access performance requirements in descending order of the data segment distribution suitability values, thereby forming a set of data segments that have completed the assignment of storage nodes or logical partitions; Based on the set of data fragments that have completed the assignment of storage nodes or logical partitions, the location information of the target storage node or logical partition where each data fragment is located and the associated power plant equipment identification are recorded one by one, index entries are created in chronological order, and a distributed diagnostic data index is generated.
[0056] Specifically, based on the previously obtained data segment distribution fitness values, we first find the data usage and access requirements of various types of power plant equipment in the reference document, and then select a fitness threshold based on the importance and load of the equipment and mark the threshold as For example, by looking up the suitability distribution of the same type of equipment in the past 30 days, we can see that the suitability values of most devices are between 100 and 600. After counting and filtering this range, we can choose to set Then, all data segments are traversed and their distribution fitness values and corresponding power plant equipment information are read to determine whether the fitness value of each data segment is not less than If the suitability value of a data fragment is greater than or equal to 300, its registration table will be marked and temporarily included in the qualified set. If the suitability value of a data fragment is less than 300, it will be temporarily included in the set to be processed later. After completing the comparison of the suitability values of all data fragments, the data fragments in the qualified set will be pre-sorted in order of numerical value. This sorting process will number each fragment from high to low according to the suitability value and append it to a pre-prepared index list. For fragments with exactly the same suitability values, they will be sorted twice according to the dictionary order of the equipment identification. After all the sorting is completed, a position-allocated pre-sorted data fragment set is obtained. At the same time, the suitability value and power plant equipment identification number of each entry in this set are recorded in sequence to form a position-allocated pre-sorted data fragment set.
[0057] The pre-sorted data segments are allocated based on the positions obtained in the previous step. The available capacity of existing target storage nodes and logical partitions in the current system is first checked. The remaining capacity and bandwidth load information for each node can be found in relevant documents or resource management records. The power plant equipment identifier and distribution suitability value of each sorted data segment are then retrieved. These distribution suitability values are read from highest to lowest, starting with the data segment at the top of the sort. For each data segment, the required storage space and write efficiency are checked. If the required capacity of the segment still has sufficient capacity within the remaining capacity of a target storage node, that node is selected. If the corresponding node capacity is insufficient or the access bandwidth exceeds a predetermined standard, such as 300MB / s, the next suitable node is selected. After the allocation is complete, the node or logical partition number where the segment finally lands is recorded. If multiple nodes meet the capacity requirements, a secondary determination based on access bandwidth or response latency is performed to determine the priority of the assignment. After all segments have been assigned to target nodes or logical partitions, the number of successfully mapped entries is counted and classified as the set of data segments that have been assigned to storage nodes or logical partitions.
[0058] Based on the data fragment set that has completed the storage node or logical partition assignment, the storage node or logical partition information where each data fragment is finally located and the corresponding power plant equipment identification are read in sequence, and index entries are created for each data fragment in chronological order. Each index entry contains key information such as the storage location number and equipment identification. When adding an entry, the start timestamp and end timestamp corresponding to the fragment are also written into the index. Then, the newly generated index entries in the whole process are cumulatively numbered to facilitate query and positioning during subsequent retrieval. Finally, all index entries are arranged in a time-ordered chain structure or a two-dimensional mapping table, and the complete set is summarized into a distributed diagnostic data index.
[0059] The steps to obtain the diagnostic data retrieval path are: Receive a target power plant equipment fault diagnosis query request, extract the device identifier field and the query time range field, use the device identifier field to match the device index primary key in the distributed diagnostic data index, and use the query time range field to perform interval matching in the time period index key to form a set of positioning data fragments; Based on the set of located data segments, the context retrieval priority of each data segment is calculated using the following formula: ; in, Indicates the The context search priority of each data segment, Indicates the The total number of times the device is referenced in the current active alarm list for each data segment. Indicates the The total number of times a data fragment has been accessed in the past 30 days. Indicates the The number of failed accesses to the power plant equipment corresponding to each data segment in the past 30 days. Indicates the The total number of device entries associated with the logical partition where the data fragment is located. Indicates the Each data fragment corresponds to the position number of the power plant equipment in the associated equipment list. Indicates the The starting timestamp value of each data segment, Indicates the start timestamp value of the time range set in the current diagnostic query request; According to the context retrieval priority, all positioning data fragments are sorted from high to low according to the context retrieval priority, and each data fragment is arranged in turn into the list of data fragments to be retrieved, and dispatched to the retrieval task execution plan queue in order to generate the diagnostic data retrieval path.
[0060] Specifically, a target power plant equipment fault diagnosis query request is received, and the device identifier field and query time range field included in the request are read. At the beginning of the query process, the device identifier field is first compared with a list of power plant equipment stored in the system. If a matching device number is found, the number is recorded and the corresponding valid information entry is checked to see if it meets the minimum record requirements. If so, the execution continues. If not, an investigation is carried out and a note is written in the investigation record that the device may have missed some content in the early data collection. The query time range field is then read and split into the start time and the end time. The time period index key for each device in the storage system is compared based on these two time values. The specific approach is to convert the query start time and end time into a standard timestamp format and search the index one by one to see if there is an overlapping time period. When a certain time is found, When the interval overlaps with the query range completely or partially, the corresponding index entries are temporarily stored in a temporary list and the corresponding start and end nodes are marked in the list. If it is found that the data entry has not exceeded the time limit of the day but still falls within the query range, it is also included in the matching results and screened in subsequent steps. When all available index keys are compared, the integrity of all entries in the temporary list is judged one by one. For example, the starting timestamp of each record is compared with the lower limit of the query range. If it is less than the specified lower limit, a time truncation is noted in the field. Similarly, if the end time is greater than the specified upper limit, a similar truncation is performed. This operation can be performed in a segmented statistical manner and the number of all truncated fields is recorded in sequence. After the truncation is completed, a batch of data entries that meet the query range are obtained. Finally, after summarizing these data entries, a set of positioning data fragments is formed.
[0061] formula: The benefit of the formula is that it can make a detailed ranking of the importance and relevance of data fragments in the current fault diagnosis context by comprehensively considering multiple factors such as the number of alarm list references, the number of fragment accesses, the number of device access failures, the number of associated device entries, the position of the device in the associated list, and the timestamp deviation. In the scenario of parallel fault diagnosis of large-scale power plant equipment, this formula can provide a refined retrieval order for query scheduling, allowing the retrieval system to prioritize fragments that are more urgent or more likely to contain critical data.
[0062] The steps to obtain the parameters are to retrieve the alarm from the current active alarm list. Each data segment corresponds to all alarm records matching the device number, and the total number of times the device is referenced in these alarm records is counted and accumulated into an integer value. For example, after counting the alarm records within 30 days, the first The device corresponding to the data fragment is referenced 12 times, that is, =12.
[0063] The steps to obtain the parameters are to count the number of The total number of times a data fragment has been externally queried or internally scanned in the past 30 days. This number can be calculated by keyword matching or log filtering on the access records. Each successful operation of reading or retrieving the data fragment leaves a record with the fragment ID and timestamp in the access log. Finally, the counts in the same time period are summed up as For example, if there are 40 requests to retrieve a certain fragment, and the scheduling system regularly scans the background for an additional 5 times, then Take 45.
[0064] The steps to obtain the parameters are as follows: The number of access failures in the power plant equipment corresponding to each data fragment in the past 30 days is counted. Access failure refers to the situation where the client or internal process attempts to read the data fragment but fails to complete it successfully. Each failure event is usually clearly recorded in the monitoring system or log system, including the timestamp and the reason for the failure. For example, if there are 3 access failure records related to the device ID corresponding to the fragment in the past 30 days, then =3.
[0065] The steps to obtain the parameters are as follows: The number of entries of all associated devices in the logical partition where the data fragment is located is accumulated, that is, the total number of device entries under the same logical partition is counted. This partition usually stores data fragments of multiple devices. At the same time, there is a fixed limit on the maximum number of entries that each partition can accommodate during system configuration. The total number of device entries under the current partition is read through the index structure. For example, if the partition contains log and alarm entries of 10 different devices, the total number of device entries may reach 500. =500.
[0066] The steps to obtain the parameters are to find the specific location number of the device in the list of devices related to the current power plant equipment. This location number is usually assigned by the system to the relevant devices and recorded in a related device table. The devices are arranged in the order of creation or registration time and marked with serial numbers, such as The device to which the data fragment belongs ranks fifth in the list of associated devices. =5.
[0067] The steps to obtain the parameters are: The actual starting timestamp of a data fragment. This timestamp comes from the operation status log or alarm record and is usually stored in the system in the format of year-month-day hour:minute:second. Convert the timestamp into a numerical form for arithmetic calculations. For example, if the timestamp is 2025-03-05 08:30:00, it can be converted into seconds for subsequent calculations.
[0068] The parameter acquisition step is to extract the set time range start timestamp from the current diagnostic query request. This time range is written into the request body when the user submits the fault diagnosis request. The units and formats are consistent. A format check will be done before the formal calculation to ensure Able to Perform subtraction and root operations. If the time range specified in the query is 2025-03-05 00:00:00, convert it to a numerical value. =The corresponding timestamp value.
[0069] Calculation process: The first step is to calculate the molecular part ,set up =12 o'clock, , , and then calculate ,set up =45, then ,set up =3, , 91125 / 4=22781.25, and then add the two to get 12.04+22781.25=22793.29.
[0070] Step 2: Calculate the right bracket ,set up =500, then , ,like =5, then , so 7.94+2.807=10.747.
[0071] The third step is to multiply the two, that is, 22793.29*10.747 245,004.40.
[0072] Step 4, denominator part, calculation ,like =1,611,163,400, =1,611,163,200, the difference between the two is 200 seconds, then square root , its square is .
[0073] Step 5: Divide the result of step 3 by the result of step 4, which is 245,004.40 / 229.286 1069.15, the result remains unchanged after taking the absolute value.
[0074] Therefore, Context search priority for each data segment About 1069.15.
[0075] This result shows that when A higher value means that the data fragment has a higher priority in alarm reference and historical access, and is relatively close to the start timestamp of the query time range. It can be queued for retrieval first when used for fault diagnosis. A value greater than 500 indicates that the fragment is dense in terms of alarm references and access times. A value lower than 10 indicates that the contextual retrieval value is not obvious. This value provides a comparative reference scale for subsequent ranking steps.
[0076] According to the context retrieval priority obtained previously, the corresponding priority values of all located data fragments are read one by one and an arrangement operation is performed from large to small. This operation will create a temporary list to store all located fragment identifiers and their priorities. When sorting, the absolute size of the priority is compared first. If the priority values of two fragments are the same, further judgment is made based on the logical partition or index order. After completing the above comparison, a list sequence with the highest priority fragment at the front is obtained. Then, each data fragment is appended to the list of data fragments to be retrieved according to the new order and dispatched in sequence. When dispatching, it is necessary to reserve a processing position for each fragment in the retrieval task execution plan queue. After all fragments are assigned an order, a diagnostic data retrieval path is formed.
[0077] The steps to obtain the retrieval task completion metric are: According to the search order and target logical partition location information set in the diagnostic data search path, the path address and storage node number of each data segment in the task are retrieved one by one, and cross-partition parallel data is pulled to generate a list of returned data segments; Based on the list of returned data fragments, the number of record entries, the number of accurate field comparison items, the number of missing field bits, the number of field bits that failed verification, the number of structural position misalignments, and the number of abnormal label fields of each returned data fragment are called to form a fragment field verification parameter set; Based on the fragment field validation parameter set, the retrieval task completion metric is calculated using the following formula: ; in, represents the retrieval task completion measure, Indicates the The number of entries with successfully matched fields in the data fragment, Indicates the The number of bits of missing fields in each data segment, Indicates the The number of fields that failed validation in each data segment, Indicates the The total number of record entries for data fragments, Indicates the The number of fields with misplaced structure in each data segment, Indicates the The number of abnormal label fields in the data fragment, Indicates the The total number of fields returned for each data fragment.
[0078] Specifically, according to the retrieval order set in the diagnostic data retrieval path and the target logical partition location information, the path address of each data fragment and its corresponding storage node number are read in sequence. When reading, the concurrency is set according to the maximum number of threads that can be accessed in parallel and the retrieval threads are allocated one by one. A corresponding request instruction is created for each path address, and the sending timestamp of the instruction is recorded in the system. After the allocation of these instructions is completed, read operations are started on different logical partitions. When the return is received, the overall length and number of fields of the data are analyzed and compared with the expected value registered previously. If the number of fields or the number of entries is inconsistent, the partition is marked as suspicious in a record table and re-checked in the subsequent processing stage. When each returned data meets the correspondence between the registered path address and the device number, the returned data is The system packages the fragment object with its original information into an identifiable fragment object and saves it to the cache. The system records the index identifier and sequence number for each object in the cache. If it is found that the fragments pulled in some parallel threads respond too slowly, the network status will be tracked on the thread and the specific delay will be reported. The threads whose delay exceeds the preset threshold will be included in the warning range. The threshold is usually set based on the internal network performance statistics of previous retrievals. For example, it is considered acceptable if it is set in the range of 200 milliseconds to 500 milliseconds, and it is marked as a high-latency connection if it exceeds 500 milliseconds. Once the high delay is confirmed, the corresponding partition will be checked first to see if it is too loaded and resources will be appropriately scheduled in subsequent rounds. Finally, after the response of all threads is completed, the results from multiple partitions are aggregated to obtain all the information that meets the items listed in the diagnostic data retrieval path, forming a list of returned data fragments.
[0079] According to the list of returned data fragments, the detailed record information corresponding to each fragment is called for field comparison. First, the number of record entries registered in the system for the fragment is read one by one and the number of entries is ensured to be within a pre-defined reference range. For example, the number of entries in a day's real-time collection operation is between 100 and 500. Then, the exact number of field comparison items is extracted. The field content is judged to be consistent with the expectation through the hash value or plain text comparison of each field. When a difference is found between the field content and the original registration, a field mismatch count is incremented in the statistics. If some fields are completely empty, they are accumulated into the missing field bit. In addition, for the system The defined validation rules will also be verified one by one. If a key field in a record does not meet the standard or is inconsistent with the template, the number of corresponding validation failed fields will be increased by one. When further checking the order of the structure, if the field position is misplaced, for example, a field that was originally expected to appear in the third position appears in the fifth position, the number of misplacements will be added and recorded in the field verification list. At the same time, the number of abnormal label fields will be identified one by one to determine whether there are fields marked with faults or special alarm marks. If so, they will be counted in the number of abnormal label fields. After completing the statistics of all fragments, these results will be summarized in sequence to finally obtain the fragment field verification parameter set.
[0080] formula: The usefulness of the formula lies in that it comprehensively considers the relationship between the number of successfully matched fields, the number of missing fields, the number of fields that failed verification, the total number of record entries, the number of structurally misplaced fields, and the number of abnormal label fields and the total number of fields. By combining multiple deviation terms and accuracy terms, it can comprehensively reflect the overall completion of the current retrieval task. In the scenario of distributed cross-partition retrieval, this formula provides a basic reference for subsequent judgments on whether the retrieval needs to be retried or whether the data needs to be re-verified.
[0081] The steps to obtain the parameters are: After a data segment is created, the number of entries in all records where the fields have been matched successfully is counted. This count can be obtained by the program when comparing with a given field template or reference structure. If the program detects that all fields in a certain record are consistent with the original registration, the record is added to the successful entry count. For example, if a segment has 200 records and all fields in 180 records are correct, the record is added to the successful entry count. Set to 180.
[0082] The steps to obtain the parameters are: The number of bits of all missing fields in a fragment. A missing field means that the field content is empty. It may appear as an empty string or invalid identifier in the log or alarm record. The statistical method can be achieved by scanning each field in each record and checking whether it has valid content. If a field is found to be empty, it is counted as 1. For example, if a total of 50 empty fields are detected in 200 records, then =50.
[0083] The parameter acquisition step is to compare the content of the field with the standard template or data dictionary according to the previous verification rules. If there is any inconsistency or error, the verification failure count will be increased by 1. Finally, the total number of fields with such errors is For example, if 25 fields are found to be mismatched or have incorrect content in 200 records, =25.
[0084] The steps to obtain the parameters are as follows: The number of entries in all records in a fragment is summed up. If the fragment contains 200 records, then =200.
[0085] The parameter acquisition step is that if the field order deviates from the expected structure during retrieval, the number of these misplaced fields is recorded and summarized in the statistical process. For example, if there are 10 records with field misalignment in the 200 records mentioned above, each misalignment may have several fields that are not in the correct order. By adding up the misaligned fields of all entries, the total number of misaligned fields in the fragment is obtained, which is recorded as If a total of 30 misplaced fields are detected, then =30.
[0086] The parameter acquisition step is to accumulate the fields with abnormal label identification found in the retrieval stage one by one. The abnormal label refers to the field with a clear fault or alarm additional description in the record. For example, if 15 fields are detected with abnormal label markings in 200 records, then =15.
[0087] The steps to obtain the parameters are: The total number of fields returned by a fragment. This parameter is slightly different from the number of record entries. The number of record entries is a count of stripes, while the field is the sum of the elements within a single record. If each record in the fragment has 10 fields and there are 200 records, then .
[0088] Calculation process: The first step is to calculate ,set up =180, =50, =25, =200, =30, then , bring it into = = .
[0089] The second step is to calculate ,set up =15, , add it to the result of the first step to get .
[0090] The third step is to calculate the denominator. ,set up =2000, .
[0091] Step 4, finally = .
[0092] The results show that the current retrieval task maintains a relatively controllable level in terms of field matching accuracy and error matching. A value greater than 10 indicates a high degree of search completion. A value less than 2 indicates that the search may have a large number of omissions or verification failures. The size of the value determines whether the data in the segment needs to be reread or further examined.
[0093] The steps to obtain the associated fault diagnosis view are as follows: Based on the retrieval task completion metric, the data segments whose retrieval task completion metric meets the preset qualification standard are compared one by one from the returned data segment list, and all the data segments that pass the comparison are retained to form a set of data segments that pass the verification; Based on the verified data fragment set, the power plant equipment identifier, query time range, current active alarms, and associated equipment list in the target power plant equipment fault diagnosis query request are called, and the aggregated data fragment content is matched with the query request context one by one to obtain the data fragment set with complete context association; Based on the set of data fragments with completed context association, all associated fields in the data fragments are called one by one, and information is assembled and formatted and output according to the device identifier and time range specified in the query request context to generate an associated fault diagnosis view.
[0094] Specifically, based on the previously obtained retrieval task completion metrics, a reference list containing preset qualification criteria values is read. This list typically records the distribution range of completion metrics statistically collected from similar past retrieval tasks. For example, in multiple tests and actual use, it was observed that most completion metrics were distributed between 2 and 25. Based on this distribution range, the technician set the qualification criteria threshold to 10. The retrieval task completion metrics of all data segments in the current list are then compared with this threshold. If a segment's completion metric is greater than or equal to 10, it is marked as passed. Otherwise, it is marked as temporarily failed and the reason is recorded. The marked passed segments are arranged according to their completion metrics and a description is attached to each segment, including the number of entries and field comparison status of the segment. Then, during the traversal process, it is recorded which segments have completion metric values greater than 20 and more detailed sampling tests are performed on these segments. If the test results are confirmed to be consistent with the system expectations, they are retained in the priority set. If any fields are still missing or erroneous, the segment is removed from the passed queue. Finally, after strict threshold comparison and sampling verification, all passed segments are retained to obtain the verified data segment set.
[0095] According to the previously formed set of verified data fragments, the device identifier and query time range in the target power plant equipment fault diagnosis query request are retrieved one by one. During on-site operation, the device identifier (for example, "DEV001") and the start and end time (for example, 2025-01-0108:00:00 and 2025-01-0120:00:00) are first split out. Then, it is checked whether the current active alarm list contains any alarm records that match the device. If so, the fault information indicated by these alarm records is read from the active alarm list and compared with the verified data fragments at the field level. If a fragment that overlaps with the alarm trigger event is found in the same time interval, the fragment is additionally marked. Signature identification, read the associated device list at the same time and determine whether there is a linkage relationship between the records of other devices and the fragment. For example, if another device in the same control loop as the device is found in the associated device list, it is necessary to aggregate the log records of the two into one scene during aggregation and indicate the mutual reference between the two. After a global scan of all fragments, alarms and associated devices, a secondary sorting is performed based on the continuity of the timestamps and it is confirmed whether the connection between the fragment content meets the query time range. If it exceeds the start or end time, the fragment content is intercepted and the corresponding start or end mark is added. Finally, all the aggregated fragments are combined into a content-rich context data entry summary, thereby obtaining a set of data fragments with complete context association.
[0096] Based on the above contextually associated data fragments, the name and value of each field in the fragment are read one by one and checked to see if they are consistent with the device identifier in the target power plant equipment fault diagnosis query request. The same check is performed on the time range. After satisfying the device identification and timestamp constraints, the relevant fields are merged. If some fields are found to correspond to multiple log sources (for example, one from the operating status log and one from the alarm record), these fields are merged into the same structure and arranged in chronological order. Fields that do not match the device identifier or are outside the specified time range are discarded. The fully aggregated field information is formatted for output. During the formatting process, the log time can be converted from "YYYY-MM-DDHH:MM:SS" to a unified second count format, or the original text can be retained but an additional time zone marker is added. The aggregated field content is then imported into the drawing process. Through references between fields and the hierarchical display of alarm records, the information related to the current diagnostic query request is visualized or structured, allowing technicians to more easily view and compare, ultimately generating a correlated fault diagnosis view.
[0097] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for intelligent diagnosis of power plant equipment faults, characterized in that: The following steps are involved: Collect power plant operating status log records and alarm records, perform timestamp alignment and device identification association processing, calculate the source data regularity score based on the data source integrity and time series continuity, and based on the source data regularity score, screen data entries with source data regularity scores that meet the standards and perform unified formatting and integration processing to establish a standardized structured diagnostic data set; Based on the structured diagnostic data set, analyzing the time attributes of the data records, the associated power plant equipment identifiers, and the access frequency statistics, combining the data age and the equipment importance to derive a data segment distribution fitness value, assigning each data segment to a target storage node or logical partition based on the data segment distribution fitness value, creating an index entry including storage location information, and generating a distributed diagnostic data index; receiving a diagnostic query request for a target power plant equipment fault, parsing the equipment identifier and time range, searching for a location data segment using the distributed diagnostic data index, calculating the query relevance of the segment in combination with the currently active alarm and the associated equipment list to obtain a contextual retrieval priority, sorting the list of data segments to be retrieved based on the contextual retrieval priority, and obtaining a diagnostic data retrieval path; According to the order and location information determined in the diagnostic data retrieval path, data fragments are retrieved from each specified storage node or logical partition, the returned data is verified against the expected data items defined by the path, and a retrieval task completion metric is calculated. Based on the retrieval task completion metric, the verified data is aggregated, the information is assembled according to the query request context, and an associated fault diagnosis view is generated.
2. The intelligent diagnosis method for power plant equipment faults according to claim 1, characterized in that: The steps for obtaining the source data regularity score are: Collect power plant operation status log records and alarm records, unify data source labels, unify sampling time bases, and linearly merge record time axes for power plant operation status log records and alarm records corresponding to the same equipment identifier to generate an equipment record fusion sequence; Based on the device record fusion sequence, the total number of fields, the number of null value fields, the sampling time interval, the frequency of field value changes, the number of abnormal alarm triggers and the cross-sequence data synchronization delay time are extracted one by one to obtain a six-dimensional parameter set of regularity; The source data regularity score is calculated based on the six-dimensional regularity parameter set.
3. The intelligent diagnosis method for power plant equipment faults according to claim 1, characterized in that: The steps for obtaining the standardized structured diagnostic data set are: Collect the source data regularity score, use the device type and the source data regularity score as screening parameters, determine and retain data entries whose source data regularity scores meet the preset standard, mark the status of each record entry that meets the standard, and generate a set of data entries whose scores meet the standard; Based on the set of data entries that meet the scoring criteria, the data field structures of the original data entries are called one by one, the field data types and field lengths are standardized, and the field positions and contents are rearranged according to the unified field sorting rules to form a set of data entries with a unified field format; According to the data entry set with a unified field format, field mapping of structured data is performed with device identification, timestamp and record category as associated primary keys to form a standardized structured diagnostic data set.
4. The intelligent diagnosis method for power plant equipment faults according to claim 1, characterized in that: The steps for obtaining the data segment distribution suitability value are as follows: Based on the structured diagnostic data set, extract the timestamp information and power plant equipment identification number corresponding to each data record, form an equipment record time series matrix through continuous time window aggregation and equipment record tracking, and generate a time equipment index matrix; According to the time device index matrix, the number of record accesses, update frequency, and active period span of each device number within a fixed statistical period are counted, and the time difference between the data generation time corresponding to each record and the current time is calculated to obtain the access behavior feature set; Based on the access behavior feature set, a data segment distribution suitability value is calculated.
5. The intelligent diagnosis method for power plant equipment faults according to claim 1, characterized in that: The steps for obtaining the distributed diagnostic data index are: Based on the data segment distribution fitness value, setting a fitness threshold for each data segment, judging whether the data segment distribution fitness value of each data segment meets the fitness threshold standard, marking the data segments that meet the standard and pre-sorting them one by one by position allocation, and generating a set of data segments with pre-sorted position allocation; Allocating a pre-sorted set of data segments according to the position, calling the power plant equipment identifier and the corresponding data segment distribution suitability value of each data segment one by one, and assigning the data segments one by one to target storage nodes or logical partitions that meet capacity and access performance requirements in descending order of the data segment distribution suitability values, thereby forming a set of data segments that have completed storage node or logical partition assignment; Based on the set of data fragments that have completed the assignment of storage nodes or logical partitions, the location information of the target storage node or logical partition where each data fragment is located and the associated power plant equipment identification are recorded one by one, index entries are created in chronological order, and a distributed diagnostic data index is generated.
6. The intelligent diagnosis method for power plant equipment faults according to claim 1, characterized in that: The steps for obtaining the diagnostic data retrieval path are: Receive a target power plant equipment fault diagnosis query request, extract a device identifier field and a query time range field, use the device identifier field to match the device index primary key in the distributed diagnostic data index, and use the query time range field to perform interval matching in the time period index key to form a set of positioning data segments; Calculating a context retrieval priority of each data segment based on the set of positioning data segments; According to the context retrieval priority, all positioning data segments are sorted from high to low according to the context retrieval priority, and each data segment is arranged in turn into the list of data segments to be retrieved, and is dispatched to the retrieval task execution plan queue in order to generate a diagnostic data retrieval path.
7. The intelligent diagnosis method for power plant equipment faults according to claim 1, characterized in that: The steps for obtaining the retrieval task completion metric are as follows: According to the search order and target logical partition location information set in the diagnostic data search path, the path address and storage node number of each data segment in the task are retrieved one by one, and cross-partition parallel data is pulled to generate a list of returned data segments; According to the returned data fragment list, the number of record entries, the number of accurate field comparison items, the number of missing field bits, the number of field bits that failed verification, the number of structural position misalignments, and the number of abnormal label fields of each returned data fragment are called to form a fragment field verification parameter set; A retrieval task completion metric is calculated based on the segment field validation parameter set.
8. The intelligent diagnosis method for power plant equipment faults according to claim 1, characterized in that: The steps for obtaining the associated fault diagnosis view are as follows: Based on the retrieval task completion metric, the data segments whose retrieval task completion metric meets the preset qualification standard are compared one by one from the returned data segment list, and all data segments that pass the comparison are retained to form a set of data segments that pass the verification; Based on the verified data segment set, the power plant equipment identifier, query time range, current active alarms, and associated equipment list in the target power plant equipment fault diagnosis query request are called, and the aggregated data segment content is matched with the query request context one by one to obtain a data segment set with complete context association; According to the data segment set with context association completed, all associated fields in the data segment are called one by one, and information is assembled and formatted and output according to the device identifier and time range specified in the query request context to generate an associated fault diagnosis view.
Citation Information
Patent Citations
Power plant data management method and system based on SIS system
CN119537358A
Distribution automation operation and maintenance work log intelligent recording platform based on CS architecture
CN119692728A
Knowledge conversion and fusion processing method and system for massive power grid operation data
CN120012884A
Data fragment storage index optimization method and system
CN120234140A
Power plant system equipment fault diagnosis method and system based on artificial intelligence
CN120408533A