Maintenance interval rating system development method based on fuzzy clustering

By processing rail vehicle operation and maintenance data using fuzzy clustering, a reasonable maintenance interval rating system was developed, which solved the problem of unreasonable maintenance intervals in existing standards, improved maintenance efficiency, and reduced costs.

CN121934818APending Publication Date: 2026-04-28CRRC CHANGCHUN RAILWAY VEHICLES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CRRC CHANGCHUN RAILWAY VEHICLES CO LTD
Filing Date
2025-11-20
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The lack of specific methods in existing standards to determine the maintenance cycle of rail vehicles leads to unreasonable maintenance intervals and problems of over-maintenance and under-maintenance.

Method used

A fuzzy clustering-based method is adopted to collect and process vehicle operation and maintenance record data, calculate fuzzy matrices and transitive closures, draw dynamic clustering diagrams, select the optimal number of classifications, and develop a preventive maintenance task interval rating system suitable for rail transit vehicles.

Benefits of technology

This has enabled more reasonable and accurate maintenance interval setting, reduced over-maintenance and missed maintenance, improved maintenance efficiency and online rate, and reduced costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121934818A_ABST
    Figure CN121934818A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of railway vehicle maintenance task interval decision making, in particular to a maintenance interval rating system development method based on fuzzy clustering, which can reasonably and effectively classify all factors influencing vehicle system, structure and equipment damage on the basis of a current maintenance interval framework, and improve the maintenance interval rating efficiency. According to the method and the system, the preventive maintenance task interval rating system suitable for the structure, the area and the like of the rail transit vehicle can be formulated conveniently, the preventive maintenance task interval of the structure, the area and the like can be formulated more reasonably and accurately, over-maintenance and missed maintenance are reduced, the online rate is improved, and cost reduction and efficiency improvement can be realized more accurately and effectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of railway vehicle maintenance task interval decision technology, and in particular to a development method for a maintenance interval rating system based on fuzzy clustering. Background Technology

[0002] The rail transit industry has developed for many years, but the compilation of maintenance procedures is still in the experience stage. Although the industry is promoting the RCM method for the formulation of maintenance procedures, not all users are using it, especially in terms of the formulation of maintenance cycles. Many companies are looking for the best method to determine or optimize maintenance cycles.

[0003] For large and complex equipment such as rail vehicles, routine maintenance and visual inspection tasks are usually carried out by area. This is conducive to the arrangement of procedures and the convenience of inspection, thereby improving maintenance efficiency and reducing maintenance costs.

[0004] Inspections conducted by area typically include checking for external damage, discoloration, and missing components such as accessible system parts, structural components, cables, and pipes. Most of these failures are caused by accidental damage sources in the external environment (passenger activity, maintenance activities, etc.) and environmental conditions (high temperature, humidity, etc.). Component failures caused by these reasons generally cannot be interval-calculated using lifespan data analysis. Therefore, it is necessary to find a suitable method to determine the characteristics of changes in the similarity, frequency, and cycle of component failures under different conditions, in order to define specific intervals (intervals forming the current maintenance framework).

[0005] Reliability-Centered Maintenance (RCM) standards such as GJB 1378A-2007, IEC 60300-3-11 2009, and S4000P contain requirements for area inspection analysis and provide recommended interval rating systems. However, these standards only provide general guidelines and lack specific analytical methods, making it difficult to effectively guide rail vehicle suppliers / users in determining the intervals for routine visual inspections.

[0006] For example, the S4000P standard provides a template for a structural analysis rating system in the recommended form for structural RCM analysis. This includes factors such as the corrosion level of the material (3 types), the level of the affected factors (3 types), and the level of the structure's own protection (3 types). Other influencing factors may also be considered. Qualitative indicators are converted into quantitative intervals in the form of a matrix. However, it does not specify which factors correspond to which intervals. This is because there are many combinations of influencing factors (e.g., 3×3×3=27 types), but the number of interval frames corresponding to these combinations is fixed, and the final interval matrix also needs to be fixed. Therefore, it is necessary to find a reasonable method to determine which similar factors correspond to the same interval.

[0007] Although there are recommended interval determination methods in international standards, they do not provide specific explanations of the principles or procedures. It is still necessary to develop specific methods to determine intervals suitable for area-based inspection tasks of rail transit vehicles based on the given methods.

[0008] Therefore, it is essential to provide a development method for a maintenance interval rating system based on fuzzy clustering to address the shortcomings of existing technologies. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a development method for a maintenance interval rating system based on fuzzy clustering. This fuzzy clustering-based maintenance interval rating system development method can, on the basis of the current maintenance interval framework, reasonably and effectively classify all factors affecting vehicle system, structure, and equipment damage, so as to formulate a rating system for preventive maintenance tasks such as structure and area maintenance applicable to rail transit vehicles. It can more reasonably and accurately formulate preventive maintenance task intervals for structure and area maintenance, reduce over-maintenance and missed maintenance, and improve the online rate, that is, it can more accurately and effectively achieve cost reduction and efficiency improvement.

[0010] The above-mentioned objectives of the present invention are achieved by the following technical means.

[0011] A method for developing a maintenance interval rating system based on fuzzy clustering is provided, including the following steps:

[0012] S1: Collect data based on vehicle operation and maintenance records, collect fault information of vehicles at different times, and collect the current vehicle maintenance interval framework.

[0013] S2: Process data. Based on the collected fault information, process the data according to the data usage requirements to obtain high-quality usable data.

[0014] S3: Standardized data, which is the data obtained from S2 that has been preprocessed to standardize the data matrix.

[0015] S4: Calculate the fuzzy matrix by using a similarity calculation method to calculate the fuzzy matrix from the standardized data matrix;

[0016] S5: Calculate the transitive closure by performing stepwise calculations on the fuzzy matrix using the flat method to obtain the transitive closure matrix.

[0017] S6: Output a dynamic clustering graph, which is drawn based on the calculated transitive closure.

[0018] S7: Select the classification result. Based on the dynamic clustering graph, select the final number of classifications according to the optimal number of classifications or the relevance principle.

[0019] S8: Output the classification results and formulate the relationship matrix of the rating system according to the selected number of categories.

[0020] Specifically, the method for step S2 is as follows:

[0021] S21: Organize the data, collect fault records and maintenance records according to the characteristics of the fault, and obtain the initial standard data form;

[0022] S22: Filter the data, remove component failure data caused by human factors, quality reasons, etc. from the initial standard data form, and obtain a standard data form that can be used for analysis;

[0023] S23: Check for duplicate data, identify duplicate data in the integrated standard data form, and obtain the standard data form for duplicate checking;

[0024] S24: Data restoration. The missing information in the plagiarism check standard data form is restored to obtain the restored standard data form, which is divided into the following cases;

[0025] When the fault is missing, the vehicle has already accumulated mileage L. f However, the time T when the fault occurred is known. f The following restoration methods can be used:

[0026] If the vehicle's utilization rate (AOR) and the start date of its use (T0) are available, the following formula is used for reconstruction:

[0027] L f =(T f -T0)·AOR

[0028] If there is vehicle utilization rate (AOR) and the date T before the data collection deadline... n and the corresponding cumulative mileage L n The following formula is used for restoration:

[0029] L f =L n -(Tn -T f )·AOR

[0030] If there is vehicle utilization rate (AOR) and other fault data information for vehicles in the same column as the fault data object, i.e., the fault time T1 and corresponding cumulative mileage L1 of another data entry, the following formula is used for restoration:

[0031] L f =L1-(T1-T f )·AOR

[0032] If there is no vehicle usage rate, but there are two other fault data points for the same vehicle in the same column as the fault data object, namely the fault times T1 and T2 of the other two data points, and the corresponding cumulative mileage L1 and L2, the following formula is used to restore the data:

[0033]

[0034] L f =L2-(T2-T) f )·AOR or L f =L1-(T1-T f )·AOR

[0035] If there is no vehicle usage rate, but there is other fault data information for the vehicle in the same column as the fault data object, namely the fault time T1 and the corresponding cumulative mileage L1 of another data, and the start date of use T0 is known, the following formula is used to restore it:

[0036]

[0037] L f =(T f -T0)·AOR

[0038] If there is other fault data information for the vehicle in the same column as the fault data object, i.e., the fault time T1 and the corresponding cumulative mileage L1 of another data entry, and also the date T of the data collection deadline. n and the corresponding cumulative mileage L n The following formula is used for restoration:

[0039]

[0040] L f =L n -(T n -T f )·AOR or L f =L1-(T1-T f )·AOR

[0041] If there is no vehicle usage rate, then there is a data collection deadline date T. n and the corresponding cumulative mileage L n Knowing the start date T0 of its use, the following formula is used for restoration:

[0042]

[0043] L f =L n -(T n -T f )·AOR

[0044] S25: Split the data. The same data record containing multiple components is split according to the component serial number to obtain a split standard data form.

[0045] S26: Identify data. For the split standard data form, identify the environmental factors and failure modes of each data item to obtain the identified standard data form.

[0046] S27: Calculate the lifespan of the faulty component. The lifespan of the component is calculated based on known data conditions, and is divided into the following modes;

[0047] Given the date of failure, the date of installation, and the vehicle's usage rate, the lifespan interval is calculated as follows:

[0048] Lifespan range = (Date of failure - Date of installation) × Vehicle usage rate

[0049] Given the vehicle's cumulative mileage at the time of the malfunction and the vehicle's cumulative mileage at the time the component was installed, the lifespan range is calculated as follows:

[0050] Lifespan range = Total mileage of vehicle when the failure occurred - Total mileage of vehicle when the component was installed. Given the failure date and installation date, the lifespan range can be directly calculated in calendar time as follows:

[0051] Lifespan range = Date of failure - Date of installation.

[0052] Specifically, the method for step S3 is as follows:

[0053] S31: Statistical analysis of fault data. According to the results identified in step S26, statistical analysis of the number of faults occurring under each environmental factor distributed across different period ranges.

[0054] S32: Establish the initial data matrix. Compile the statistical results from step S31 into a matrix format to form the initial data matrix X = (x ij ), 1≤i≤n, 1≤j≤m;

[0055] in:

[0056] x ij The value in the i-th row and j-th column of the matrix represents the number of faults in the j-th cycle under the i-th factor.

[0057] n represents the number of combinations of environmental factors;

[0058] m is the set number of cycles;

[0059] S33: Calculate the maximum value parameter by removing the maximum value from each column of matrix X; maximum value parameter M j For: M j =max(x 1j x 2j , ..., x nj ), 1≤j≤m;

[0060] S34: Calculate the normalized data matrix, based on X determined in step S32 and M calculated in step S33. j Calculate the standardized data matrix;

[0061] The standardized data matrix is: X′=(x′ ij ), 1≤i≤n, 1≤j≤m

[0062] in:

[0063]

[0064] Specifically, the method for step S4 is as follows:

[0065] For the standardized data matrix X′ obtained in step S34, the similarity value is calculated using the maximum-minimum similarity calculation method to form a fuzzy matrix;

[0066] The fuzzy matrix R is: R = (r ij ), 1≤i, j≤n;

[0067] in:

[0068]

[0069] in:

[0070] r ij The value in the i-th row and j-th column of the fuzzy matrix;

[0071] x ik Let be the value in the i-th row and k-th column of matrix X;

[0072] x jk Let be the value in the j-th row and k-th column of matrix X.

[0073] Specifically, the method for step S5 is as follows:

[0074] For the fuzzy matrix R obtained in step S4, the transitive closure matrix is ​​calculated using the squaring method; the transitive closure matrix t(R) is: t(R) = R 2(m+1)

[0075] in:

[0076] m is the m-th calculation, and satisfies R 2(m+1) =R 2m

[0077] in:

[0078]

[0079] in:

[0080]

[0081] in:

[0082] ∧ indicates that the minimum value is taken from both sides;

[0083] ∨(…) indicates taking the maximum value within the parentheses.

[0084] Specifically, the method for step S6 is as follows:

[0085] S61: Obtain the clustering threshold. Based on the transitive closure matrix t(R) obtained in step S5, identify the different values ​​in the matrix. These values ​​are the clustering threshold λ.

[0086] S62: Calculate the λ-intercept matrix. Based on the clustering threshold λ obtained in step S61, construct the λ-intercept matrix using the λ-cut-set method.

[0087] λ-intercept matrix R λ for:

[0088] in:

[0089]

[0090] In R λ In the middle, if Then it is assumed that data points i and j belong to the same class at the same level.

[0091] S63: Draw dynamic clustering graphs based on R. λ Factors belonging to the same category under different λ values ​​are converged at a single point to form a dynamic clustering graph.

[0092] Specifically, the method for step S7 is as follows:

[0093] S71: Determine the optimal clustering threshold by comparing F statistics, calculating F statistics under different thresholds, and selecting the threshold corresponding to the largest F statistics as the optimal clustering threshold;

[0094] The statistic F is:

[0095] in:

[0096]

[0097]

[0098] in:

[0099] F λ The F-statistic is the value of the clustering threshold λ.

[0100] SSB is the sum of squared deviations between groups;

[0101] SSW is the within-group squared deviation;

[0102] c λ The number of categories under the clustering threshold λ;

[0103] n g Let c be the number of data in the g-th class, where g ≤ c λ ;

[0104] x gi For the i-th data in class g, i≤n g ;

[0105] x i This represents the i-th group of data, or the i-th row of data;

[0106] n is the total number of samples, which is the number of combinations of environmental factors in step S32;

[0107] S72: Select the appropriate number of categories, based on the determined optimal threshold, and in combination with the interval framework, etc.;

[0108] When selecting the number of categories, the following principles should be met:

[0109] Principle 1: When there is no gap frame constraint, and the number of categories is required to be greater than or equal to 2, the optimal number of categories is selected as the final number of categories;

[0110] Principle 2: When there are gap frames, the optimal number of categories should be selected as the final number of categories, provided that the number of categories is greater than or equal to the number of gap frames.

[0111] Principle 3: If the maximum number of categories is less than the number of gap frames, the conditions for using gap frames can be adjusted, some frames can be discarded, and the discarded frames can be merged into lower-level frames.

[0112] Specifically, the method of step S8 is as follows: according to the classification results selected in step S72, the output information corresponding to the combination of environmental factors identified in step S26 is set to the same value to form a rating matrix of the interval corresponding to the combination of environmental factors, so as to realize the development of the interval rating system. The purpose of the development is to formulate an interval determination rating system that is adapted to the current situation of the rail transit industry and can automatically output maintenance intervals by having analysts select environmental factors based on the analysis object.

[0113] This invention can identify factors affecting vehicle system, structure, and equipment damage based on existing maintenance records, operation records, and other input data, considering the current maintenance interval framework, and make similarity judgments. Finally, it can make reasonable and effective classifications, and develop a qualitative-to-quantitative preventive maintenance task interval rating system suitable for the current situation of the rail transit industry. This system can more reasonably and accurately formulate preventive maintenance task intervals for structures and areas, reduce over-maintenance and missed maintenance, and improve the online rate, thus achieving cost reduction and efficiency improvement more accurately and effectively. Attached Figure Description

[0114] The invention will be further described with reference to the accompanying drawings, but the contents of the drawings do not constitute any limitation on the invention.

[0115] Figure 1 This is a flowchart of the development method of the maintenance interval rating system based on fuzzy clustering according to the present invention.

[0116] Figure 2 This is a schematic diagram of fault data recording in the maintenance interval rating system development method based on fuzzy clustering of the present invention.

[0117] Figure 3 This is an example diagram of the initial data matrix for the development method of the maintenance interval rating system based on fuzzy clustering of this invention.

[0118] Figure 4 This is a schematic diagram of the dynamic clustering diagram of the maintenance interval rating system development method based on fuzzy clustering of the present invention.

[0119] Figure 5 This is a schematic diagram of the rating matrix considering the number of interval frames in the development method of the maintenance interval rating system based on fuzzy clustering of the present invention.

[0120] Figure 6 This is a schematic diagram of the rating matrix used in the adjustment interval framework of the maintenance interval rating system development method based on fuzzy clustering of this invention. Detailed Implementation

[0121] The present invention will be further described in conjunction with the following embodiments.

[0122] Example 1:

[0123] like Figure 1-6 As shown, the development method of the maintenance interval rating system based on fuzzy clustering includes the following steps: The overall flowchart is as follows. Figure 1 As shown.

[0124] S1: Data collection, based on vehicle operation and maintenance records, to collect information on vehicle malfunctions occurring at different times, such as... Figure 2 As shown; at the same time, the current vehicle maintenance interval frame is collected, and the maintenance interval frame has been determined to be 1 month, 3 months, 6 months, 1 year, and 5 years.

[0125] S2: Process data. Based on the collected fault information, process the data according to the data usage requirements to obtain high-quality usable data.

[0126] The specific method for step S2 is as follows:

[0127] S21: Organize the data, collect fault records and maintenance records according to the characteristics of the fault, and obtain the initial standard data form;

[0128] S22: Filter the data, remove component failure data caused by human factors, quality reasons, etc. from the initial standard data form, and obtain a standard data form that can be used for analysis;

[0129] S23: Check for duplicate data, identify duplicate data in the integrated standard data form, and obtain the standard data form for duplicate checking;

[0130] S24: Data restoration. The missing information in the plagiarism check standard data form is restored to obtain the restored standard data form, which is divided into the following cases;

[0131] When the fault is missing, the vehicle has already accumulated mileage L. f However, the time T when the fault occurred is known. f The following restoration methods can be used:

[0132] If the vehicle's utilization rate (AOR) and the start date of its use (T0) are available, the following formula is used for reconstruction:

[0133] L f =(T f -T0)·AOR

[0134] If there is vehicle utilization rate (AOR) and the date T before the data collection deadline... n and the corresponding cumulative mileage L n The following formula is used for restoration:

[0135] L f =L n -(T n -T f )·AOR

[0136] If there is vehicle utilization rate (AOR) and other fault data information for vehicles in the same column as the fault data object, i.e., the fault time T1 and corresponding cumulative mileage L1 of another data entry, the following formula is used for restoration:

[0137] L f =L1-(T1-T f )·AOR

[0138] If there is no vehicle usage rate, but there are two other fault data points for the same vehicle in the same column as the fault data object, namely the fault times T1 and T2 of the other two data points, and the corresponding cumulative mileage L1 and L2, the following formula is used to restore the data:

[0139]

[0140] L f =L2-(T2-T) f )·AOR or L f =L1-(T1-T f )·AOR

[0141] If there is no vehicle usage rate, but there is other fault data information for the vehicle in the same column as the fault data object, namely the fault time T1 and the corresponding cumulative mileage L1 of another data, and the start date of use T0 is known, the following formula is used to restore it:

[0142]

[0143] L f =(T f -T0)·AOR

[0144] If there is other fault data information for the vehicle in the same column as the fault data object, i.e., the fault time T1 and the corresponding cumulative mileage L1 of another data entry, and also the date T of the data collection deadline. n and the corresponding cumulative mileage L n The following formula is used for restoration:

[0145]

[0146] L f =L n -(T n -T f )·AOR or L f=L1-(T1-T f )·AOR

[0147] If there is no vehicle usage rate, then there is a data collection deadline date T. n and the corresponding cumulative mileage L n Knowing the start date T0 of its use, the following formula is used for restoration:

[0148]

[0149] L f =L n -(T n -T f )·AOR

[0150] S25: Split the data. The same data record containing multiple components is split according to the component serial number to obtain a split standard data form.

[0151] S26: Identify data. For the split standard data form, identify the environmental factors and failure modes of each data item to obtain the identified standard data form.

[0152] S27: Calculate the lifespan of the faulty component. The lifespan of the component is calculated based on known data conditions, and is divided into the following modes;

[0153] Given the date of failure, the date of installation, and the vehicle's usage rate, the lifespan interval is calculated as follows:

[0154] Lifespan range = (Date of failure - Date of installation) × Vehicle usage rate

[0155] Given the vehicle's cumulative mileage at the time of the malfunction and the vehicle's cumulative mileage at the time the component was installed, the lifespan range is calculated as follows:

[0156] Lifespan range = Total mileage of vehicle when the failure occurred - Total mileage of vehicle when the component was installed. Given the failure date and installation date, the lifespan range can be directly calculated in calendar time as follows:

[0157] Lifespan range = Date of failure - Date of installation.

[0158] S3: Standardized data, which is the data obtained from S2 that has been preprocessed to standardize the data matrix.

[0159] The specific method for step S3 is as follows:

[0160] S31: Statistical analysis of fault data. Based on the results identified in step S26, statistically analyze the number of faults occurring under each environmental factor across different period ranges. The statistical results are shown below. Figure 3 ;

[0161] S32: Establish the initial data matrix. Compile the statistical results from step S31 into a matrix format to form the initial data matrix X = (x ij ), 1≤i≤n, 1≤j≤m;

[0162] in:

[0163] x ij The value in the i-th row and j-th column of the matrix represents the number of faults in the j-th cycle under the i-th factor.

[0164] n represents the number of combinations of environmental factors;

[0165] m is the set number of cycles;

[0166] The final initial data matrix is

[0167]

[0168] S33: Calculate the maximum value parameter by taking the maximum value of each column of matrix X;

[0169] Maximum value parameter M j For: M j =max(x 1j x 2j , ..., x nj ), 1≤j≤m;

[0170] The calculation yields M = (9, 8, 10, 8, 11).

[0171] S34: Calculate the normalized data matrix based on X determined in step S32 and the values ​​calculated in step S33.

[0172] M j Calculate the standardized data matrix;

[0173] The standardized data matrix is: X′=(x′ ij ), 1≤i≤n, 1≤j≤m

[0174] in:

[0175]

[0176] Calculated

[0177]

[0178] S4: Calculate the fuzzy matrix by using a similarity calculation method to calculate the fuzzy matrix from the standardized data matrix;

[0179] The specific method for step S4 is as follows:

[0180] For the standardized data matrix X′ obtained in step S34, the similarity value is calculated using the maximum-minimum similarity calculation method to form a fuzzy matrix;

[0181] The fuzzy matrix R is: R = (r ij ), 1≤i, j≤n;

[0182] in:

[0183]

[0184] in:

[0185] r ij The value in the i-th row and j-th column of the fuzzy matrix;

[0186] x ik Let be the value in the i-th row and k-th column of matrix X;

[0187] x jk Let be the value in the j-th row and k-th column of matrix X;

[0188] The fuzzy matrix R is obtained through calculation, where

[0189]

[0190] Similarly, it can be calculated that

[0191]

[0192] S5: Calculate the transitive closure by performing stepwise calculations on the fuzzy matrix using the flat method to obtain the transitive closure matrix.

[0193] The specific method for step S5 is as follows:

[0194] For the fuzzy matrix R obtained in step S4, the transitive closure matrix is ​​calculated by the squaring method.

[0195] The transitive closure matrix t(R) is: t(R) = R 2(m+1)

[0196] in:

[0197] m is the m-th calculation, and satisfies R 2(m+1) =R 2m

[0198] in:

[0199]

[0200] in:

[0201]

[0202] in:

[0203] ∧ indicates that the minimum value is taken from both sides;

[0204] ∨(…) indicates taking the maximum value within the parentheses;

[0205] Calculated

[0206]

[0207] in

[0208]

[0209] Similarly, other values ​​can be calculated, thus obtaining...

[0210]

[0211] Therefore, the transitive closure is obtained as follows:

[0212]

[0213] S6: Output a dynamic clustering graph. Based on the calculated transitive closure, draw the dynamic clustering graph. The specific method for step S6 is as follows:

[0214] S61: Obtain the clustering threshold. Based on the transitive closure matrix t(R) obtained in step S5, identify the different values ​​in the matrix. These values ​​are the clustering threshold λ.

[0215] Based on t(R) in step S5, the clustering thresholds λ = 1, 0.6969, 0.67, 0.5786, 0.5777 can be obtained.

[0216] S62: Calculate the λ-intercept matrix. Based on the clustering threshold λ obtained in step S61, construct the λ-intercept matrix using the λ-cut-set method.

[0217] λ-intercept matrix R λ for:

[0218] in:

[0219]

[0220] In R λ In the middle, if Then it is assumed that data points i and j belong to the same class at the same level.

[0221] The following R can be obtained through calculation. λ .

[0222]

[0223] S63: Draw dynamic clustering graphs based on R. λ Factors belonging to the same category under different λ values ​​are clustered together at a single point to form a dynamic clustering graph. See the results below. Figure 4 ;

[0224] S7: Select the classification result. Based on the dynamic clustering graph, select the final number of classifications according to the optimal number of classifications or the relevance principle.

[0225] The specific method for step S7 is as follows:

[0226] S71: Determine the optimal clustering threshold by comparing F statistics, calculating F statistics under different thresholds, and selecting the threshold corresponding to the largest F statistics as the optimal clustering threshold;

[0227] The statistic F is:

[0228] in:

[0229]

[0230] in:

[0231] F λ The F-statistic is the value of the clustering threshold λ.

[0232] SSB is the sum of squared deviations between groups;

[0233] SSW is the within-group squared deviation;

[0234] c λ The number of categories under the clustering threshold λ;

[0235] n g Let c be the number of data in the g-th class, where g ≤ c λ ;

[0236] x gi For the i-th data in class g, i≤n g ;

[0237] x i This represents the i-th group of data, or the i-th row of data;

[0238] n is the total number of samples, which is the number of combinations of environmental factors in step S32;

[0239] The matrix obtained based on step S32 is:

[0240]

[0241] The values ​​for the first set of data are X1 = 5 + 8 + 9 + 1 + 2 = 25. Similarly, X2 = 25, X3 = 33, X4 = 32, and X5 = 27. Therefore, the calculated mean is...

[0242]

[0243] When λ = 0.6969, the data is divided into 4 categories: {X1}, {X2}, {X3}, and {X4,X5}, with the following group means:

[0244]

[0245] The sum of squares of the deviations between groups is:

[0246] SSB λ=0.6969 = 1 × (25 - 28.4) 2 +1×(25-28.4) 2 +1×(33-28.4) 2 +2×(29.5-28.4) 2 = 11.56 + 11.56 + 21.16 + 2.42 = 46.7

[0247] The sum of squared deviations within a group is:

[0248] SSW λ=0.6969 =0+0+0+(32-29.5) 2 +(27-29.5) 2 =12.5

[0249] Therefore, the F-statistic can be calculated as follows:

[0250]

[0251] Similarly, when λ = 0.67, there are 3 categories: {X2}, {X1,X3}, and {X4,X5}. In this case, F... λ=0.67 =0.33; when λ = 0.5786, there are 2 classes, namely {X2} and {X1,X3,X4,X5}, and F in this case λ=0.5777 =0.97.

[0252] Therefore, the F value is maximized when λ = 0.6969, and the optimal number of classifications is 4.

[0253] S72: Select the appropriate number of categories, based on the determined optimal threshold, and in combination with the interval framework, etc.;

[0254] When selecting the number of categories, the following principles should be met:

[0255] Principle 1: When there is no gap frame constraint, and the number of categories is required to be greater than or equal to 2, the optimal number of categories is selected as the final number of categories;

[0256] Principle 2: When there are gap frames, the optimal number of categories should be selected as the final number of categories, provided that the number of categories is greater than or equal to the number of gap frames.

[0257] Principle 3: If the maximum number of categories is less than the number of gap frames, the conditions for using gap frames can be adjusted, some frames can be discarded, and the discarded frames can be merged into lower-level frames.

[0258] The current interval frames are 1 month, 3 months, 6 months, 1 year, and 5 years. As a constraint, they can be divided into 4 categories: retaining 1 month, 3 months, 6 months, and 1 year, while the 5-year frame is considered as 1 year in actual application. If the number of interval frames is not considered, they can be directly divided into 5 categories: 1 month, 3 months, 6 months, 1 year, and 5 years.

[0259] S8: Output the classification results and formulate the relationship matrix of the rating system according to the selected number of categories.

[0260] The specific method of step S8 is as follows: according to the classification results selected in step S72, the output information corresponding to the combination of environmental factors identified in step S26 is set to the same value to form a rating matrix of the interval corresponding to the combination of environmental factors, so as to realize the development of the interval rating system. The purpose of the development is to formulate an interval determination rating system that is adapted to the current situation of the rail transit industry and can automatically output maintenance intervals by having analysts select environmental factors based on the analysis object.

[0261] By outputting the data according to 5 categories and 4 categories respectively, we can obtain matrices for the two rating systems, see [link / reference]. Figure 5 and Figure 6 .

Claims

1. A method for developing a maintenance interval rating system based on fuzzy clustering, characterized in that: Includes the following steps: S1: Collect data based on vehicle operation and maintenance records, collect fault information of vehicles at different times, and collect the current vehicle maintenance interval framework. S2: Process data. Based on the collected fault information, process the data according to the data usage requirements to obtain high-quality usable data. S3: Standardized data, which is the data obtained from S2 that has been preprocessed to standardize the data matrix. S4: Calculate the fuzzy matrix by using a similarity calculation method to calculate the fuzzy matrix from the standardized data matrix; S5: Calculate the transitive closure by performing stepwise calculations on the fuzzy matrix using the flat method to obtain the transitive closure matrix. S6: Output a dynamic clustering graph, which is drawn based on the calculated transitive closure. S7: Select the classification result. Based on the dynamic clustering graph, select the final number of classifications according to the optimal number of classifications or the relevance principle. S8: Output the classification results and formulate the relationship matrix of the rating system according to the selected number of categories.

2. The development method of the maintenance interval rating system based on fuzzy clustering according to claim 1, characterized in that: The specific method for step S2 is as follows: S21: Organize the data, collect fault records and maintenance records according to the characteristics of the fault, and obtain the initial standard data form; S22: Filter the data, remove component failure data caused by human factors, quality reasons, etc. from the initial standard data form, and obtain a standard data form that can be used for analysis; S23: Check for duplicate data, identify duplicate data in the integrated standard data form, and obtain the standard data form for duplicate checking; S24: Data restoration. The missing information in the plagiarism check standard data form is restored to obtain the restored standard data form, which is divided into the following cases; When the fault is missing, the vehicle has already accumulated mileage L. f However, the time T when the fault occurred is known. f The following restoration methods can be used: If the vehicle's utilization rate (AOR) and the start date of its use (T0) are available, the following formula is used for reconstruction: L f =(T f -T0)·AOR If there is vehicle utilization rate (AOR) and the date T before the data collection deadline... n and the corresponding cumulative mileage L n The following formula is used for restoration: L f =L n -(T n -T f )·AOR If there is vehicle utilization rate (AOR) and other fault data information for vehicles in the same column as the fault data object, i.e., the fault time T1 and corresponding cumulative mileage L1 of another data entry, the following formula is used for restoration: L f =L1-(T1-T f )·AOR If there is no vehicle usage rate, but there are two other fault data points for the same vehicle in the same column as the fault data object, namely the fault times T1 and T2 of the other two data points, and the corresponding cumulative mileage L1 and L2, the following formula is used to restore the data: L f =L2-(T2-T) f )·AOR or L f =L1-(T1-T f )·AOR If there is no vehicle usage rate, but there is other fault data information for the vehicle in the same column as the fault data object, namely the fault time T1 and the corresponding cumulative mileage L1 of another data, and the start date of use T0 is known, the following formula is used to restore it: L f =(T f -T0)·AOR If there is other fault data information for the vehicle in the same column as the fault data object, i.e., the fault time T1 and the corresponding cumulative mileage L1 of another data entry, and also the date T of the data collection deadline. n and the corresponding cumulative mileage L n The following formula is used for restoration: L f =L n -(T n -T f )·AOR or L f =L1-(T1-T f )·AOR If there is no vehicle usage rate, then there is a data collection deadline date T. n and the corresponding cumulative mileage L n Knowing the start date T0 of its use, the following formula is used for restoration: L f =L n -(T n -T f )·AOR S25: Split the data. The same data record containing multiple components is split according to the component serial number to obtain a split standard data form. S26: Identify data. For the split standard data form, identify the environmental factors and failure modes of each data item to obtain the identified standard data form. S27: Calculate the lifespan of the faulty component. The lifespan of the component is calculated based on known data conditions, and is divided into the following modes; Given the date of failure, the date of installation, and the vehicle's usage rate, the lifespan interval is calculated as follows: Lifespan range = (Date of failure - Date of installation) × Vehicle usage rate Given the vehicle's cumulative mileage at the time of the malfunction and the vehicle's cumulative mileage at the time the component was installed, the lifespan range is calculated as follows: Lifespan = Total mileage of vehicle at the time of failure - Total mileage of vehicle at the time of component installation Given the date of the failure and the date of installation, the lifespan range can be directly calculated using the calendar date as follows: Lifespan range = Date of failure - Date of installation.

3. The development method of the maintenance interval rating system based on fuzzy clustering according to claim 2, characterized in that: The specific method for step S3 is as follows: S31: Statistical analysis of fault data. According to the results identified in step S26, statistical analysis of the number of faults occurring under each environmental factor distributed across different period ranges. S32: Establish the initial data matrix. Compile the statistical results from step S31 into a matrix format to form the initial data matrix X = (x ij ), 1≤i≤n, 1≤j≤m; in: x ij The value in the i-th row and j-th column of the matrix represents the number of faults in the j-th cycle under the i-th factor. n represents the number of combinations of environmental factors; m is the set number of cycles; S33: Calculate the maximum value parameter by taking the maximum value from each column of matrix X; Maximum value parameter M j For: M j =max(x 1j ,x 2j ,…,x nj ), 1≤j≤m; S34: Calculate the normalized data matrix, based on X determined in step S32 and the values ​​calculated in step S33. M j Calculate the standardized data matrix; The standardized data matrix is: X′=(x′ ij ), 1≤i≤n, 1≤j≤m in:

4. The development method of the maintenance interval rating system based on fuzzy clustering according to claim 3, characterized in that: The specific method for step S4 is as follows: For the standardized data matrix X′ obtained in step S34, the similarity value is calculated using the maximum-minimum similarity calculation method to form a fuzzy matrix; The fuzzy matrix R is: R = (r ij ), 1≤i,j≤n; in: in: r ij The value in the i-th row and j-th column of the fuzzy matrix; x ik Let be the value in the i-th row and k-th column of matrix X; x jk Let be the value in the j-th row and k-th column of matrix X.

5. The development method of the maintenance interval rating system based on fuzzy clustering according to claim 4, characterized in that: The specific method for step S5 is as follows: For the fuzzy matrix R obtained in step S4, the transitive closure matrix is ​​calculated by the squaring method. The transitive closure matrix t(R) is: t(R) = R 2(m+1) in: m is the m-th calculation, and satisfies R 2(m+1) =R 2m in: in: in: ∧ indicates that the minimum value is taken from both sides; ∨(…) indicates taking the maximum value within the parentheses.

6. The development method of the maintenance interval rating system based on fuzzy clustering according to claim 5, characterized in that: The specific method for step S6 is as follows: S61: Obtain the clustering threshold. Based on the transitive closure matrix t(R) obtained in step S5, identify the different values ​​in the matrix. These values ​​are the clustering threshold λ. S62: Calculate the λ-intercept matrix. Based on the clustering threshold λ obtained in step S61, construct the λ-intercept matrix using the λ-cut-set method. λ-intercept matrix R λ for: in: In R λ In the middle, if Then it is assumed that data points i and j belong to the same class at the same level; S63: Draw dynamic clustering graphs based on R. λ Factors belonging to the same category under different λ values ​​are converged at a single point to form a dynamic clustering graph.

7. The development method of the maintenance interval rating system based on fuzzy clustering according to claim 6, characterized in that: The specific method for step S7 is as follows: S71: Determine the optimal clustering threshold by comparing F statistics, calculating F statistics under different thresholds, and selecting the threshold corresponding to the largest F statistics as the optimal clustering threshold; The statistic F is: in: in: F λ The F-statistic is the value of the clustering threshold λ. SSB is the sum of squared deviations between groups; SSW is the within-group squared deviation; c λ The number of categories under the clustering threshold λ; n g Let c be the number of data in the g-th class, where g ≤ c λ ; x gi For the i-th data in class g, i≤n g ; x i This represents the i-th group of data, or the i-th row of data; n is the total number of samples, which is the number of combinations of environmental factors in step S32; S72: Select the appropriate number of categories, based on the determined optimal threshold, and in combination with the interval framework, etc.; When selecting the number of categories, the following principles should be met: Principle 1: When there is no gap frame constraint, and the number of categories is required to be greater than or equal to 2, the optimal number of categories is selected as the final number of categories; Principle 2: When there are gap frames, the optimal number of categories should be selected as the final number of categories, provided that the number of categories is greater than or equal to the number of gap frames. Principle 3: If the maximum number of categories is less than the number of gap frames, the conditions for using gap frames can be adjusted, some frames can be discarded, and the discarded frames can be merged into lower-level frames, and finally the lower-level frames can be retained.

8. The method for developing a maintenance interval rating system based on fuzzy clustering according to claim 7, characterized in that: The specific method of step S8 is as follows: according to the classification results selected in step S72, the output information corresponding to the combination of environmental factors identified in step S26 is set to the same value to form a rating matrix of the interval corresponding to the combination of environmental factors, so as to realize the development of the interval rating system.