A contact net defect internal cause analysis method based on Cox model

By constructing an intrinsic cause analysis model for overhead contact line defects using the Cox model and an improved partial likelihood function, this model solves the problems of relying on operation and maintenance experience and limited data in existing technologies. It enables multi-dimensional and time-dimensional defect analysis, improving the scientific rigor and accuracy of the analysis.

CN115757409BActive Publication Date: 2026-05-29CHENGDU ZHIGU YUNXING INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU ZHIGU YUNXING INFORMATION TECH CO LTD
Filing Date
2022-11-21
Publication Date
2026-05-29

Smart Images

  • Figure CN115757409B_ABST
    Figure CN115757409B_ABST
Patent Text Reader

Abstract

The application discloses a contact net defect internal cause analysis method based on a Cox model, relates to the technical field of contact net defect cause analysis of track traffic, and comprises the following steps: acquiring a historical contact net defect detailed record data table and a contact net defect internal cause factor detailed data table; pre-processing data in the defect detailed record data table and the internal cause factor detailed data table; using the pre-processed data to construct a proportional hazards model of defect occurrence and each internal cause factor data; using regression coefficients of the proportional hazards model to calculate an influence weight of each internal cause factor on the defect occurrence and a defect occurrence probability of the contact net. The contact net defect internal cause analysis method provided by the application not only considers a multi-dimension and multi-factor condition of internal cause factors, but also brings a time factor of defect discovery into the model, so that the whole analysis process is more comprehensive and scientific; and an improved partial likelihood function is adopted, so that calculation resources are saved and calculation accuracy is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of analysis of the causes of defects in rail transit overhead contact lines, and more specifically to a method for analyzing the intrinsic causes of defects in overhead contact lines based on the Cox model. Background Technology

[0002] The overhead contact line is a crucial power supply facility for the entire rail transit system. During train operation, high-speed relative motion occurs between the pantograph on the train roof and the contact line. This relative motion subjects the entire contact line system to significant impacts, making it prone to frequent defects in various parts. In contact line maintenance, analysts typically need to investigate the underlying causes of these defects (i.e., internal factor analysis). Effective internal factor analysis can guide the development of improvement plans for contact line operation, thereby increasing the efficiency of contact line maintenance.

[0003] Existing defect root cause analysis primarily relies on defect records and the professional experience of overhead contact line maintenance personnel, employing simple data analysis. It mainly includes the following steps: Step 1, analyzing defect record data; Step 2, proposing hypotheses about the causes of the defects based on the actual situation; Step 3, changing the actual state of the relevant causative factors; Step 4, re-examining whether the changes to the causative factors affect the occurrence of the defects; Step 5, determining whether there is an impact. If yes, obtain the root cause analysis conclusion; if no, return to Step 2.

[0004] This method has the following drawbacks: First, it relies excessively on the professional experience of operations and maintenance personnel. Second, the data source is too singular, resulting in an insufficiently comprehensive range of defect analysis dimensions. Third, it cannot perform defect analysis from a time perspective. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this invention discloses a method for analyzing the intrinsic causes of overhead contact line defects based on an improved Cox model. The purpose of this invention is to address the problems of existing technologies that rely on defect records and the professional experience of overhead contact line maintenance personnel, and involve simple data analysis, resulting in insufficient comprehensive defect analysis dimensions and an inability to analyze defects from a temporal perspective. This invention adopts survival analysis from statistics as its overall approach, using the Cox model and its improved partial likelihood function as tools. Through multi-dimensional factor analysis of intrinsic factors such as manufacturing factors, design factors, and maintenance factors, it achieves the analysis of the intrinsic causes of overhead contact line defects.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] A method for analyzing the intrinsic causes of defects in overhead contact lines based on the Cox model includes the following steps:

[0008] 1. Data Acquisition

[0009] S1. Obtain a detailed record table of historical overhead contact line defects and a detailed data table of internal factors causing overhead contact line defects;

[0010] (1) Obtain detailed historical defect record data table

[0011] During the operation of the overhead contact system, a detailed record table can be collected, which includes information such as the time of occurrence, location (e.g., support number and anchor section number), and handling time of each type of electrical or mechanical defect.

[0012] In this invention, historical detailed records of overhead contact line defects are obtained and recorded in a detailed overhead contact line defect record data table. Before building the model, detailed defect record data is obtained through the collection and processing of historical overhead contact line defect record data for model training.

[0013] (2) Obtain detailed data table of historical internal factors

[0014] Preferably, the internal factors include design factors, manufacturing factors, construction factors, operation and maintenance factors, environmental cumulative factors, and other factors.

[0015] The internal causes of overhead contact line defects can be categorized into six dimensions: design factors, manufacturing factors, construction factors, operation and maintenance factors, environmental cumulative factors, and other factors (other factors may include haze, fog, strong winds, hail, heavy snow, heavy rain, freezing rain, frost, and lightning). Each dimension corresponds to a data table. Each dimension's data table records detailed factor-related values ​​for each location (support post or anchor section). Each dimension's data table must at least include the support post number or anchor section number, and the corresponding factor-related values ​​for that dimension.

[0016] In this invention, detailed data on the internal causes of historical catenary defects can be obtained and recorded in a detailed data table of internal causes of catenary defects. Based on the nature of the internal factors, a corresponding data collection plan is developed to acquire the data; for example, temperature can be collected using temperature sensors, and vehicle speed can be obtained by searching vehicle operation data.

[0017] 2. Data Preprocessing

[0018] S2. Preprocess the data in the detailed defect record data table and the detailed internal factor data table;

[0019] In the above steps, the data in the detailed defect record data table and the detailed internal factor data table (hereinafter referred to as the two defect internal factor data tables) are preprocessed, and the preprocessed data is used as modeling data to facilitate subsequent modeling. The preprocessing includes the following steps:

[0020] (1) Handling missing values

[0021] S21. Handle the missing values ​​in the two data tables for the intrinsic causes of defects;

[0022] Preferably, in the missing value handling, the missing value handling threshold is 'a'. If the proportion of missing data for a certain factor to the total data of a field is greater than or equal to 'a', then the field is deleted; if the proportion is less than 'a', then the mean of the field is used to replace the missing values ​​of the field.

[0023] In the missing value handling steps described above, the missing value handling threshold is set to 'a'. Missing value handling is performed on the data field of each factor. If the proportion of missing data for a factor to the total data in a field is greater than or equal to 'a', then the field is deleted; if the proportion is less than 'a', then the missing values ​​in the field are replaced with the mean of the field.

[0024] The purpose of the missing value handling in this invention is to process the data field of each factor to obtain complete data that can be used for modeling.

[0025] (2) Joint query of two data tables for internal causes of defects

[0026] S22. After handling missing values, use the support number and anchor segment number as the associated fields to perform a joint query in the two data tables of defect internal factors to obtain the internal factor data of the defect location.

[0027] In this invention, the joint query of two tables refers to obtaining the internal factor data of the defect location by using the support number and anchor segment number as the associated fields. The query result table includes at least the support number and anchor segment number, the defect name, the discovery time, and the relevant values ​​of each factor.

[0028] In this invention, the purpose of jointly querying the two data tables of defect internal causes is to associate the support number with the anchor segment number in order to obtain the internal factor data of the defect location.

[0029] (3) Data standardization

[0030] S23. Standardize the data for each acquired internal factor.

[0031] In the data standardization step, data for each internal factor is standardized. Specifically, the standardization of the guide height design value among the internal factors is as follows:

[0032]

[0033] in, To standardize high design values, This represents the original value of the guide height design value corresponding to support i. The arithmetic mean of the high design values ​​among all modeling samples. The standard deviation is denoted as .

[0034] In this invention, the data of each acquired internal factor is standardized, which can control the gradient from changing drastically during the modeling process and make the subsequent calculation of the impact weight of defects more accurate.

[0035] (4) Constructing label data and time representations

[0036] S24. After standardization, use the standardized data to generate modeling data, and add a defect occurrence label field and time characterization quantity to the modeling data.

[0037] In this invention, the generated modeling data includes standardized internal factor data.

[0038] Preferably, in the defect occurrence label field, if a defect occurs at a certain location, the label value at that location is 1; if no defect occurs, the label value at that location is 0.

[0039] Preferably, among the time representation quantities, if the analysis is performed on the same line, the time between the defect discovery time and the most recent maintenance time of each support pillar is counted and used as the time representation quantity; if the analysis is performed on different lines, the number of trains passing each support pillar before the defect discovery time of each support pillar is counted and used as the time representation quantity.

[0040] In this invention, the purpose of adding label data is to quickly verify whether the constructed Cox model has achieved the target effect, and the purpose of adding time characterization data is to quickly determine the duration of the defect occurrence.

[0041] 3. Cox model based on improved likelihood function

[0042] S3. Using the preprocessed data, construct a proportional risk model of defect occurrence and various internal factors.

[0043] In this invention, the total amount of pillar data is assumed to be N. The total number of internal factors involved in the modeling is P. A proportional hazards model (i.e., a Cox model) is constructed using the modeling data to correlate defect occurrence with the data of each internal factor.

[0044] Preferably, the proportional risk model is:

[0045]

[0046] in, The proportional risk value output by the model. These are the regression coefficients of the model. As an internal factor, It is the basic risk proportion function.

[0047] In this invention, the basic risk proportion function can be obtained by querying relevant national or industry standards. If no relevant data is available in the standard documents, it can be obtained through nonparametric statistics using the current modeling sample, as shown in the following formula:

[0048] Preferably, the basic risk proportion function is:

[0049]

[0050] in,

[0051] ,

[0052] I represents the label data, and z represents the observed pillar position. Let be the set of pillar locations where defects occur after a time period t. Let be the set of all observed pillar positions after a time dimension t. For set The total number of central pillar positions.

[0053] Preferably, the regression coefficients of the model for coefficient at maximum The value, for:

[0054]

[0055] in, It is the partial likelihood function; z represents the set of pillar locations where defects occur after a time dimension t; z and m represent the sets... One of the pillar positions inside; Let l be the set of all observed pillar positions after a time dimension t; l represents the set. One of the pillar positions inside; express The number of pillar positions in the set The time when the defect occurs at position m of the support column, and h is the time from 0 to Card(D) t A loop iterating through 1.

[0056] In this invention, the regression coefficients of the model The calculation employs the maximum likelihood method. The maximum likelihood method obtains the coefficient values ​​by maximizing the partial likelihood function corresponding to the Cox model. Generally, the Cox model coefficients are estimated using the Breslow partial likelihood function; however, due to the large number of dimensions in the data on the internal factors of the overhead contact system, the computational workload is significant, and higher accuracy is required. Therefore, this invention chooses the Efron partial likelihood function, which offers higher computational accuracy. The Efron partial likelihood function is as follows:

[0057]

[0058] Logarithmic transformation of the above Efron partial likelihood function yields the above... Then, using the Newton-Raphson algorithm, we find the function that makes... coefficient at maximum The model coefficients obtained in this modeling process are as follows:

[0059]

[0060] Preferably, when constructing the model, after obtaining the model coefficient values, the Concordance Index is used to check whether the constructed model has achieved the target effect. When the Concordance Index is less than or equal to 0.5, the model is completely invalid, and when it is equal to 1, the model prediction is completely correct.

[0061] In this invention, after obtaining the model coefficient values, the Concordance Index (C-index) is used to verify whether the constructed Cox model has achieved the target effect. The C-index ranges from 0 to 1. When the C-index is less than or equal to 0.5, it indicates that the Cox model is completely ineffective. When the C-index is equal to 1, it indicates that the Cox model's predictions are completely correct.

[0062] Preferably, the threshold of the Concordance Index is set before building the model. After each model is built, the Concordance Index value corresponding to that model is calculated using the modeling data. ,like If the model and its coefficients are not found, then the model and its coefficients are retained. The purpose of this step is to verify whether the Cox model constructed in this study has achieved the target effect.

[0063] 4. Output the results of the internal cause analysis of defects.

[0064] S4. Using the regression coefficients of the proportional hazards model, calculate the weight of each internal factor on the occurrence of defects, as well as the probability of defects occurring in the overhead contact line.

[0065] After establishing the Cox model, the regression coefficients of the Cox model are used to calculate the influence weight of each intrinsic factor on the occurrence of defects. Assuming the total number of intrinsic factors involved in the modeling is P, the regression coefficients after establishing the Cox model are... After normalizing the regression coefficients of the Cox model, the influence weights of each intrinsic factor on the occurrence of defects are obtained.

[0066] Preferably, the influence weight is:

[0067]

[0068] in, The weight of the influence of internal factor i on the occurrence of missing terms. is the regression coefficient of the model corresponding to internal factor i, and P is the total number of internal factors involved in the modeling.

[0069] In this invention, the purpose of calculating the influence weight of each of the above-mentioned internal factors on the occurrence of defects is to obtain the internal factor with the greatest correlation to the occurrence of defects.

[0070] Preferably, the probability of a defect occurring at this support location is:

[0071]

[0072] in, This represents the probability of a defect occurring. For survival function, Among all the pillars studied, the pillar with the largest time characterization t corresponds to the time characterization quantity. Based on the basic risk proportion function, These are the regression coefficients of the model. This data is standardized based on internal factors. In this invention, the pillar is a geographical marker of the defect location; other markers such as anchor segments or kilometer markers can also be used.

[0073] In this invention, in order to predict the probability of a defect occurring in a certain support z after a time characterization t, the following two steps are required.

[0074] The first step is to collect data values ​​of the relevant intrinsic factors of pillar k, that is, the intrinsic factors corresponding to pillar k. The value of is determined, and the data is standardized using the mean and standard deviation of each factor. The standardized data is denoted as . .

[0075] The second step is to obtain the survival function using the following formula:

[0076]

[0077] The survival function expresses the probability that the support structure has not developed a defect after a time interval t. Subtracting the survival function value from 1 gives the probability that the support structure has developed a defect after time interval t, as shown in the following formula:

[0078]

[0079] in, The baseline survival function is constructed using the basic risk proportional function and the Kaplan-Meier algorithm. The construction process is as follows:

[0080]

[0081] Will Substituting the baseline survival function into the formula for the probability of a pillar defect as described above, we obtain the probability of a pillar defect as described in this invention.

[0082] The beneficial effects of this invention are:

[0083] 1. The data analysis in the existing technology is relatively crude. The contact network defect internal cause analysis method provided by the present invention not only considers the multi-dimensional and multi-factor situation of internal factors, but also incorporates the time factor of defect discovery into the model, making the whole analysis process more comprehensive and scientific.

[0084] 2. The contact wire defect intrinsic cause analysis method provided by the present invention adopts an improved partial likelihood function, which saves computing resources and ensures calculation accuracy.

[0085] 3. The contact wire defect intrinsic cause analysis method provided by the present invention uses normalized regression coefficients as the output of the intrinsic cause analysis process, and the results are more intuitive and more interpretable.

[0086] 4. The contact network defect intrinsic cause analysis method provided by the present invention standardizes the data of each factor during the data preprocessing process, which can control the gradient from changing drastically during the modeling process and ensure that the influence weight of defect occurrence is calculated more accurately.

[0087] 5. The contact wire defect intrinsic cause analysis method provided by the present invention obtains the influence weight of each intrinsic factor on the occurrence of defects by further normalizing the regression coefficients of the Cox model. Attached Figure Description

[0088] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0089] The following will provide a clear and complete description of the concept, specific structure, and technical effects of the present invention in conjunction with the embodiments and accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention.

[0090] Example 1

[0091] A method for analyzing the intrinsic causes of catenary defects based on the Cox model, such as Figure 1 As shown, it includes the following steps:

[0092] S1. Obtain a detailed record table of historical overhead contact line defects and a detailed data table of internal factors causing overhead contact line defects;

[0093] S2. Preprocess the data in the detailed defect record data table and the detailed internal factor data table;

[0094] S3. Using the preprocessed data, construct a proportional risk model of defect occurrence and various internal factors.

[0095] S4. Using the regression coefficients of the proportional hazards model, calculate the weight of each internal factor on the occurrence of defects, as well as the probability of defects occurring in the overhead contact line.

[0096] Example 2

[0097] This embodiment further elaborates on step S1 based on embodiment 1. Step S1 includes the steps of obtaining a detailed defect record data table and obtaining a detailed internal factor data table, as detailed below:

[0098] (1) Obtain detailed historical defect record data table

[0099] During the operation of the overhead contact system, a detailed record table can be collected, which includes information such as the time of occurrence, location (e.g., support number and anchor section number), and handling time of each type of electrical or mechanical defect. The table structure is shown in the table below.

[0100]

[0101] (2) Obtain detailed data table of historical internal factors

[0102] The internal causes of overhead contact line defects can be categorized into six dimensions: design factors, manufacturing factors, construction factors, operation and maintenance factors, environmental cumulative factors, and other factors. Each dimension corresponds to a data table. Each dimension's data table records detailed factor-related values ​​for each location (support post or anchor section). Each dimension's data table must at least include the support post number or anchor section number, and the corresponding factor-related values ​​for that dimension.

[0103] Taking the design factor table as an example, the structure of the table is shown in the following table:

[0104]

[0105] Example 3

[0106] This embodiment further elaborates on step S2 based on embodiment 2. Step S2 includes the following steps:

[0107] S21. Handle the missing values ​​in the two data tables for the intrinsic causes of defects;

[0108] S22. Using the support number and anchor segment number as the associated fields, perform a joint query in the two data tables of defect internal causes to obtain the internal factor data of the defect location;

[0109] S23. Standardize the data for each acquired internal factor.

[0110] S24. Use the standardized data to generate modeling data, and add the defect occurrence label field and time characterization quantity to the modeling data.

[0111] The S2 step described above includes data preprocessing, joint querying of the two data tables on defect causes, data standardization, and construction of label data and time representation metrics, as detailed below:

[0112] (1) Handling missing values

[0113] Let the threshold for handling missing values ​​be 'a'. Missing value handling is performed on the data field of each factor. If the proportion of missing data for a factor to the total data in a field is greater than or equal to 'a', then the field is deleted; if the proportion is less than 'a', then the missing values ​​in the field are replaced with the mean of the field.

[0114] (2) Joint query of two data tables for internal causes of defects

[0115] A two-table join query refers to retrieving internal factor data related to the location of a defect by using the support number and anchor segment number as the join fields. The query result table should at least contain the support number and anchor segment number, defect name, discovery time, and relevant values ​​for each factor. Its table structure can be illustrated as follows:

[0116]

[0117] (3) Data standardization

[0118] The data for each internal factor is standardized. Specifically, the standardization of the guide height design value among the internal factors is as follows:

[0119]

[0120] in, To standardize high design values, This represents the original value of the guide height design value corresponding to support i. The arithmetic mean of the high design values ​​among all modeling samples. The standard deviation is denoted as .

[0121] (4) Constructing label data and time representations

[0122] Finally, a defect occurrence label field and a time representation are added to the generated modeling data. If a defect occurs at a certain location, the label value for that location is 1; if no defect occurs, the label value for that location is 0.

[0123] Time-related metrics can be represented in two ways. If the analysis is conducted on the same route, the time between the time the defect was discovered and the time of the most recent maintenance for each support pillar can be counted and used as a time-related metric. If the analysis is conducted on different routes, the number of trains passing each support pillar before the time the defect was discovered can be counted.

[0124] The table structure for modeling data can be shown in the following table.

[0125]

[0126] Example 4

[0127] This embodiment further elaborates on step S3 based on embodiment 3, assuming the total amount of pillar data is N. The total number of internal factors involved in the modeling is P. A proportional hazards model (i.e., a Cox model) is constructed using the modeling data to correlate defect occurrence with the data of each internal factor.

[0128] Specifically, the proportional risk model is as follows:

[0129]

[0130] in, The proportional risk value output by the model. These are the regression coefficients of the model. As an internal factor, It is the basic risk proportion function.

[0131] In this embodiment, the basic risk proportion function can be obtained by querying relevant national or industry standards. If no relevant data is available in the standard file, it can be obtained using nonparametric statistics from the current modeling sample, with the calculation formula as follows:

[0132] The basic risk proportion function is:

[0133]

[0134] in,

[0135] ,

[0136] I represents the label data, and z represents the observed pillar position. Let be the set of pillar locations where defects occur after a time period t. Let be the set of all observed pillar positions after a time dimension t. For set The total number of central pillar positions.

[0137] In this embodiment, the regression coefficients of the model The calculation employs the maximum likelihood method. The maximum likelihood method obtains the coefficient values ​​by maximizing the partial likelihood function corresponding to the Cox model. Generally, the Breslow partial likelihood function is used for Cox model coefficient estimation; however, due to the large number of dimensions in the data on internal factors of the overhead contact system, the computational workload is significant, and higher accuracy is required. Therefore, this embodiment selects the Efron partial likelihood function, which offers higher computational accuracy. The Efron partial likelihood function is as follows:

[0138]

[0139] in, It is the partial likelihood function; z represents the set of pillar locations where defects occur after a time dimension t; z and m represent the sets... One of the pillar positions inside; Let l be the set of all observed pillar positions after a time dimension t; l represents the set. One of the pillar positions inside; express The number of pillar positions in the set The time when the defect occurs at position m of the support column, and h is the time from 0 to Card(D) t A loop iterating through 1.

[0140] Logarithmic transformation of the above Efron partial likelihood function yields:

[0141]

[0142] Then, using the Newton-Raphson algorithm, we find the function that makes... coefficient at maximum The model coefficients obtained in this modeling process are as follows:

[0143]

[0144] In this embodiment, after obtaining the model coefficient values, the Concordance Index (C-index) is used to verify whether the constructed Cox model has achieved the target effect. The C-index ranges from 0 to 1. When the C-index is less than or equal to 0.5, it indicates that the Cox model is completely ineffective. When the C-index is equal to 1, it indicates that the Cox model's predictions are completely correct.

[0145] Before building the model, first set the threshold for the Concordance Index. After each model is built, the C-index value corresponding to that model is calculated using the modeling data. ,like If so, then the model and its coefficients are retained.

[0146] Example 5

[0147] This embodiment further elaborates on step S4 based on embodiment 4. After the Cox model is established, the regression coefficients of the Cox model are used to calculate the influence weight of each intrinsic factor on the occurrence of defects. Assuming the total number of intrinsic factors involved in the modeling is P, and the regression coefficients after the Cox model is established are... .

[0148] After normalizing the regression coefficients of the Cox model, the influence weights of each intrinsic factor on the occurrence of the defect are obtained. Specifically, the influence weights are:

[0149]

[0150] in, The weight of the influence of internal factor i on the occurrence of missing terms. is the regression coefficient of the model corresponding to internal factor i, and P is the total number of internal factors involved in the modeling.

[0151] In this embodiment, in order to predict the probability of a defect occurring in a certain support z after a time characterization t, the following two steps are required.

[0152] The first step is to collect data values ​​of the relevant intrinsic factors of pillar k, that is, the intrinsic factors corresponding to pillar k. The value of is determined, and the data is standardized using the mean and standard deviation of each factor. The standardized data is denoted as . .

[0153] The second step is to obtain the survival function using the following formula:

[0154]

[0155] The survival function expresses the probability that the support structure has not developed a defect after a time interval t. Subtracting the survival function value from 1 gives the probability that the support structure has developed a defect after time interval t, as shown in the following formula:

[0156]

[0157] in, The baseline survival function is constructed using the basic risk proportional function and the Kaplan-Meier algorithm. The construction process is as follows:

[0158]

[0159] Will Substituting the baseline survival function into the formula for the probability of a pillar defect occurring, we obtain the probability of a pillar defect occurring in this embodiment:

[0160]

[0161] in, This represents the probability of a defect occurring. For survival function, Among all the pillars studied, the pillar with the largest time characterization t corresponds to the time characterization quantity. Based on the basic risk proportion function, These are the regression coefficients of the model. The data are standardized based on internal factors.

[0162] The embodiments of the present invention have been described in detail above, but the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention, and these equivalents or substitutions are all included within the scope defined by the claims of the present invention.

Claims

1. A method for analyzing the intrinsic causes of defects in overhead contact lines based on the Cox model, characterized in that, Includes the following steps: Obtain historical detailed records of overhead contact line defects and detailed data tables of internal factors causing overhead contact line defects; wherein, the detailed data of internal factors causing overhead contact line defects includes data of six dimensions, namely design factors, manufacturing factors, construction factors, operation and maintenance factors, environmental cumulative factors, and other factors; The data in the detailed defect record data table and the detailed internal factor data table are preprocessed, and the preprocessing includes the following steps: Handle missing values ​​in the two data tables for defect causes; Using the support number and anchor segment number as the associated fields, perform a joint query in the two data tables of defect internal factors to obtain the internal factor data of the defect location; The data for each acquired internal factor are standardized. Modeling data is generated using standardized data, and defect occurrence label fields and time representation quantities are added to the modeling data. When adding time representation quantities, if the analysis is conducted on the same route, the time between the defect discovery time and the most recent maintenance time of each support is counted and used as the time representation quantity. If the analysis is conducted on different routes, the number of trains passing each support before the defect discovery time is counted and used as the time representation quantity. Using the preprocessed data, a proportional hazards model is constructed to correlate defect occurrence with various internal factors. The proportional hazards model uses an improved partial likelihood function for parameter estimation. The improved partial likelihood function is the Efron partial likelihood function, which is used to process high-dimensional internal factor data and improve computational accuracy. Using the regression coefficients of the proportional hazards model, we calculate the influence weight of each intrinsic factor on the occurrence of defects, and calculate the probability of a defect occurring at a specific support location after a given time dimension.

2. The method for analyzing the intrinsic causes of catenary defects as described in claim 1, characterized in that, The detailed record data of the contact network defects includes relevant data on various types of electrical or mechanical defects. The relevant data includes the defect name, occurrence time, occurrence location, and handling time. The occurrence location includes the support column number and anchor section number.

3. The method for analyzing the intrinsic causes of catenary defects as described in claim 1, characterized in that, In missing value handling, the missing value handling threshold is 'a'. If the proportion of missing data for a certain factor to the total data of a field is greater than or equal to 'a', then the field is deleted; if the proportion is less than 'a', then the mean of the field is used to replace the missing values ​​of the field.

4. The method for analyzing the intrinsic causes of catenary defects as described in claim 1, characterized in that, In the standardization process, the standardization of the guide height design value among the internal factors is as follows: in, To standardize high design values, This represents the original value of the guide height design value corresponding to support i. The arithmetic mean of the high design values ​​among all modeling samples. Standard deviation; In the defect occurrence label field, if a defect occurs at a certain location, the label value for that location is 1; if no defect occurs, the label value for that location is 0.

5. The method for analyzing the intrinsic causes of catenary defects as described in claim 1, characterized in that, The proportional risk model is as follows: in, The proportional risk value output by the model. These are the regression coefficients of the model. As an internal factor, It is the basic risk proportion function.

6. The method for analyzing the intrinsic causes of catenary defects as described in claim 5, characterized in that, The basic risk proportion function is: in, , I represents the label data, and z represents the observed pillar position. Let be the set of pillar locations where defects occur after a time period t. Let be the set of all observed pillar positions after a time dimension t. For set The total number of central pillar positions.

7. The method for analyzing the intrinsic causes of catenary defects as described in claim 5, characterized in that, The model regression coefficients for coefficient at maximum The value, for: in, It is the partial likelihood function; z represents the set of pillar locations where defects occur after a time dimension t; z and m represent the sets... One of the pillar positions inside; Let l be the set of all observed pillar positions after a time dimension t; l represents the set. One of the pillar positions inside; express The number of pillar positions in the set The time when the defect occurs at position m of the support column, and h is the time from 0 to Card(D) t A loop iterating through 1.

8. The method for analyzing the intrinsic causes of catenary defects as described in claim 5, characterized in that, When building a model, after obtaining the model coefficient values, the Concordance Index is used to check whether the model has achieved the target effect. When the Concordance Index is less than or equal to 0.5, the model is completely invalid, and when it is equal to 1, the model prediction is completely correct. Before building the model, first set the threshold for the Concordance Index. After each model is built, the Concordance Index value corresponding to that model is calculated using the modeling data. ,like If so, then the model and its coefficients are retained.

9. The method for analyzing the intrinsic causes of catenary defects as described in claim 1, characterized in that, The influence weight is: in, The weight of the influence of internal factor i on the occurrence of missing terms. is the regression coefficient of the model corresponding to internal factor i, and P is the total number of internal factors involved in the modeling.

10. The method for analyzing the intrinsic causes of catenary defects as described in claim 1, characterized in that, The probability of a defect occurring at this support post location is as follows: in, This represents the probability of a defect occurring. For survival function, Among all the pillars studied, the pillar with the largest time characterization t corresponds to the time characterization quantity. Based on the basic risk proportion function, These are the regression coefficients of the model. The data are standardized based on internal factors.