A spacecraft anomaly correlation model training method based on on-orbit anomaly information
Through natural language processing and artificial intelligence technology, the classification and correlation model of spacecraft anomaly description information is solved, and the identification omissions and diagnosis errors in the correlation research of spacecraft anomaly information and space environment information in the existing technology are solved, and the accuracy and availability of correlation models are improved.
Patent Information
- Application Number
- CN202111491592.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The existing expert system has omissions in identification or diagnosis errors in the correlation between spacecraft anomaly information and space environment information, resulting in a decrease in the accuracy of the correlation model.
Natural language processing technology is used to classify the abnormal performance in the spacecraft anomaly description information database, and an correlation model between each anomaly and the spatial environment factor is established based on artificial intelligence technology to build an automatically generated more flexible and universal correlation model.
The accuracy and availability of the correlation model of spacecraft in-orbit anomalies and space environmental factors is improved, artificial errors are reduced, and the correlation between stand-alone and anomalies and complex space environmental factors is clarified.
Smart Images

Figure CN114330103B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of spacecraft space environment data information processing, and in particular to a spacecraft anomaly correlation model training method based on on-orbit anomaly information. Background Art
[0002] Large space agencies such as NASA and ESA have conducted a lot of correlation research based on manually evaluated spacecraft anomaly information and space environment information, and have formed relatively mature expert systems. For example, NASA currently uses the Spacecraft Anomaly Expert Real-time System (SEAES) and the SATCAST system launched by the Air Force Research Laboratory (AFR). However, the current expert system's recognition of satellite anomaly information is mainly based on professional engineers who conduct real-time evaluation of anomalies based on unified rules (categories). Therefore, there may be omissions or diagnostic errors in the evaluation process, which greatly reduces the accuracy of the correlation model. Summary of the invention
[0003] The purpose of the present invention is to provide a spacecraft anomaly correlation model training method based on on-orbit anomaly information, taking a spacecraft anomaly description information library as the analysis object, using natural language processing technology to classify the spacecraft anomaly performance in the anomaly description information library in a large amount of redundant information, and establishing a correlation model between each anomaly and space environmental factors based on artificial intelligence technology; it aims to effectively utilize spacecraft anomaly description information to construct a more flexible and universal correlation model that can be automatically generated, so as to provide corresponding guarantees and ground management references for the safe and reliable operation of spacecraft in the on-orbit space environment.
[0004] To solve the above technical problems, the present invention provides a technical solution: a method for training a spacecraft anomaly correlation model based on on-orbit anomaly information, and a method for constructing a spacecraft anomaly correlation model based on on-orbit anomaly description information, comprising:
[0005] (1) Build an information database of on-orbit anomaly description information;
[0006] (2) Generate a correlation model between multiple space environmental factors and spacecraft anomaly information;
[0007] In step (2), generating a correlation model includes:
[0008] 1) Construction of environment and anomaly fusion database;
[0009] 2) Construction of causal model;
[0010] 3) Construction of correlation model;
[0011] 4) Model training; for model training, a k-fold method is adopted, and constants k and n are set, where k>n>1. The time interval [0, T] is divided into k parts, and n of them are selected as training sets, and the rest are used as test sets. Secondly, the results of the training and test sets can be used to obtain quantitative indicators such as the accuracy, AUC, F1 score, and parameters of the selected model.
[0012] The present invention adopts the above technical scheme, takes the spacecraft anomaly description information base as the analysis object, uses natural language processing technology to classify the spacecraft anomaly performance in the anomaly description information base in a large amount of redundant information, and establishes a correlation model between each anomaly and space environmental factors based on artificial intelligence technology; it aims to effectively use the spacecraft anomaly description information to build a more flexible and universal correlation model that can be automatically generated, and provide corresponding guarantees and ground management references for the safe and reliable operation of the spacecraft in the orbital space environment. Further, the accuracy and availability of the correlation model between the spacecraft on-orbit anomaly and the space environmental factors are improved. Based on natural language processing and artificial intelligence technology, the present invention constructs a spacecraft anomaly description information base through word segmentation, clustering and other methods, and constructs a directed graph model of space environment and spacecraft anomaly based on Granger causality to generate a universal space environment and spacecraft risk correlation model, which improves the accuracy and availability of predicting the space environment anomaly that may be generated by the spacecraft through the space environmental factors on orbit.
[0013] Furthermore, in step (1), building an information base of on-orbit anomaly description information includes:
[0014] 1) Text preprocessing; 2) Dimensionality reduction of feature word set; 3) Phenomenon and single machine classification; 4) Repeat the previous 3 steps.
[0015] Furthermore, text preprocessing includes: proper noun replacement, word segmentation, and stop word removal;
[0016] Before text preprocessing, it is necessary to prepare a text set of spacecraft on-orbit anomaly description information and a proper noun vocabulary, and then perform proper noun replacement, that is, replace the words with the same meaning in all documents with the reference vocabulary to form a spacecraft anomaly description text set with a unified naming rule;
[0017] Then, the word segmentation process is performed. First, the words in the proper noun vocabulary are considered without being divided. Then, based on the word segmentation technology such as Jiaba, the text is processed into a set of feature words with words or phrases as the smallest unit;
[0018] Finally, stop words are removed, and punctuation, numbers, English, and function words are removed based on the feature word set in the previous step. Among them, before preprocessing, it is necessary to prepare a text set of spacecraft on-orbit anomaly description information and a proper noun vocabulary; this aspect is based on the information in the document library of spacecraft on-orbit anomaly description information, using the abnormal name, time, phenomenon description and other parts in the document, to capture the name, time, and phenomenon description information of each document to form a single text, and finally form a spacecraft on-orbit anomaly description information text set with a time mark. The proper noun vocabulary is manually constructed, so it can be specified in advance for the single-machine name table and other abnormal phenomenon vocabulary that need to be specified in advance. After the preprocessing is completed, the document becomes a feature word set with a time mark, a single-machine feature word, and one to several abnormal description feature words.
[0019] Furthermore, the dimension reduction of the feature word set includes: TF-IDF value dimension reduction of candidate feature words and principal component analysis dimension reduction;
[0020] Dimensionality reduction of candidate feature word TF-IDF value For a specific document and a specific word t, the TF-IDF formula of word t is expressed as follows:
[0021] V tgidf (t,d i )=f tf (t,d i )×f idf (t,d i );
[0022] Principal component analysis dimensionality reduction uses the PCA data dimensionality reduction algorithm, which transforms the high-dimensional original feature data into a set of linearly independent vector representations in each dimension through linear transformation, while retaining the variance of the data as much as possible.
[0023] TF-IDF (Term Frequency–Inverse Document Frequency) is a commonly used weighting technique for information retrieval and data mining. TF stands for term frequency, and IDF stands for inverse document frequency.
[0024] For a specific document and a specific word t, the TF-IDF formula of word t is as follows:
[0025] V tgidf (t,d i )=f tf (t,d i )×f idf (t,d i )
[0026] f tf (t,di ) is the frequency of the word in the text set, and the formula is as follows:
[0027]
[0028] in, is the number of word t in the text set, is the total number of words in the text set.
[0029] f idf (t,d i ) is the inverse document frequency, and its formula is as follows:
[0030]
[0031] Where D represents the total number of documents, M t Represents the number of documents containing word t.
[0032] The TF-IDF weight is used to remove unimportant feature words, and the high-dimensional and sparse feature words are reduced in dimension through principal component analysis to compress the feature information.
[0033] Furthermore, the phenomenon and single-machine classification includes: discrimination against word lists, manual correction and extraction of phenomenon word lists;
[0034] First, compare it with the single-machine name vocabulary constructed previously to determine whether it is a single-machine feature word. If it is not in the vocabulary, it is considered to be a phenomenon feature word.
[0035] Secondly, all phenomenon feature words were manually corrected, unimportant feature words were eliminated, and errors in segmentation judgment were manually corrected;
[0036] Finally, the corrected feature words are used as the phenomenon vocabulary, added to the proper noun vocabulary, and the text is segmented and dimension reduction is performed again.
[0037] Furthermore, the construction of the environment and anomaly fusion database includes: the construction of the space environment fusion database and the fusion of the environment and anomaly database;
[0038] Firstly, the continuous environmental factors in the time interval corresponding to the spacecraft anomaly database are selected, and the space environment fusion database is constructed with high-energy electron flux, geomagnetic index AE, solar activity index F10.7, heavy ion LET, high-energy proton flux, solar wind, interplanetary magnetic field and other environmental factors as multiple environmental factors;
[0039] Secondly, based on the interpolation and standardization processing method of unified time markers, the anomaly database and the environmental database are merged.
[0040] Furthermore, the causal model construction includes: Granger-like causal lag value calculation, factor screening and p-value test based on lag value and Granger-like causal model construction;
[0041] First, take the appropriate time characteristic value δt, extend the spatial environment data at time t0 to the interval [t0-δt, t0], and perform standardization to satisfy the value interval of [0, 1];
[0042] Afterwards, the Granger-like causal lag value between any two factors X and Y in the merged database is calculated based on indicators such as Pearson correlation coefficient, t-test, AIC, and BIC;
[0043] Finally, the factors of the above database are traversed and the causal results are constructed into a Granger-like causal directed graph.
[0044] Furthermore, the correlation model construction includes: extraction and screening of cause nodes and construction of the correlation model based on the machine learning model;
[0045] First, select specific abnormal factors from the spacecraft anomaly database, extract all the causes Xi (i∈1..N) of all factors, and extract their corresponding time lag values ti; when extracting, you can simultaneously view the correlation between these N nodes. If there is a relevant causal relationship, cut out the one with a higher p value or a larger time lag value;
[0046] Then, a correlation model is constructed. The model can be constructed using machine learning models such as generalized linear correlation model, random forest model, extreme tree model, etc. as the basic model. The relationship of the model is constructed as follows: t∈[0,T];
[0047] Afterwards, a preliminary training sampling of 70% of the data was performed for factor screening. The factors with weaker correlation were screened out through p-value or contribution weight and the variance was ensured to be optimal. A preliminary correlation model between untrained space environmental factors and typical spacecraft anomalies was obtained.
[0048] Furthermore, 1) the k-fold method is used to create training sets and test sets. When there are fewer abnormal data, different n values n1 and n0 can be used for the abnormal data (i.e., {Y|Y>0}) and the non-abnormal data (i.e., {Y|Y=0}), and k>n1>n0>1 to ensure the balance of the training set. 2) Indicators and models are selected for model testing. The results of the training and test sets can be used to obtain quantitative indicators such as the accuracy, AUC, F1 score, and parameters of the selected model. When the result is good (for example, the accuracy is greater than 0.8), the model can be accepted. Otherwise, other machine learning models can be replaced for testing. Finally, a correlation model between the space environment and spacecraft abnormal phenomena is obtained.
[0049] The present invention also achieves the following beneficial effects: compared with previous studies on the correlation between spacecraft anomalies and space environment, this method is based on a specific description of spacecraft anomalies rather than a simple anomaly classification, makes more effective use of spacecraft anomaly information, and eliminates the influence of human errors, which is beneficial to clarifying the correlation between the list machine and abnormal phenomena and complex space environmental factors, so that the established model of the correlation between the space environment and spacecraft anomalies has a higher accuracy rate and more information, which can provide a reference for design and zeroing. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a schematic diagram of constructing a correlation model based on on-orbit anomaly description information of the present invention;
[0051] Figure 2 It is a schematic diagram of the cause-effect model of the present invention. DETAILED DESCRIPTION
[0052] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0053] Reference Figure 1 to Figure 2 As shown, a method for training a spacecraft anomaly correlation model based on on-orbit anomaly information, and a method for constructing a spacecraft anomaly correlation model based on on-orbit anomaly description information, comprising:
[0054] (1) Build an information database of on-orbit anomaly description information;
[0055] (2) Generate a correlation model between multiple space environmental factors and spacecraft anomaly information;
[0056] In step (2), generating a correlation model includes:
[0057] 1) Construction of environment and anomaly fusion database;
[0058] 2) Construction of causal model;
[0059] 3) Construction of correlation model;
[0060] 4) Model training; for model training, a k-fold method is adopted, and constants k and n are set, where k>n>1. The time interval [0, T] is divided into k parts, and n of them are selected as training sets, and the rest are used as test sets. Secondly, the results of the training and test sets can be used to obtain quantitative indicators such as the accuracy, AUC, F1 score, and parameters of the selected model.
[0061] In step (1), building an information base of on-orbit anomaly description information includes:
[0062] 1) Text preprocessing; 2) Dimensionality reduction of feature word set; 3) Phenomenon and single machine classification; 4) Repeat the previous 3 steps.
[0063] In this embodiment, text preprocessing includes: proper noun replacement, word segmentation, and stop word removal;
[0064] Before text preprocessing, it is necessary to prepare a text set of spacecraft on-orbit anomaly description information and a proper noun vocabulary, and then perform proper noun replacement, that is, replace the words with the same meaning in all documents with the reference vocabulary to form a spacecraft anomaly description text set with a unified naming rule;
[0065] Then, the word segmentation process is performed. First, the words in the proper noun vocabulary are considered without being divided. Then, based on the Jiaba word segmentation technology, the text is processed into a set of feature words with words or phrases as the smallest unit.
[0066] Finally, stop words are removed. Based on the feature word set in the previous step, punctuation, numbers, English and function words are removed.
[0067] In this embodiment, the dimension reduction of the feature word set includes: TF-IDF value dimension reduction of candidate feature words and principal component analysis dimension reduction;
[0068] Dimensionality reduction of candidate feature word TF-IDF value For a specific document and a specific word t, the TF-IDF formula of word t is expressed as follows:
[0069] V tgidf (t,d i )=f tf (t,d i )×f idf (t,d i );
[0070] f tf (t,d i ) is the frequency of the word in the text set, and the formula is as follows:
[0071]
[0072] in, is the number of word t in the text set, is the total number of all words in the text set;
[0073] f idf (t,d i ) is the inverse document frequency, and its formula is as follows:
[0074]
[0075] Where D represents the total number of documents, M t Represents the number of documents containing word t;
[0076] Use TF-IDF weights to remove unimportant feature words, and use principal component analysis to reduce the dimension of high-dimensional and sparse feature words to compress feature information;
[0077] Principal component analysis dimensionality reduction uses the PCA data dimensionality reduction algorithm, which transforms the high-dimensional original feature data into a set of linearly independent vector representations in each dimension through linear transformation, while retaining the variance of the data as much as possible.
[0078] In this embodiment, the phenomenon and single machine classification includes: judging by comparing the word list, manually correcting and extracting the phenomenon word list;
[0079] First, compare it with the single-machine name vocabulary constructed previously to determine whether it is a single-machine feature word. If it is not in the vocabulary, it is considered to be a phenomenon feature word.
[0080] Secondly, all phenomenon feature words were manually corrected, unimportant feature words were eliminated, and errors in segmentation judgment were manually corrected;
[0081] Finally, the corrected feature words are used as the phenomenon vocabulary, added to the proper noun vocabulary, and the text is segmented and dimension reduction is performed again.
[0082] In this implementation, the construction of the environment and anomaly fusion database includes: construction of a spatial environment fusion database and fusion of the environment and anomaly database;
[0083] Firstly, the continuous environmental factors in the time interval corresponding to the spacecraft anomaly database are selected as multiple environmental factors to construct the space environment fusion database;
[0084] Secondly, based on the interpolation and standardization processing method of unified time markers, the anomaly database and the environmental database are merged.
[0085] In this implementation, the causal model construction includes: Granger-like causal lag value calculation, factor screening and p-value test based on the lag value, and Granger-like causal model construction;
[0086] First, take the appropriate time characteristic value δt, extend the spatial environment data at time t0 to the interval [t0-δt, t0], and perform standardization to satisfy the value interval of [0, 1];
[0087] Afterwards, the Granger-like causal lag value between any two factors X and Y in the merged database is calculated based on indicators such as Pearson correlation coefficient, t-test, AIC, and BIC;
[0088] Finally, the factors of the above database are traversed and the causal results are constructed into a Granger-like causal directed graph.
[0089] In this implementation, the correlation model construction includes: extraction and screening of cause nodes and construction of the correlation model based on a machine learning model;
[0090] First, select specific anomaly factors from the spacecraft anomaly database and extract all the causes X of all factors. i (i∈1..N), and extract its corresponding time lag value t i ; When extracting, you can check the correlation between these N nodes at the same time. If there is a relevant causal relationship, the one with a higher p value or a larger lag value will be cut off;
[0091] Then, a correlation model is constructed. The model can be constructed using machine learning models such as generalized linear correlation model, random forest model, extreme tree model, etc. as the basic model. The relationship of the model is constructed as follows:
[0092] Afterwards, a preliminary training sampling of 70% of the data was performed for factor screening. The factors with weaker correlation were screened out through p-value or contribution weight and the variance was ensured to be optimal. A preliminary correlation model between untrained space environmental factors and typical spacecraft anomalies was obtained.
[0093] 1) Use the k-fold method to create training sets and test sets. When there are fewer abnormal data, different n values n1 and n0 can be used for abnormal data (i.e., {Y|Y>0}) and non-abnormal data (i.e., {Y|Y=0}), and k>n1>n0>1 to ensure the balance of the training set. 2) Select indicators and models for model testing. The results of the training and test sets can be used to obtain quantitative indicators such as the accuracy, AUC, F1 score, and model parameters of the selected model. When the result is good (for example, the accuracy is greater than 0.8), the model can be accepted. Otherwise, other machine learning models can be replaced for testing. Finally, a correlation model between the space environment and spacecraft anomalies is obtained.
[0094] Model training can be divided into: 1) construction of training set and test set; 2) selection of indicators and models for model testing. First, the k-fold method is adopted, and constants k and n are set, where k>n>1. The time interval [0, T] is divided into k parts, and n of them are selected as training sets, and the rest are used as test sets. For the case of less abnormal data, the abnormal data (i.e., {Y|Y>0}) and the non-abnormal data (i.e., {Y|Y=0}) can be respectively assigned different n values n1 and n0, and k>n1>n0>1 to ensure the balance of the training set. Secondly, the results of the training and test sets can be used to obtain quantitative indicators such as the accuracy, AUC, F1 score of the selected model and the parameters of the model. When the result is good (for example, the accuracy is greater than 0.8), the model can be accepted, otherwise other machine learning models can be replaced for testing. Finally, a correlation model between the space environment and spacecraft abnormal phenomena is obtained.
[0095] Figure 1 The schematic diagram of the construction of the correlation model based on the on-orbit anomaly description information is shown in Figure 1. The method uses the on-orbit anomaly description information document library of spacecraft, which contains the anomaly name, occurrence time, phenomenon description and other parts. First, the on-orbit anomaly description information document library 1 is subjected to text preprocessing 3. The preprocessing is divided into three steps, including
[0096] 1) Replacement of proper nouns: Using the artificially constructed professional terminology vocabulary 2, replace the words with the same meaning in all documents with the reference vocabulary to form a set of spacecraft anomaly description texts with unified naming rules;
[0097] 2) Word segmentation: Based on the professional terminology vocabulary 2 and Jiaba word segmentation technology, the text is processed into a set of feature words with words or phrases as the smallest unit, and the magnetism is annotated;
[0098] 3) Remove stop words: Remove punctuation, numbers, English and function words (auxiliary words, adverbs, prepositions, conjunctions) based on magnetic annotation, and further remove them based on manually created stop word lists.
[0099] After the text preprocessing 3 is completed, the dimension of the feature word set is reduced 4 based on the calculation of TF-IDF value and principal component analysis, that is, words with too low word frequency or too high document frequency are ignored, and words that often appear together are combined into phrases. The magnetic strip after dimensionality reduction is compared with the single-machine word list in the professional terminology word list 2 to determine whether it is a single-machine feature word 5, and the feature words other than the single-machine feature word 5 are classified as phenomenon feature words 6. After the phenomenon feature word 6 is manually corrected 7, it is returned to the text preprocessing 3, dimensionality reduction 4 and compared with the single-machine word list in the professional terminology word list 2 as input for classification again, forming the final version of the single-machine feature word 5 and the phenomenon feature word 6. The finalized feature word is added with a time tag to form a spacecraft anomaly database 8. Each anomaly information in the database 8 contains time, single machine, and anomaly description.
[0100] After the construction of the spacecraft anomaly database 8 is completed, various space environment data sets 9 (including high-energy electron flux, geomagnetic index AE, solar activity index F10.7, heavy ion LET, high-energy proton flux, solar wind, interplanetary magnetic field and other data) in the corresponding time zone are selected to build a space environment fusion database 10. Based on the interpolation and standardization processing of the unified time mark 11, the unified time series is used as the benchmark for data cleaning, missing values and outliers are removed, and interpolation and standardization processing are performed, and finally a space environment and spacecraft anomaly information fusion database 12 containing single machine, anomaly description and space environment information is formed within the time range of [0, T].
[0101] Traverse any two factors in the space environment and spacecraft anomaly information fusion database 12 to construct a causal directed graph model 13, including the causal model of Bo Aohan's abnormal single machine and phenomenon, the causal model of abnormal single machine, and the causal model of abnormal phenomenon. The process of constructing the causal relationship between the two factors X(t) and Y(t) is as follows: Use indicators such as Pearson correlation coefficient, t test, AIC, BIC, etc. to construct indicators that can judge their correlation, and find a certain time tL that makes the sequence X(t-tL) and Y(t) have the highest correlation. The traversal process can be carried out in a sampling manner, setting the parameter k as a fixed value, randomly selecting the time tr and performing the above solution within the time period [tr,tr+round(T / k)] to obtain tL(r), and repeat the sampling multiple times to obtain the median or mean of tL(r) as the causal lag value t of factors X and Y. XY .t XY If the value exceeds a certain limit (such as the 27-day solar rotation period or the longest time for particle transport between the sun and the earth), or the p value in the test is large and the result is not significant, it is considered that there is no significant Granger-like causality between X and Y. XY The positive and negative phases of the two determine the causal relationship. If t XY is positive, then X is the cause of Y. If t XYIf it is negative, it means that Y is the cause of X. Based on the correlation between abnormal single machine and phenomenon and environment, abnormal single machine and environment, abnormal phenomenon and environment, and t XY Positive and negative phases, build causal models of abnormal single machines and phenomena, causal models of abnormal single machines, and causal models of abnormal phenomena. In the actual construction process, unnecessary factors can be cut to form the subgraphs described above. It is not necessary to traverse all pairs of factors to build causal relationships, or to perform causal relationship analysis only on the corresponding relationship between the environment and the anomaly.
[0102] Based on the causal directed graph model13, a correlation model between the space environment and the spacecraft abnormal information18 can be constructed. First, according to the abnormal single machine and phenomenon factor Y insphe 14. Abnormal single machine factor Y ins 15. Abnormal Phenomenon Factor Y phe 16 Extracting Factor All Causes X in Causal Model i (t)(i∈1..N) and its corresponding time lag value t i 17. Subsequently, machine learning models such as generalized linear correlation model, random forest model, extreme tree model, etc. were used as basic models to construct a correlation model between space environment and spacecraft abnormal information18, which includes the correlation model of abnormal single machine and phenomenon, the correlation model of abnormal single machine, and the correlation model of abnormal phenomenon.
[0103] After the correlation model is constructed 18, model training is performed 19. First, factor screening is performed, and 70% of the data is sampled. The factors with weaker correlation are screened out through p-value or contribution weight, and the variance is ensured to be optimal. Then the k-fold method is adopted, and constants k and n are set; where k>n>1, and the time interval [0, T] is divided into k parts, and n of them are selected as training sets, and the rest are used as test sets. The accuracy of the selected model can be obtained through the results of the training and test sets. When the result is good (for example, the accuracy is greater than 0.8), the model can be accepted, otherwise other machine learning models can be replaced for testing. Model training is performed on different anomalies, and finally a correlation model between the space environment and spacecraft anomalies is obtained 20.
[0104] In summary, the present invention has been made into actual samples and tested for multiple times as described in the specification and the drawings. From the results of the test, it can be proved that the present invention can achieve its intended purpose, and its practical value is beyond doubt. The above embodiments are only used to illustrate the present invention, and are not intended to limit the present invention in any form. Any person with ordinary knowledge in the technical field, if it does not depart from the scope of the technical features of the present invention, uses the equivalent embodiments of the technical content disclosed by the present invention to make partial changes or modifications, and does not depart from the technical features of the present invention, all still fall within the scope of the technical features of the present invention.
Claims
1. A method for training a spacecraft anomaly correlation model based on on-orbit anomaly information, characterized in that: The method of constructing a spacecraft anomaly correlation model based on on-orbit anomaly description information includes: (1) Build an information database of on-orbit anomaly description information; (2) Generate a correlation model between multiple space environmental factors and spacecraft anomaly information; In step (2), generating a correlation model includes: 1) Construction of environment and anomaly fusion database; 2) Construction of causal model; The causal model construction includes: Granger-like causal lag value calculation, factor screening and p-value test based on lag value and Granger-like causal model construction; First, take the appropriate time characteristic value δt, extend the spatial environment data at time t0 to the interval [t0-δt, t0], and perform standardization to satisfy the value interval of [0, 1]; Afterwards, the Granger-like causal lag value between any two factors X and Y in the merged database is calculated based on the Pearson correlation coefficient, t-test, AIC and BIC indicators; Finally, the factors in the environment and anomaly fusion database are traversed, and the causal results are constructed into a Granger-like causal directed graph; 3) Construction of correlation model; The correlation model construction includes: extraction and screening of cause nodes and construction of a correlation model based on a machine learning model; First, select specific anomaly factors from the spacecraft anomaly database and extract all the causes X of all factors. i , where i∈1..N, and extract its corresponding time lag value t i ; When extracting, you can simultaneously check the correlation between these N nodes. If there is a relevant causal relationship, the factors with high p-values or large lag values are trimmed; Then, a correlation model is constructed. The model can be constructed using machine learning models such as generalized linear correlation model, random forest model and extreme tree model as basic models. The relationship of the model is constructed as follows: t∈[0,T]; Among them, Y(t) is the abnormal data of the spacecraft at time t, ΣX i (tt i ) is the time lag value t i Corrected total environmental factors, ΣX i (t) is the environmental factor at time t; After that, 70% of the data were sampled for preliminary training to screen factors. The factors with weak correlation were screened out by p-value or contribution weight and the variance was ensured to be optimal. The correlation model between untrained space environment factors and typical spacecraft anomalies was initially obtained. 4) Model training; for model training, a k-fold method is adopted, and constants k and n are set, where k>n>1. The time interval [0, T] is divided into k parts, and n of them are selected as training sets, and the rest are used as test sets. Secondly, the accuracy, AUC, F1 score and parameters of the selected model can be obtained through the results of the training set and the test set.
2. The method for training a spacecraft anomaly correlation model based on on-orbit anomaly information according to claim 1, characterized in that: In step (1), building an information base of on-orbit anomaly description information includes: 1) Text preprocessing; 2) Dimensionality reduction of feature word set; 3) Phenomenon and single machine classification; 4) Repeat the previous 3 steps.
3. The spacecraft anomaly correlation model training method based on on-orbit anomaly information according to claim 2 is characterized in that: Text preprocessing includes: proper noun replacement, word segmentation, and stop word removal; Before text preprocessing, it is necessary to prepare a text set of spacecraft on-orbit anomaly description information and a proper noun vocabulary, and then perform proper noun replacement, that is, replace the words with the same meaning in all documents with the reference vocabulary to form a spacecraft anomaly description text set with a unified naming rule; Then, the word segmentation process is performed. First, the words in the proper noun vocabulary are considered without being divided. Then, based on the Jiaba word segmentation technology, the text is processed into a set of feature words with words or phrases as the smallest unit. Finally, stop words are removed. Based on the feature word set in the previous step, punctuation, numbers, English and function words are removed.
4. The method for training a spacecraft anomaly correlation model based on on-orbit anomaly information according to claim 3, characterized in that: The dimension reduction of feature word set includes: TF-IDF value dimension reduction of candidate feature words and principal component analysis dimension reduction; Dimensionality reduction of candidate feature word TF-IDF value For a specific document and a specific word t, the TF-IDF formula of word t is expressed as follows: V tgidf (t,d i )=f tf (t,d i )×f idf (t,d i ); Principal component analysis dimensionality reduction uses the PCA data dimensionality reduction algorithm, which transforms the high-dimensional original feature data into a set of linearly independent vector representations in each dimension through linear transformation, while retaining the variance of the data as much as possible.
5. The method for training a spacecraft anomaly correlation model based on on-orbit anomaly information according to claim 4, characterized in that: Phenomenon and single-machine classification include: discrimination against word lists, manual correction and extraction of phenomenon word lists; First, compare it with the single-machine name vocabulary constructed previously to determine whether it is a single-machine feature word. If it is not in the vocabulary, it is considered to be a phenomenon feature word. Secondly, all phenomenon feature words were manually corrected, unimportant feature words were eliminated, and errors in segmentation judgment were manually corrected; Finally, the corrected feature words are used as the phenomenon vocabulary, added to the proper noun vocabulary, and the text is segmented and dimension reduction is performed again.
6. The method for training a spacecraft anomaly correlation model based on on-orbit anomaly information according to claim 1, characterized in that: The construction of environment and anomaly fusion database includes: construction of space environment fusion database and fusion of environment and anomaly database; Firstly, the continuous environmental factors in the time interval corresponding to the spacecraft anomaly database are selected as multiple environmental factors to construct the space environment fusion database; Secondly, based on the interpolation and standardization processing method of unified time markers, the anomaly database and the environmental database are merged.
7. The method for training a spacecraft anomaly correlation model based on on-orbit anomaly information according to claim 1, characterized in that: 1) The k-fold method is used to create training sets and test sets. When there are few abnormal data, the abnormal data {Y|Y>0} and the non-abnormal data {Y|Y=0} can be taken as different n values n1 and n0, and k>n1>n0>1 to ensure the balance of the training set. 2) Indicators and models are selected for model testing. The accuracy, AUC, F1 score and parameters of the selected model can be obtained through the results of the training set and the test set. When the accuracy is greater than 0.8, the model can be accepted. Otherwise, other machine learning models can be replaced for testing. Finally, a correlation model between the space environment and spacecraft anomalies is obtained.
Citation Information
Patent Citations
Spacecraft space environment abnormity and influence forecast method and system, storage medium, and server
CN110017866A
Space environment risk index construction method and device and storage medium
CN111738604A