Telemetering data anomaly detection method and anomaly detection equipment based on satellite telemetering system
By building a large language model and customized data processing flow, the low efficiency and insufficient accuracy of traditional satellite telemetry data anomaly detection methods are solved, and efficient and accurate anomaly detection of satellite telemetry data is achieved, which is suitable for telemetry data of different scales and types.
Patent Information
- Application Number
- CN202510590943.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional satellite telemetry data anomaly detection methods have difficulty accurately identifying anomaly types and locating abnormal parameters when processing high-dimensional, long time series and complexly correlated telemetry data. Existing methods are inefficient and the detection accuracy needs to be improved, especially when dealing with complex correlations between multiple parameters.
Build a large language model, combine the characteristics of satellite telemetry data, and realize automated anomaly detection through data preprocessing, model training and evaluation, including data cleaning, standardization, feature extraction, parameter correlation analysis and model tuning, to generate anomaly detection benchmark patterns and use the adaptive learning ability of the large language model to identify unknown abnormal patterns.
It achieves efficient and accurate anomaly detection of satellite telemetry data, reduces manual intervention, improves data processing efficiency and model generalization ability, and adapts to telemetry data of different scales and types.
Smart Images

Figure CN120671028A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of satellite communication technology, and in particular relates to a method and device for detecting anomalies in telemetry data based on a satellite telemetry system. Background Art
[0002] With the rapid development of satellite applications, satellite telemetry systems play a vital role in aerospace, meteorological monitoring, communications, and navigation. These systems collect large amounts of parameter data through telemetry equipment, providing crucial support for satellite operational status monitoring and health assessment. However, telemetry data often has high dimensionality, long time series, and complex correlations. Traditional anomaly detection methods face significant challenges in processing this data, particularly in accurately identifying anomaly types and locating abnormal parameters. Furthermore, anomalies in telemetry data often involve more than just single parameter anomalies; they may also involve complex correlations between multiple parameters. Existing methods are inefficient in handling such correlations, and detection accuracy needs to be improved. Therefore, utilizing advanced data processing and machine learning techniques to efficiently and accurately detect anomalies in telemetry data remains a hot topic and a key challenge in current research. Summary of the Invention
[0003] This application aims to provide an effective telemetry data anomaly detection technology. By building and training a large language model and combining it with the characteristics of satellite telemetry data, it can efficiently detect anomalies in the data and ensure data accuracy and reliability. The method of the present invention includes multiple steps, including telemetry data preprocessing, model training and evaluation, and anomaly detection. The method also implements automated anomaly detection through an anomaly detection device.
[0004] On one hand, the present application provides a method for detecting anomalies in telemetry data based on a satellite telemetry system, the method comprising: S1, preprocesses the telemetry data provided by the satellite telemetry system to generate customized data; S2, generating a training data set, a test data set, and a validation data set based on the customized data; S3, building a large language model based on the training data set, training and evaluating the large language model, and generating a tested model; S4. Based on the detected model, perform anomaly prediction on the verification data set to obtain an anomaly detection result corresponding to the telemetry data.
[0005] Furthermore, step S1 includes the following sub-steps: S11, converting the telemetry data into structured data; S12, performing data cleaning and post-cleaning processing on the structured data; S13, performing standardization processing on the structured data after data cleaning and processing to generate standardized data; S14, generating the customized data according to the standardized data; S15, checking whether there is dirty data in the customized data, if yes, re-execute the sub-steps S12 to S15, if no, execute the step S2.
[0006] Furthermore, in the sub-step S12, The data cleaning of the structured data includes deduplication processing, missing value interpolation processing, and outlier removal or correction of the structured data; The cleaning and post-processing of the structured data includes splicing the cleaned structured data into a preset data format.
[0007] Furthermore, the sub-step S14 includes: extracting features related to anomaly detection from the standardized data; According to the features related to anomaly detection, the standardized data are arranged in a preset customized order to generate the customized data.
[0008] Furthermore, step S3 includes: Analyzing the correlation between several parameters of preset categories in the training data set; Calculating the importance weights of the plurality of parameters through correlation analysis between the parameters and constructing a parameter relationship diagram between the plurality of parameters; generating an anomaly detection benchmark pattern according to the parameter relationship graph; According to the anomaly detection benchmark mode, the several parameters are tuned to train the large language model.
[0009] Furthermore, the step S2 further includes: generating a test data set according to the customized data; The step S3 further comprises: The large language model is evaluated using the test dataset.
[0010] Furthermore, the step S3 further includes: Selecting evaluation indicators from the test data set to evaluate the anomaly detection performance of the trained large language model; Based on the evaluation result of the anomaly detection performance, determine whether the large language model needs to be trained again; if so, re-execute the S3 step; if not, generate the detected model.
[0011] Furthermore, the step S4 includes: Inputting the validation data set into the tested model to perform prediction calculations; According to the prediction calculation result and in combination with the preset abnormality standard, determining whether there are abnormal data points in the verification data set; If there are abnormal data points, the abnormal data points are classified and analyzed, and an abnormality detection result is generated based on the classification analysis result.
[0012] In a second aspect, the present application provides a telemetry data anomaly detection device based on a satellite telemetry system, which is used to implement the above-mentioned telemetry data anomaly detection method based on a satellite telemetry system. The system includes: Data customization module, used to pre-process the telemetry data provided by the satellite telemetry system and generate customized data; A data distribution module is used to generate a training data set, a test data set and a validation data set according to the customized data; A model training module, configured to construct a large language model based on the training data set, train and evaluate the large language model, and generate a tested model; A result acquisition module is used to perform anomaly prediction on the verification data set based on the detected model to obtain an anomaly detection result corresponding to the telemetry data.
[0013] In a third aspect, the present application provides a computer-readable storage medium, in which a computer program is stored, and the computer program is used to execute the above-mentioned telemetry data anomaly detection method based on a satellite telemetry system.
[0014] Compared to existing technologies, this application's advantage lies in its efficient and accurate anomaly detection in satellite telemetry data, achieved by combining a large language model with a customized data processing pipeline. This method does not rely on a predefined rule base, but instead automatically learns and adapts to complex relationships in the data, enabling it to identify unknown anomalous patterns and reducing manual intervention. Furthermore, the automated data processing pipeline improves data processing efficiency, and the scalability of the large language model enables it to adapt to telemetry data of varying sizes and types, resulting in stronger generalization and higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 This is a flowchart of a method for detecting anomalies in telemetry data based on a satellite telemetry system provided by an embodiment of the present application; Figure 2 This is a functional module diagram of a telemetry data anomaly detection device based on a satellite telemetry system provided by an embodiment of the present application. DETAILED DESCRIPTION
[0017] Specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. It is apparent that the described embodiments are only some of the embodiments of the present application, and not all of them. All other embodiments derived by persons of ordinary skill in the art based on the description of this application without inventive effort are intended to fall within the scope of protection of this application.
[0018] Example 1 One embodiment of the present application provides a method for detecting anomalies in telemetry data based on a satellite telemetry system. Figure 1 , the method comprises the following steps: S1: Preprocess the telemetry data provided by the satellite telemetry system to generate customized data.
[0019] In one embodiment, telemetry data, typically generated by various sensors onboard satellites, contains a large amount of complex time series data and multi-dimensional parameters. Raw data may contain noise, missing values, outliers, and other issues. Therefore, data preprocessing is necessary to improve data quality and meet the requirements of subsequent operations. Data preprocessing is a key step in improving the accuracy of subsequent anomaly detection and can include a series of operations such as data cleaning, standardization, and feature extraction.
[0020] In one embodiment, step S1 includes: S11: Convert the telemetry data into structured data.
[0021] Alternatively, raw telemetry data is often time-series data, containing multi-dimensional sensor outputs. First, this time-series data needs to be converted into a structured data format. For example, multiple sensor data points (such as temperature, voltage, and attitude angle) at each time point can be consolidated into structured data in a tabular format for easier analysis.
[0022] S12: performing data cleaning and post-cleaning processing on the structured data.
[0023] Optionally, after converting raw data into structured data, data cleaning is a core step in preprocessing. Data cleaning includes operations such as deduplication, missing value interpolation, and outlier removal or correction to improve data accuracy and completeness.
[0024] In one embodiment, the specific steps of S12 are as follows: Deduplication: Remove duplicate records to ensure data uniqueness and accuracy.
[0025] Missing value interpolation: For missing data, interpolation methods (such as linear interpolation, mean interpolation, etc.) are used to fill the missing parts to reduce the impact of data loss on model training.
[0026] Outlier removal or correction: Detect and remove outliers that do not fit within the normal data range, or correct unreasonable data points through correction methods. For example, correction is performed based on the relationship with adjacent data points to ensure data rationality.
[0027] After data cleaning, post-processing steps are performed. Here, the parameters, data, status, and status-related parameters in the structured data are concatenated into a specific input format for model processing. The specific input format is input = parameter + data + status + status-related parameters. For example, if there are five parameters, the value of parameter 3 is 50, and the anomaly is caused by parameters 4 and 5, then input = parameter 3, 50, anomaly, parameter 4, parameter 5. The final model can predict the masked word in the sentence based on the input (that is, we need to know the parameter part of the anomaly). This customized data processing process optimizes the validity and consistency of satellite telemetry data, ensuring that the model can identify anomalous parameters and features. After data cleaning, processed data is generated.
[0028] S13: Standardizing the structured data after data cleaning and processing to generate standardized data.
[0029] Optionally, to eliminate the impact of different parameter dimensions on data analysis, the processed data needs to be normalized. Normalization converts parameters of different dimensions to a unified scale to ensure that all parameters have equal weight during training and prevent certain features from dominating the model due to excessive scale.
[0030] S14: Generate the customized data according to the standardized data.
[0031] Optionally, after normalization, features relevant to anomaly detection can be further extracted. For example, statistical features (such as mean and standard deviation) at different time steps can be extracted from time series data and arranged in a pre-set order to generate customized datasets. These datasets provide rich input information for subsequent anomaly detection models.
[0032] S15: Check whether there is dirty data in the customized data. If so, re-execute the sub-steps S12 to S15; if not, execute step S2.
[0033] Optionally, after generating customized data, dirty data may still exist. This dirty data can be caused by factors such as data input errors and processing errors, so automated methods are needed to detect and identify dirty data. If dirty data is found, the cleaning and standardization process is re-executed; if the data is correct, proceed to the next step.
[0034] S2: Generate a training dataset and a validation dataset based on the customized data.
[0035] In one embodiment, after data preprocessing, data splitting is required. After splitting, three datasets are generated: a training dataset, a test dataset, and a validation dataset. Each dataset plays a different role at different stages of the model, as follows: Training dataset: used to train a large language model, enabling the model to learn the characteristics and patterns of telemetry data and understand the relationship between different parameters.
[0036] Test dataset: used to evaluate the performance of the model, verify the accuracy of the model on known data, and ensure that the model can effectively identify known abnormal patterns.
[0037] Validation dataset: used to predict anomalies on real data and generate anomaly detection results after model training is completed and the tested model is generated.
[0038] The purpose of dataset generation is to ensure that large language models can be trained and evaluated on different datasets, avoiding overfitting and improving the generalization ability of the model.
[0039] S3: Build a large language model based on the training data set, train and evaluate the large language model, and generate a tested model.
[0040] In one embodiment, after completing data preprocessing and dataset generation, the next step is to use this data to build and train a model. By deeply analyzing and learning from the training dataset, the large language model can learn the complex multi-dimensional parameter relationships in telemetry data, effectively solving the problem that traditional threshold methods and expert systems cannot identify complex anomalies. The specific steps of S3 are as follows: Parameter Correlation Analysis: Before model training, the correlation between various parameters in the training dataset must be analyzed. This step aims to reveal the interrelationships between different telemetry parameters. By calculating the correlations between parameters, we can generate a parameter relationship graph. This graph shows which parameters have close relationships and which parameters may be key to anomaly detection. The resulting parameter relationship graph serves as an important reference for subsequent model training, helping the model better understand the complex multi-dimensional relationships in the data.
[0041] Generate anomaly detection baseline patterns: Based on the parameter relationship graph, a baseline pattern for anomaly detection is generated. This baseline pattern represents the behavior of normal data and serves as the learning target for the large language model. Using this pattern, the model can identify data that deviates from the normal range and thus identifies it as an anomaly.
[0042] Tuning the Large Language Model: After generating a baseline pattern, the large language model needs to be tuned based on the baseline pattern and parameter relationship graph. The goal of tuning is to enable the model to better capture the complex relationships between various parameters in the data. By adjusting the model's training parameters, learning rate, and other hyperparameters, the model can gradually learn the subtle differences between abnormal and normal patterns, thereby improving the accuracy of anomaly detection.
[0043] Model Evaluation and Tuning: After fine-tuning, large language models need to be evaluated on a test dataset to verify their performance on real-world data. Model evaluation includes calculating metrics such as accuracy, recall, F1 score, BLEU (Bilingual Evaluation Understudy), and ROUGE (Recall-Oriented Understudy for Gisting Evaluation). Recall and F1 score, in particular, provide a deeper understanding of the model's ability to identify anomalies. Recall and F1 score are important metrics for evaluating classification model performance, especially when working with imbalanced datasets, where they provide more meaningful information than accuracy alone. Recall measures the model's ability to identify all examples that are actually positive. In other words, recall focuses on the proportion of all true positive examples that the model correctly identifies. It reflects the model's under-detection rate, that is, the number of positive examples that the model has missed. The positive class refers to the category that the model needs to pay special attention to or identify in the task. It usually represents the object of "interest" or the target to be predicted. Positive class examples usually indicate the occurrence of a certain event or situation. The negative class is the opposite of the positive class and usually represents a category that is of no interest or does not require attention. Negative class examples represent the non-occurrence of positive class events, meaning that the model does not need to make a special prediction.
[0044] The recall calculation formula is:
[0045] in: True Positives (TP): The number of samples correctly classified as positive.
[0046] False Negatives (FN): The number of positive samples that are incorrectly classified as negative.
[0047] The F1 score is the harmonic mean of recall and precision, combining the model's precision and recall. Precision measures how many of the samples predicted as positive by the model are actually positive, while the F1 score attempts to strike a balance between these two metrics.
[0048] The calculation formula of F1 value is:
[0049] in: Precision is : , indicates how many samples predicted as positive are actually positive. False positives (FP) refer to the number of negative samples that are mistakenly classified as positive.
[0050] Recall is: , which represents the proportion of all positive classes successfully identified by the model.
[0051] BLEU is calculated based on precision and is primarily used to evaluate the quality of machine translation. BLEU scores machine-generated translations by comparing the degree of n-gram (n consecutive words) matching between the machine-generated translation and the reference translation. Higher BLEU values indicate better machine translation quality.
[0052] ROUGE is a set of evaluation metrics used primarily for text summarization and text generation tasks, particularly in automatic summary generation. ROUGE measures coverage by calculating the recall of n-grams in the generated text compared to the reference text, placing a greater emphasis on recall.
[0053] If the evaluation results do not meet expectations, adjust the model structure or training parameters based on the evaluation results and retrain. By evaluating the model's performance, we can detect whether the large language model is overfitting or underfitting, and further adjust the model based on the evaluation results.
[0054] S4: Based on the detected model, perform anomaly prediction on the verification data set to obtain an anomaly detection result corresponding to the telemetry data.
[0055] In one embodiment, after the large language model is fully trained, it enters the final anomaly detection phase, where the model predicts anomalies based on the data by inputting a validation dataset.
[0056] In one embodiment, the specific steps of S4 are as follows: Input validation dataset: Input the processed validation dataset into the trained large language model for prediction calculation.
[0057] Anomaly determination and classification analysis: Based on the model's prediction results and preset anomaly criteria (such as deviation threshold and anomaly level), the data point is determined to be anomaly. If a data point is identified as an anomaly, further classification analysis is performed, such as determining the type of anomaly (such as equipment failure or environmental interference).
[0058] Generate anomaly detection results: Based on the classification analysis results, generate the final anomaly detection results and output the results. The results include specific information about each abnormal data point, such as time, parameters, and anomaly type, to facilitate subsequent response measures.
[0059] Example 2 One embodiment of the present application provides a device for detecting anomalies in telemetry data based on a satellite telemetry system. Figure 2 , used to implement the telemetry data anomaly detection method based on a satellite telemetry system, the telemetry data anomaly detection device 100 based on a satellite telemetry system includes: Data customization module 101: used to preprocess telemetry data, generate customized data, and provide high-quality input data for subsequent training and anomaly detection.
[0060] Data distribution module 102: Generates training data set, test data set and validation data set according to customized data, and passes these data sets to corresponding modules for processing.
[0061] Model training module 103: Build a large language model based on the training data set, perform training and evaluation, and ensure that the model has accurate anomaly detection capabilities.
[0062] Result acquisition module 104: performs anomaly prediction on the verification dataset based on the trained large language model and generates a detection result.
[0063] Example 3 This application also relates to a computer-readable storage medium storing a computer program for executing the above-described method. This program, executable in a computer system, includes steps such as data preprocessing, model training, and anomaly detection. By executing this program, automatic anomaly detection of satellite telemetry data can be achieved, and anomaly detection results can be generated.
[0064] This application solves several key problems in traditional satellite telemetry data anomaly detection methods: first, traditional anomaly detection methods based on preset thresholds cannot identify complex abnormal relationships between multiple parameters; second, expert system-based methods have difficulty building an accurate and complete telemetry data rule base and cannot cope with undefined abnormal patterns; finally, physical model-based methods require a lot of time and expert experience to build models and are not suitable for complex large-scale systems.
[0065] Compared with the existing technology, the improvements of this application are reflected in: the use of a large language model for multi-parameter relationship learning can capture the complex dependencies between parameters in telemetry data, thereby identifying more accurate abnormal patterns; through the adaptive learning ability of the large language model, it can automatically discover unknown abnormal patterns from the data without relying on a pre-defined rule base; combined with customized data processing procedures, the automation level of anomaly detection is improved, the dependence on expert experience is reduced, and at the same time it has higher scalability and can adapt to the needs of complex and large-scale systems.
[0066] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims. It should be understood that the present disclosure is not limited to the structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for detecting anomalies in telemetry data based on a satellite telemetry system, characterized in that: The following steps are involved: S1, preprocesses the telemetry data provided by the satellite telemetry system to generate customized data; S2, generating a training data set and a validation data set based on the customized data; S3, building a large language model based on the training data set, training and evaluating the large language model, and generating a tested model; S4. Based on the detected model, perform anomaly prediction on the verification data set to obtain an anomaly detection result corresponding to the telemetry data.
2. The method for detecting anomalies in telemetry data based on a satellite telemetry system according to claim 1, wherein: The step S1 includes the following sub-steps: S11, converting the telemetry data into structured data; S12, performing data cleaning and post-cleaning processing on the structured data; S13, performing standardization processing on the structured data after data cleaning and processing to generate standardized data; S14, generating the customized data according to the standardized data; S15, checking whether there is dirty data in the customized data, if yes, re-execute the sub-steps S12 to S15, if no, execute the step S2.
3. The method for detecting anomalies in telemetry data based on a satellite telemetry system according to claim 2, wherein: In the sub-step S12, The data cleaning of the structured data includes deduplication processing, missing value interpolation processing, and outlier removal or correction of the structured data; The cleaning and post-processing of the structured data includes splicing the cleaned structured data into a preset data format.
4. The method for detecting anomalies in telemetry data based on a satellite telemetry system according to claim 2, wherein: The sub-step S14 includes: extracting features related to anomaly detection from the standardized data; According to the features related to anomaly detection, the standardized data are arranged in a preset customized order to generate the customized data.
5. The method for detecting anomalies in telemetry data based on a satellite telemetry system according to claim 1, wherein: The step S3 comprises: Analyzing the correlation between several parameters of preset categories in the training data set; Calculating the importance weights of the plurality of parameters through correlation analysis between the parameters and constructing a parameter relationship diagram between the plurality of parameters; generating an anomaly detection benchmark pattern according to the parameter relationship graph; According to the anomaly detection benchmark mode, the several parameters are tuned to train the large language model.
6. The method for detecting anomalies in telemetry data based on a satellite telemetry system according to claim 5, wherein: The step S2 further includes: generating a test data set according to the customized data; The step S3 further comprises: The large language model is evaluated using the test dataset.
7. The method for detecting anomalies in telemetry data based on a satellite telemetry system according to claim 6, wherein: The step S3 further comprises: Selecting evaluation indicators from the test data set to evaluate the anomaly detection performance of the trained large language model; Based on the evaluation result of the anomaly detection performance, determine whether the large language model needs to be trained again; if so, re-execute the S3 step; if not, generate the detected model.
8. The method for detecting anomalies in telemetry data based on a satellite telemetry system according to claim 1, wherein: The step S4 comprises: Inputting the validation data set into the tested model to perform prediction calculations; According to the prediction calculation result and in combination with the preset abnormality standard, determining whether there are abnormal data points in the verification data set; If there are abnormal data points, the abnormal data points are classified and analyzed, and an abnormality detection result is generated based on the classification analysis result.
9. A telemetry data anomaly detection device based on a satellite telemetry system, characterized in that: The device is used to implement the method for detecting anomalies in telemetry data based on a satellite telemetry system according to any one of claims 1 to 8, comprising: Data customization module, used to pre-process the telemetry data provided by the satellite telemetry system and generate customized data; A data distribution module, configured to generate a training data set and a validation data set based on the customized data; A model training module, configured to construct a large language model based on the training data set, train and evaluate the large language model, and generate a tested model; A result acquisition module is used to perform anomaly prediction on the verification data set based on the detected model to obtain an anomaly detection result corresponding to the telemetry data.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 8.