Intelligent identification method for abnormal state of meteorological numerical mode
Through principal component analysis and isolated forest methods, the abnormal state recognition of meteorological numerical modes is solved, and the problems of low efficiency and false alarms and misreports in the existing technology are achieved, and efficient and accurate meteorological abnormality detection and classification are achieved.
Patent Information
- Application Number
- CN202510071677.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
Existing meteorological anomaly detection technology is inefficient when dealing with complex multi-dimensional meteorological data, it is difficult to accurately identify abnormal states, it is prone to false alarms or misreports, and it is impossible to adapt to the changing meteorological environment in real time.
Principal component analysis is used to reduce the dimensionality of multi-dimensional meteorological numerical data, obtain the principal component feature set, and then gradually divide the meteorological sample set through the isolated forest method, identify high isolation points, form potential anomaly sample set, and abnormality marking and classification are performed through adaptive scoring criteria.
It improves the efficiency and calculation accuracy of meteorological data, ensures accurate identification of abnormal states, reduces false alarms and missed reports, can flexibly adapt to diversified meteorological scenarios, and realizes structured classification and sorting of meteorological abnormal data.
Smart Images

Figure CN119989221A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of meteorological anomaly detection, and in particular to an intelligent identification method for abnormal states of meteorological numerical models. Background Art
[0002] The field of meteorological anomaly detection technology aims to identify, judge and warn abnormal conditions in meteorological numerical models through intelligent algorithms and pattern recognition technologies. As an important part of meteorological forecasting and disaster prevention and control, it improves the detection accuracy and response speed of weather and climate anomalies, and provides effective support for meteorological forecasting and risk prevention.
[0003] The intelligent identification method of abnormal states of meteorological numerical models aims to accurately identify abnormal states from massive meteorological numerical data. It can quickly and accurately detect abnormal meteorological events, effectively reduce missed reports and false alarms, ensure that meteorological abnormal states are identified and processed in a timely manner, enhance the reliability and efficiency of meteorological warnings, and provide technical support for meteorological forecasts and disaster prevention decisions.
[0004] Existing meteorological anomaly detection technology has shortcomings when dealing with complex multi-dimensional meteorological data. Due to the high dimension of numerical data and the large number of characteristic variables, direct analysis will increase the complexity of data processing and affect data processing efficiency. When determining abnormal conditions, it relies on fixed thresholds and simple statistical methods, which makes it difficult to effectively deal with abnormal conditions under complex weather conditions and is prone to false alarms or omissions. The fixed threshold method cannot adapt to the changing meteorological environment in real time, and different meteorological events and states are easily confused. It is impossible to perform fine-grained layered processing on data isolation and multi-dimensional changes. There is a lack of dynamic adaptive detection and structured classification of potential abnormal data, resulting in a lack of precision and classification accuracy in meteorological anomaly identification results. Summary of the invention
[0005] The purpose of the present invention is to solve the shortcomings in the prior art and to propose an intelligent identification method for abnormal states of meteorological numerical models.
[0006] In order to achieve the above object, the present invention adopts the following technical solution: an intelligent identification method of abnormal state of meteorological numerical model, comprising the following steps:
[0007] Step 1: Based on multidimensional meteorological numerical data, principal component analysis is used to calculate the elements of the covariance matrix in turn, perform principal component screening operations on the matrix elements, select eigenvectors with larger variances, sort features according to variance values, map high-dimensional data to the low-dimensional space of principal component features, and obtain the principal component feature set;
[0008] Step 2: Based on the principal component feature set, the isolation forest is used to gradually segment the meteorological sample set, the segmentation times of each sample are calculated, the isolated data points appearing after segmentation are judged, high isolation points are identified according to the segmentation times, and a potential abnormal sample set is formed through the high isolation point set;
[0009] Step 3: Based on the potential abnormal sample set, each sample is segmented and deeply scored, and the score value is determined one by one through the adaptive scoring standard, and the samples with scores higher than the preset threshold are marked as abnormal, and the abnormal data score records are compiled to obtain the abnormal score list;
[0010] Step 4: Based on the anomaly score list, perform an anomaly marking operation on high-scoring samples, filter samples with scores exceeding a threshold, combine them into a meteorological anomaly identification set, and mark them to obtain a meteorological anomaly marking set;
[0011] Step 5: Based on the meteorological anomaly tag set, the abnormal samples are grouped according to meteorological characteristics, the data are marked as different meteorological types to form classification tags, the tag set is classified and summarized according to the meteorological feature category, and an abnormal type classification set is established.
[0012] As a further solution of the present invention, the specific steps of generating the principal component feature set are:
[0013] Based on multidimensional meteorological numerical data, principal component analysis is used to extract meteorological characteristic variables, calculate the covariance matrix item by item, perform covariance operation on each pair of characteristics, complete the matrix construction, and obtain the covariance matrix;
[0014] Based on the covariance matrix, the eigenvalue and eigenvector of each feature are calculated, and the eigenvector that can reflect the trend of data change is screened out by sorting the eigenvalues to obtain the principal component feature;
[0015] Based on the principal component features, the original data is linearly transformed through the principal component matrix, the high-dimensional data is projected into a low-dimensional space, the meteorological data after dimensionality reduction is obtained, and a principal component feature set is obtained;
[0016] The principal component feature set includes temperature features, humidity features and air pressure features after dimensionality reduction.
[0017] As a further solution of the present invention, the principal component analysis is performed according to the formula:
[0018]
[0019] Among them: Cov w (X,Y) represents the weighted covariance between variables X and Y, X and Y are meteorological characteristic variables, n is the total number of data samples, and w i is the weight coefficient of the i-th sample, Xi is the specific observation value of the i-th sample on variable X, Y i is the specific observation value of the i-th sample on variable Y, is the mean of variable X, is the mean of variable Y, α is the empirical adjustment coefficient, V x is the variance of variable X, V y is the variance of variable Y.
[0020] As a further solution of the present invention, the specific steps of generating the potential abnormal sample set are:
[0021] Based on the principal component feature set, the meteorological data is segmented, the path depth of each data point in the segmentation tree is calculated, the segmentation times of each data point are recorded, and the segmentation depth data is obtained;
[0022] Based on the segmentation depth data, an isolation forest is used to evaluate the isolation degree of each data point, and high isolation points are identified according to the number of segmentation times of the data point, and are screened to form a potential abnormal sample set;
[0023] Based on the potential abnormal sample set, extract data points with strong isolation and variation, and integrate them to obtain a potential abnormal sample set;
[0024] The potential abnormal sample set includes identified abnormal temperature points, abnormal humidity points and abnormal air pressure points.
[0025] As a further solution of the present invention, the isolation forest is according to the formula:
[0026]
[0027] Where: s(x) is the isolation score value, x is the specific data point to be evaluated, E(h(x)) is the expected value of the path length, d(x) is the distance from the data point x to the nearest neighbor, c(n) is the normalization constant, w1 is the isolation score adjustment weight, w2 is the path length weight, w3 is the distance impact weight, and w4 is the normalization coefficient adjustment weight.
[0028] As a further solution of the present invention, the method for calculating the change amplitude gradually calculates the change rate of the meteorological numerical model within a set time window, captures the fluctuation of the meteorological data, and uses a sliding window to calculate the sliding average of the change rate to obtain a stable fluctuation amplitude, which is convenient for identifying abnormal conditions. When the sliding average change rate exceeds a preset threshold, it is marked as an abnormal state.
[0029] As a further solution of the present invention, the specific steps of generating the abnormality score list are:
[0030] Based on the potential abnormal sample set, a deep segmentation calculation is performed on each sample, the path depth of the sample in the segmentation tree is calculated item by item, and the segmentation times of each sample are recorded to obtain a segmentation depth data set;
[0031] Based on the segmentation depth data set, an adaptive scoring standard is set, the segmentation depth of each sample is scored, the score value is compared with the set threshold, whether it exceeds the threshold is determined one by one, and a score value list is generated;
[0032] Based on the score value list, samples with scores higher than a set threshold are screened out, and abnormal marks are performed, and the abnormal scores are sorted and recorded to obtain an abnormal score list;
[0033] The abnormality score list includes the abnormality score of temperature, the abnormality score of humidity and the abnormality score of air pressure.
[0034] As a further solution of the present invention, the specific steps of generating the meteorological anomaly mark set are:
[0035] Based on the anomaly score list, samples with scores higher than a set threshold are screened, and samples are marked as abnormal one by one, and saved as abnormal data records to obtain a screened abnormal sample set;
[0036] Based on the screening abnormal sample set, the samples marked as abnormal are classified according to meteorological characteristics, and grouped according to temperature and humidity characteristics to obtain an abnormal meteorological feature set;
[0037] Based on the abnormal meteorological feature set, the abnormal samples are classified and sorted according to the marking status to generate a final abnormal identification data set to obtain a meteorological abnormal marking set;
[0038] The meteorological anomaly mark set includes temperature data, humidity data and air pressure data marked as abnormal.
[0039] As a further solution of the present invention, the specific steps of generating the abnormal type classification set are:
[0040] Based on the meteorological anomaly tag set, the abnormal samples are grouped one by one according to the meteorological characteristics, and the meteorological characteristic value of each sample is classified into a corresponding meteorological category to form a grouped sample set;
[0041] Based on the grouped sample set, each sample is marked with a meteorological type according to the meteorological feature category of each sample, and the classification mark of each sample is saved to obtain a classification mark set;
[0042] Based on the classification mark set, all samples marked as abnormal are further summarized and sorted according to meteorological types to obtain an abnormal type classification set;
[0043] The abnormal type classification set includes temperature mutation, humidity anomaly and air pressure fluctuation.
[0044] As a further solution of the present invention, samples are classified according to labels, abnormal samples are divided into multiple sets by type, the number, distribution characteristics and specific abnormal manifestations of samples are counted, and all classified abnormal types and statistical data are summarized to form an abnormal type classification set.
[0045] Compared with the prior art, the advantages and positive effects of the present invention are:
[0046] 1. In the present invention, the multidimensional meteorological numerical data is subjected to dimensionality reduction processing by principal component analysis, and the high-dimensional data is compressed into the principal component feature space, so that the core features of the data are retained, while redundant information is reduced, and the efficiency of data processing and the calculation accuracy are improved;
[0047] 2. In the present invention, the principal component feature set after dimensionality reduction is gradually segmented by the isolation forest method, which can effectively identify abnormal feature points in meteorological data, and locate the abnormal state in meteorological data by calculating the number of segmentations and isolation degree of each data point, thereby ensuring the accuracy of detection;
[0048] 3. In the present invention, potential abnormal samples are segmented and deeply scored, and an adaptive scoring standard is introduced to determine the score value one by one. Abnormal samples are marked according to preset thresholds to achieve accurate anomaly detection that is flexible and adaptable to diverse meteorological scenarios. Based on samples with scores higher than the threshold, anomaly marking and grouping classification are performed. By summarizing and organizing meteorological features, anomaly type classification sets are generated by category, so that abnormal data can be managed in a structured manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic diagram of the main steps of the present invention. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0051] Embodiment 1
[0052] See also Figure 1 The present invention provides a technical solution: an intelligent identification method for abnormal state of meteorological numerical model, comprising the following steps:
[0053] Step 1: Based on multidimensional meteorological numerical data, principal component analysis is used to calculate the elements of the covariance matrix in turn, perform principal component screening operations on the matrix elements, select eigenvectors with larger variances, sort features according to variance values, map high-dimensional data to the low-dimensional space of principal component features, and obtain the principal component feature set;
[0054] Step 2: Based on the principal component feature set, the isolation forest is used to gradually segment the meteorological sample set, calculate the number of segmentations of each sample, determine the isolated data points that appear after segmentation, identify high isolation points according to the number of segmentations, and form a potential abnormal sample set through the collection of high isolation points;
[0055] Step 3: Based on the potential abnormal sample set, each sample is segmented and deeply scored, and the score value is determined one by one through the adaptive scoring standard. Samples with scores higher than the preset threshold are marked as abnormal, and the abnormal data score records are compiled to obtain the abnormal score list;
[0056] Step 4: Based on the anomaly score list, perform anomaly marking operations on high-scoring samples, filter samples with scores exceeding the threshold, combine them into a meteorological anomaly identification set, and mark them to obtain a meteorological anomaly marking set;
[0057] Step 5: Based on the meteorological anomaly tag set, group the abnormal samples according to meteorological characteristics, mark the data into different meteorological types to form classification tags, classify and summarize the tag set according to meteorological feature categories, and establish an anomaly type classification set.
[0058] The specific steps to generate the principal component feature set are:
[0059] Based on multidimensional meteorological numerical data, principal component analysis is used to extract meteorological characteristic variables, calculate the covariance matrix item by item, perform covariance operation on each pair of characteristics, complete the matrix construction, and obtain the covariance matrix;
[0060] Based on the covariance matrix, the eigenvalues and eigenvectors of each feature are calculated. By sorting the eigenvalues, the eigenvectors that can reflect the trend of data changes are screened out to obtain the principal component features.
[0061] Based on the principal component characteristics, the original data is linearly transformed through the principal component matrix, the high-dimensional data is projected into the low-dimensional space, the meteorological data after dimensionality reduction is obtained, and the principal component feature set is obtained;
[0062] Based on multidimensional meteorological numerical data, the principal component analysis method is used to process meteorological characteristic variables. Specifically, the covariance matrix is first constructed, the covariance of each pair of characteristic variables is calculated item by item, the covariance operation is performed and the matrix construction process is completed. The covariance matrix is obtained by calculating the covariance between each meteorological characteristic variable item by item.
[0063] Based on the covariance matrix, the eigenvalues and eigenvectors of the matrix are calculated. Specifically, the eigenvalue decomposition method is used to extract the eigenvalues and eigenvectors. The eigenvalues and eigenvectors are calculated item by item according to the row-column relationship in the matrix. By sorting by the eigenvalue size, the eigenvectors that can effectively reflect the trend of meteorological data changes are screened out to generate the principal component features.
[0064] Based on the principal component characteristics, the original meteorological data is linearly transformed. Specifically, the linear mapping method is used to project the data to a low-dimensional space through the principal component matrix, and a representation based on the principal component is constructed in the low-dimensional space. The obtained principal component matrix is used to project the high-dimensional meteorological data to the low-dimensional space to obtain the meteorological data after dimensionality reduction, and the principal component feature set is generated;
[0065] Among them, the principal component feature set includes temperature features, humidity features and air pressure features after dimensionality reduction.
[0066] Principal component analysis, according to the formula:
[0067]
[0068] Among them: Cov w (X,Y) represents the weighted covariance between variables X and Y, X and Y are meteorological characteristic variables, n is the total number of data samples, and w i is the weight coefficient of the i-th sample, X i is the specific observation value of the i-th sample on variable X, Y i is the specific observation value of the i-th sample on variable Y, is the mean of variable X, is the mean of variable Y, α is the empirical adjustment coefficient, V x is the variance of variable X, V y is the variance of variable Y;
[0069] Execution process: First, calculate the mean of the sample data and Represents the average value of variables X and Y in all samples, and then according to the observed value X on X and Y i and Y i , calculate the deviation for each sample point i, that is and And introduce the weight coefficient w i , the deviation product for each sample The weighted deviation product terms are then summed and divided by n-1 to complete the weighted calculation of the covariance, generate the initial value of the covariance, and then add the empirical adjustment coefficient α to adjust the variance product term V between the features. x V yAfter fine-tuning, the weighted covariance of all feature variables is finally calculated to form a complete improved covariance matrix.
[0070] The specific steps to generate a potential abnormal sample set are:
[0071] Based on the principal component feature set, the meteorological data is segmented, the path depth of each data point in the segmentation tree is calculated, the number of segmentations of each data point is recorded, and the segmentation depth data is obtained;
[0072] Based on the segmentation depth data, the isolation forest is used to evaluate the isolation degree of each data point, and high isolation points are identified according to the number of segmentation times of the data point, and then screened to form a potential abnormal sample set;
[0073] Based on the potential abnormal sample set, data points with strong isolation and variation amplitude are extracted and integrated to obtain the potential abnormal sample set;
[0074] Based on the principal component feature set, the meteorological data is segmented. A series of segmentation trees are constructed using the segmentation tree algorithm to segment the data. Specifically, each data point is segmented in the tree nodes in turn, and the appropriate node segmentation points are selected according to the data distribution characteristics. The segmentation path is deeply calculated, and the number of segmentations of each data point is recorded to generate segmentation depth data.
[0075] Based on the segmentation depth data, the isolation forest algorithm is used to evaluate the isolation degree of each data point. Specifically, the segmentation depth of each data point is calculated and accumulated through each tree in the isolation forest, and the isolation score is calculated according to the path depth and the number of segmentations. Based on the isolation score, data points with higher isolation degrees are identified and screened to form a potential abnormal sample set.
[0076] Based on the potential abnormal sample set, data points with high isolation and strong variation are extracted. The isolation ranking method is used to sort the data points with high isolation scores by scores. Data points with large variation are selected according to the scores, and then integrated to generate the potential abnormal sample set.
[0077] Among them, the potential abnormal sample set includes the identified temperature abnormal points, humidity abnormal points and air pressure abnormal points.
[0078] Isolation forest, according to the formula:
[0079]
[0080] Where: s(x) is the isolation score value, x is the specific data point to be evaluated, E(h(x)) is the expected value of the path length, d(x) is the distance from the data point x to the nearest neighbor, c(n) is the normalization constant, w1 is the isolation score adjustment weight, w2 is the path length weight, w3 is the distance influence weight, and w4 is the normalization coefficient adjustment weight;
[0081] Execution process: First, generate the initial isolation score for each data point, adjust the value of weight w1 by calculating the average value, so that the score result is distributed within the expected range, then calculate the expected value of path length E(h(x)), the average number of steps required for data point x to be split into leaf nodes in the decision tree of the isolation forest, and multiply it with weight w2 to enhance the influence of path length on the isolation score. The value of weight w2 can be determined by the ratio of the average value of path length to the path length required for the expected outlier point. Then measure the distance d(x) between the data point x and the nearest neighbor point to reflect the degree of difference between the data point x and the surrounding data distribution, and multiply it with weight w3 to make the influence of distance on the isolation score more significant. The value of weight w3 is determined by the ratio of the mean value of the distance distribution between the sampling points to the mean value of the initial isolation score. Finally, multiply the normalization constant c(n) by weight w4 to adjust the adaptability of the overall isolation score. The value of weight w4 is determined according to the upper limit of the target score range and the proportional relationship of the natural logarithm of the sample size n.
[0082] The method for calculating the amplitude of change is to gradually calculate the rate of change of the meteorological numerical model within a set time window, capture the fluctuation of meteorological data, and use a sliding window to calculate the sliding average of the rate of change to obtain a stable fluctuation amplitude, which is convenient for identifying abnormal conditions. When the sliding average change rate exceeds the preset threshold, it is marked as an abnormal state.
[0083] The specific steps to generate anomaly score list are:
[0084] Based on the potential abnormal sample set, perform deep segmentation calculation on each sample, calculate the path depth of the sample in the segmentation tree item by item, and record the number of segmentations of each sample to obtain the segmentation depth dataset;
[0085] Based on the segmentation depth dataset, an adaptive scoring standard is set to score the segmentation depth of each sample, and the score value is compared with the set threshold to determine whether it exceeds the threshold one by one, and generate a score value list;
[0086] Based on the score value list, samples with scores higher than the set threshold are screened out, marked as abnormal, and sorted and recorded to obtain an abnormal score list;
[0087] Based on the potential abnormal sample set, a deep segmentation calculation is performed on each sample, and the segmentation tree algorithm is used to calculate the path depth of the sample in the segmentation tree item by item. Specifically, the segmentation tree is constructed and the tree nodes are traversed. The tree nodes are binary-divided according to the eigenvalue distribution of the sample. The path depth of the sample is recorded at each node, and the number of segmentations of the sample in the tree is recorded at the same time. Finally, the deep segmentation calculation is completed and a segmentation depth data set is generated.
[0088] Based on the segmentation depth dataset, an adaptive scoring standard is set. Specifically, a scoring function is used to score the segmentation depth of each sample. The path length of the sample in each segmentation tree is calculated according to the set scoring rules. A score value list is generated based on the path depth and set parameters. Each score value is compared with the set threshold item by item. The comparison process checks whether the sample exceeds the threshold standard item by item, and finally a score value list is generated.
[0089] Based on the score value list, samples with scores higher than the set threshold are screened. Specifically, the score value list is screened and samples exceeding the threshold are extracted. The samples are marked using the abnormal marking method. The marking status of each abnormal sample is recorded in the marking list. The abnormal data is sorted and summarized to generate an abnormal score list.
[0090] The abnormality score list includes the abnormality score of temperature, the abnormality score of humidity and the abnormality score of air pressure.
[0091] The specific steps to generate the meteorological anomaly tag set are:
[0092] Based on the anomaly score list, samples with scores higher than the set threshold are screened, and samples are marked as abnormal one by one and saved as abnormal data records to obtain a screened abnormal sample set;
[0093] Based on the screening of abnormal sample sets, samples marked as abnormal are classified according to meteorological characteristics, and grouped according to temperature and humidity characteristics to obtain an abnormal meteorological feature set;
[0094] Based on the abnormal meteorological feature set, the abnormal samples are classified and sorted according to the marking status to generate the final anomaly recognition data set and obtain the meteorological anomaly marking set;
[0095] Based on the abnormal score list, filter out samples with scores higher than the set threshold, use the screening algorithm to traverse the score list one by one, extract samples exceeding the threshold through conditional screening operations, use the marking function to mark each high-scoring sample as abnormal, and save the marked abnormal samples in a data structured manner to the abnormal data record to generate a filtered abnormal sample set;
[0096] Based on the screening of abnormal sample sets, samples marked as abnormal are classified according to meteorological characteristics, and a classification and grouping algorithm is used to group them according to meteorological characteristics. Abnormal samples are divided into temperature and humidity categories according to their characteristics, and are classified into corresponding sets according to their characteristic categories. The grouping operation of meteorological characteristic variables is completed to generate an abnormal meteorological feature set.
[0097] Based on the abnormal meteorological feature set, the abnormal samples are classified and sorted according to the marking status. The state classification algorithm is used to classify the samples according to the marking status. The samples are screened and classified into the corresponding marking set through the classification function. Different state categories are assigned according to the markings. Finally, all groups are integrated to complete the generation of the data set and obtain the meteorological anomaly marking set.
[0098] The meteorological anomaly mark set includes temperature data, humidity data and air pressure data marked as abnormal.
[0099] The specific steps to generate an exception type classification set are:
[0100] Based on the meteorological anomaly tag set, the abnormal samples are grouped one by one according to the meteorological characteristics, and the meteorological characteristic value of each sample is classified into the corresponding meteorological category to form a grouped sample set;
[0101] Based on the grouped sample set, the meteorological type of each sample is marked according to the meteorological feature category of each sample, and the classification mark of each sample is saved to obtain a classification mark set;
[0102] Based on the classification tag set, all samples marked as abnormal are further summarized and sorted by meteorological type to obtain an abnormal type classification set;
[0103] Based on the meteorological anomaly tag set, the abnormal samples are grouped one by one according to the meteorological characteristics, and the feature grouping algorithm is used to classify the samples according to the specific meteorological feature values of each sample. The samples are judged according to the temperature, humidity, and air pressure feature values, and the meteorological category to which the samples belong is determined according to the range or type of the meteorological feature values. Each sample is classified into the corresponding meteorological category set through cyclic iteration to form a grouped sample set;
[0104] Based on the grouped sample set, according to the meteorological feature category of each sample, a category labeling algorithm is used to label the meteorological type of each sample one by one. Specifically, the category information of each sample is retrieved through the feature classification function, and the corresponding category label is assigned according to the grouping feature. The classification label is saved in the sample data through the assignment operation. After traversing all samples, a classification label set is generated;
[0105] Based on the classification tag set, all samples marked as abnormal are further summarized, and the type summary algorithm is used to summarize the abnormal samples according to the meteorological type. The abnormal samples of each meteorological type are extracted and stored in the corresponding type set. The abnormal samples in the set are deduplicated and merged into a unified data structure, and finally an abnormal type classification set is generated;
[0106] Among them, the anomaly type classification set includes temperature mutation class, humidity anomaly class and air pressure fluctuation class.
[0107] Classify samples according to labels, divide abnormal samples into multiple sets by type, count the number of samples, distribution characteristics and specific abnormal manifestations, and summarize all classified abnormal types and statistical data to form an abnormal type classification set.
[0108] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.
Claims
1. An intelligent identification method for abnormal state of meteorological numerical model, characterized in that: The following steps are involved: Step 1: Based on multidimensional meteorological numerical data, principal component analysis is used to calculate the elements of the covariance matrix in turn, perform principal component screening operations on the matrix elements, select eigenvectors with larger variances, sort features according to variance values, map high-dimensional data to the low-dimensional space of principal component features, and obtain the principal component feature set; Step 2: Based on the principal component feature set, the isolation forest is used to gradually segment the meteorological sample set, the segmentation times of each sample are calculated, the isolated data points appearing after segmentation are judged, high isolation points are identified according to the segmentation times, and a potential abnormal sample set is formed through the high isolation point set; Step 3: Based on the potential abnormal sample set, each sample is segmented and deeply scored, and the score value is determined one by one through the adaptive scoring standard, and the samples with scores higher than the preset threshold are marked as abnormal, and the abnormal data score records are compiled to obtain the abnormal score list; Step 4: Based on the anomaly score list, perform an anomaly marking operation on high-scoring samples, filter samples with scores exceeding a threshold, combine them into a meteorological anomaly identification set, and mark them to obtain a meteorological anomaly marking set; Step 5: Based on the meteorological anomaly tag set, the abnormal samples are grouped according to meteorological characteristics, the data are marked as different meteorological types to form classification tags, the tag set is classified and summarized according to the meteorological feature category, and an abnormal type classification set is established.
2. The intelligent identification method of abnormal state of meteorological numerical model according to claim 1 is characterized in that: The specific steps of generating the principal component feature set are: Based on multidimensional meteorological numerical data, principal component analysis is used to extract meteorological characteristic variables, calculate the covariance matrix item by item, perform covariance operation on each pair of characteristics, complete the matrix construction, and obtain the covariance matrix; Based on the covariance matrix, the eigenvalue and eigenvector of each feature are calculated, and the eigenvector that can reflect the trend of data change is screened out by sorting the eigenvalues to obtain the principal component feature; Based on the principal component features, the original data is linearly transformed through the principal component matrix, the high-dimensional data is projected into a low-dimensional space, the meteorological data after dimensionality reduction is obtained, and a principal component feature set is obtained; The principal component feature set includes temperature features, humidity features and air pressure features after dimensionality reduction.
3. The intelligent identification method of abnormal state of meteorological numerical model according to claim 1 is characterized in that: The principal component analysis is based on the formula: Among them: Cov w (X,Y) represents the weighted covariance between variables X and Y, X and Y are meteorological characteristic variables, n is the total number of data samples, and w i is the weight coefficient of the i-th sample, X i is the specific observation value of the i-th sample on variable X, Y i is the specific observation value of the i-th sample on variable Y, is the mean of variable X, is the mean of variable Y, α is the empirical adjustment coefficient, V x is the variance of variable X, V y is the variance of variable Y.
4. The intelligent identification method of abnormal state of meteorological numerical model according to claim 1 is characterized in that: The specific steps of generating the potential abnormal sample set are: Based on the principal component feature set, the meteorological data is segmented, the path depth of each data point in the segmentation tree is calculated, the segmentation times of each data point are recorded, and the segmentation depth data is obtained; Based on the segmentation depth data, an isolation forest is used to evaluate the isolation degree of each data point, and high isolation points are identified according to the number of segmentation times of the data point, and are screened to form a potential abnormal sample set; Based on the potential abnormal sample set, extract data points with strong isolation and variation, and integrate them to obtain a potential abnormal sample set; The potential abnormal sample set includes identified abnormal temperature points, abnormal humidity points and abnormal air pressure points.
5. The intelligent identification method of abnormal state of meteorological numerical model according to claim 1 is characterized in that: The isolation forest is based on the formula: Where: s(x) is the isolation score value, x is the specific data point to be evaluated, E(h(x)) is the expected value of the path length, d(x) is the distance from the data point x to the nearest neighbor, c(n) is the normalization constant, w1 is the isolation score adjustment weight, w2 is the path length weight, w3 is the distance impact weight, and w4 is the normalization coefficient adjustment weight.
6. The intelligent identification method of abnormal state of meteorological numerical model according to claim 4 is characterized in that: The method for calculating the change amplitude gradually calculates the change rate of the meteorological numerical model within a set time window, captures the fluctuation of meteorological data, and uses a sliding window to calculate the sliding average of the change rate to obtain a stable fluctuation amplitude, which is convenient for identifying abnormal conditions. When the sliding average change rate exceeds a preset threshold, it is marked as an abnormal state.
7. The intelligent identification method of abnormal state of meteorological numerical model according to claim 1 is characterized in that: The abnormality score list includes an abnormality score of temperature, an abnormality score of humidity, and an abnormality score of air pressure.
8. The intelligent identification method of abnormal state of meteorological numerical model according to claim 1 is characterized in that: The meteorological anomaly mark set includes temperature data, humidity data and air pressure data marked as abnormal.
9. The intelligent identification method of abnormal state of meteorological numerical model according to claim 1 is characterized in that: The specific steps of generating the abnormal type classification set are: Based on the meteorological anomaly tag set, the abnormal samples are grouped one by one according to the meteorological characteristics, and the meteorological characteristic value of each sample is classified into a corresponding meteorological category to form a grouped sample set; Based on the grouped sample set, each sample is marked with a meteorological type according to the meteorological feature category of each sample, and the classification mark of each sample is saved to obtain a classification mark set; Based on the classification mark set, all samples marked as abnormal are further summarized and sorted according to meteorological types to obtain an abnormal type classification set; The abnormal type classification set includes temperature mutation class, humidity abnormal class and air pressure fluctuation class.
10. The intelligent identification method of abnormal state of meteorological numerical model according to claim 9 is characterized in that: Classify samples according to labels, divide abnormal samples into multiple sets by type, count the number of samples, distribution characteristics and specific abnormal manifestations, and summarize all classified abnormal types and statistical data to form an abnormal type classification set.
Citation Information
Cited By
Agrometeorological decision service system and method based on artificial intelligence large model
CN120634003A
Data standardization and quality detection method for dynamic environment monitoring of cloud side end
CN121691506A