Nevzorov water content probe data quality control processing method and system based on airborne in-situ observation
By classifying, quality-controlling, and optimizing airborne in-situ observation data, the problem of insufficient accuracy in liquid water content inversion in airborne in-situ observation data processing was solved, and the reliability and accuracy of the data were improved, supporting the scientific analysis of cloud microphysical processes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA METEOROLOGICAL ADMINISTRATION WEATHER MODIFICATION CENT
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-08
AI Technical Summary
In existing technologies, the data processing of the Nevzorov water content probe for airborne in-situ observation has failed to form a complete collaborative processing logic, resulting in insufficient accuracy in liquid water content inversion.
By acquiring airborne in-situ observation data, we perform data classification, quality control processing, cloud entry judgment, and inversion optimization. We use adaptive normalization technology to process parameters of different dimensions, and combine multi-condition collaborative matching and inversion model optimization to improve the reliability and accuracy of the data.
This improves the accuracy and reliability of liquid water content inversion, ensures the comparability and computability of data in subsequent analyses, and enhances the scientific rigor of cloud microphysical process analysis.
Smart Images

Figure CN121997238A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of meteorological data analysis technology, specifically relating to a data quality control processing method and system based on airborne in-situ observations from the Nevzorov water content probe. Background Technology
[0002] Airborne in-situ observation technology is one of the core methods for cloud microphysics research. Among them, the Nevzorov water content probe is a key device for directly collecting atmospheric liquid water content and total water content. The reliability of its observation data directly determines the scientific nature of cloud microphysics process analysis.
[0003] In existing technologies, airborne in-situ observation data processing typically encompasses data acquisition, preliminary screening, cloud ingress determination, data correction, and inversion calculation. Specifically, the data acquisition stage involves simultaneously acquiring multi-source data, including Nevzorov probe data, cloud particle probe data, meteorological environmental data, and navigation and positioning data, through airborne equipment. The preliminary screening stage generally involves simple data classification based on the acquisition period or equipment identification. The cloud ingress determination stage often relies on a single optical scattering or temperature and humidity parameter for threshold determination. The data correction stage often uses a fixed baseline or a single environmental parameter for deviation correction. The inversion calculation stage typically relies on preset empirical formulas or simple linear models to retrieve liquid water content. Furthermore, some existing technologies perform localized optimizations for specific stages, such as improving data reliability by refining quality control rules or increasing calculation accuracy by optimizing the inversion model. However, such optimizations are mostly focused on a single stage and do not form a collaborative processing logic across the entire chain. Therefore, improving the accuracy of liquid water content inversion is a technical problem that needs to be addressed. Summary of the Invention
[0004] This application provides a method and system for quality control processing of Nevzorov water content probe data based on airborne in-situ observation.
[0005] This application provides a data quality control processing method for the Nevzorov water content probe based on airborne in-situ observations, applied to the Nevzorov water content probe data quality control processing system. The method includes: acquiring airborne in-situ observation data; performing data classification processing based on the acquisition flight association information of the airborne in-situ observation data to obtain an initial data group; performing merging processing on the initial data group based on flight time threshold conditions to generate an observation dataset; performing quality control processing on the observation dataset according to the Nevzorov probe data quality control task, identifying and deleting abnormal data subsets in abnormal working states from the observation dataset to obtain a quality control dataset; extracting environmental feature parameters and location feature parameters from the quality control dataset and performing multi-condition joint analysis of cloud entry, through multi-condition... Collaborative matching is used to select data collected during the cloud entry period as valid observation data. In response to the water content baseline correction task, baseline reference features are extracted from the valid observation data to construct correction benchmark conditions. The water content data in the valid observation data is then corrected using the correction benchmark conditions to obtain corrected observation data. Using the corrected observation data, particle spectrum number analysis is performed according to the differences in particle spectrum features to obtain the grading results. Particle spectrum distribution information is extracted based on the grading results. An inversion model is constructed based on the correlation between particle spectrum distribution and liquid water content. The particle spectrum distribution information is imported into the inversion model to perform liquid water content inversion optimization. The inversion model parameters are adjusted based on the feature matching results during the inversion optimization process. The inversion optimization results are generated by combining the inversion calculation results and parameter adjustment records.
[0006] This application provides a Nevzorov water content probe data quality control processing system, which includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the above method.
[0007] This application provides a computer-readable storage medium including a computer program. When the computer program is run on a Nevzorov moisture content probe data quality control processing system, the computer program is used to cause the Nevzorov moisture content probe data quality control processing system to perform the steps of the above-described method. Attached Figure Description
[0008] Figure 1 This is a flowchart illustrating a data quality control processing method for a Nevzorov water content probe based on airborne in-situ observation, provided in an embodiment of this application.
[0009] Figure 2 This is a schematic diagram of the structure of a Nevzorov moisture content probe data quality control processing system provided in an embodiment of this application. Detailed Implementation
[0010] See Figure 1 This is a data quality control processing method for Nevzorov water content probe based on airborne in-situ observation provided in the embodiments of this application. This method can be applied to the Nevzorov water content probe data quality control processing system. The specific process is as follows: steps 110-160.
[0011] The multi-source data processed in this application's embodiments (such as environmental parameters, location information, and particle spectrum characteristics) exhibit significant differences in physical meaning and units of measurement. Direct mathematical operations may lead to dimensional mismatch issues, causing the results to lose their physical meaning. Those skilled in the art can effectively overcome this challenge using conventional adaptive normalization techniques. This technique transforms parameters with different dimensions into dimensionless standardized values or unifies them to a common benchmark, for example, by employing min-max normalization or Z-score normalization, thereby eliminating the influence of units and ensuring the comparability and calculability of the data in subsequent steps such as merging, quality control, cloud integration assessment, and model inversion.
[0012] Adaptive normalization of parameters with different physical meanings and dimensions is a routine operation performed by those skilled in the art based on existing technologies. In atmospheric science and data analysis, this type of preprocessing is a standard procedure with mature methods (such as linear scaling and statistical normalization) and extensive tool support. Technicians can flexibly choose methods based on parameter distribution characteristics (such as using Z-score normalization for approximately normally distributed parameters and logarithmic normalization for parameters with large spans) and easily integrate them into the processing flow. This ensures a solid mathematical foundation for each step (such as signal stability screening, water content baseline correction, particle spectrum clustering and inversion model optimization), preventing certain parameters from dominating the calculation results due to excessively large dimensions. Through adaptive normalization processing, which is well-known in the art, it is ensured that all parameters participate in the calculation at a consistent scale, thereby guaranteeing the feasibility of the entire quality control and inversion process and the effectiveness of the results.
[0013] Step 110: Acquire airborne in-situ observation data, and perform data classification processing based on the acquisition flight association information of the airborne in-situ observation data to obtain an initial data group.
[0014] The Nevzorov water content probe data quality control and processing system first reads all raw airborne in-situ observation data from the storage module of the airborne observation equipment. This data comes from sources including water content detection data collected by the Nevzorov water content probe, cloud droplet particle observation data collected by the CDP probe, particle spectrum observation data collected by the CIP probe, and environmental parameter data collected by airborne meteorological sensors. All data includes corresponding acquisition sortie numbers, acquisition timestamps, and equipment identification information. The system extracts the sortie association information from each data entry, including the sortie number, sortie start time, sortie end time, and sortie flight area label. Data collected within the same sortie are grouped into the same subset by sortie number, while data from different sorties are divided into independent subsets, ultimately forming initial data groups based on sortie numbers. Each initial data group contains multiple types of observation data from the corresponding sortie, with the acquisition times of each data point arranged chronologically.
[0015] Step 120: Perform merging processing on the initial data set based on flight time threshold conditions to generate an observation dataset.
[0016] The system extracts the time information of each sortie data in the initial data set, obtaining the start and end times of each sortie. The system presets a sortie time threshold, classifying adjacent sortie data within this threshold as segments of the same continuous flight mission. The system sequentially compares the time interval between the end time of two adjacent sorties and the start time of the next sortie. If this time interval is less than the preset time threshold, the data from the later sortie is merged into the dataset of the previous sortie. The acquisition times in the merged dataset are still arranged in chronological order, and the start time of the merged sortie is updated to the earliest start time, and the end time to the latest end time. If the time interval is greater than or equal to the preset time threshold, the two sorties are retained as independent. The system performs the above merging operation on all sortie data in the initial data set, integrating all merged sortie data and independent sortie data that do not meet the merging conditions into a unified observation dataset. This observation dataset contains all airborne in-situ observation data that has undergone sortie merging processing, and each data point retains the original sortie association information, acquisition timestamp, and equipment identification information.
[0017] Step 130: Perform quality control processing on the observation dataset according to the Nevzorov probe data quality control task, identify and delete the abnormal data subset that is in an abnormal working state from the observation dataset, and obtain the quality control dataset.
[0018] Step 131: Based on the Nevzorov probe data quality control task, determine the quality control dimensions and the analysis priority of each dimension, extract probe working signal data from the observation dataset to perform signal stability screening, construct adaptive threshold judgment conditions for signal fluctuations through time series pattern analysis, and mark signal data that exceeds the adaptive threshold judgment conditions as potential abnormal signals.
[0019] The system first reads the specific requirements of the Nevzorov probe data quality control task from the preset quality control task configuration module, determining the quality control dimensions, including probe operating signal stability, data noise level, water content detection drift, environmental parameter matching, physical rationality, data mutation, and data integrity. Simultaneously, it reads the analysis priority configuration for each dimension, with probe operating signal stability, physical rationality, and data integrity having the highest priority, and the remaining dimensions having secondary priority. The system filters all operating signal data from the Nevzorov probe from the observation dataset, including TWC voltage signal sequences, LWC voltage signal sequences, and probe status parameter signal sequences, all arranged in chronological order of acquisition time. For each signal sequence, the system uses a time-series pattern analysis method to construct an adaptive threshold judgment condition for signal fluctuations. Specifically, it traverses the signal sequence using a fixed-length sliding window, calculating the mean and standard deviation of the signal within each sliding window. The mean plus or minus three times the standard deviation is used as the dynamic threshold interval corresponding to that window. As the sliding window moves, the dynamic threshold interval adjusts according to the real-time fluctuations of the signal. The system compares the signal data at each acquisition moment with the dynamic threshold range of the corresponding sliding window. If the signal data exceeds the range, the signal data at that moment is marked as a potential abnormal signal. At the same time, the acquisition time, signal type, and specific value exceeding the threshold of the signal data are recorded.
[0020] Step 132: Noise identification is performed on the potential abnormal signals. The noise signal components are separated by comparing the signal frequency characteristics with the standard working frequency range. At the same time, drift detection is performed on the water content detection data in the observation dataset to track the changing trend of the detection data in different collection periods and calculate the trend fitting deviation.
[0021] The system performs noise identification on the marked potential anomalous signals. First, it uses a Fourier transform algorithm to extract the frequency characteristics of the potential anomalous signals, obtaining the frequency component distribution of the signals. The system reads the standard operating frequency range of the Nevzorov probe from its equipment parameter library, compares the frequency components of the potential anomalous signals with the standard operating frequency range, and separates the signal components that do not belong to the standard operating frequency range. These components are identified as noise signals, and the system marks them as noise data, while recording the acquisition time and frequency characteristics of the noise signals. Simultaneously, the system extracts all water content detection data collected by the Nevzorov probe from the observation dataset, including TWC and LWC detection data sequences. The data is divided into multiple continuous subsequences according to the acquisition time period, with each subsequence corresponding to a fixed-duration acquisition time period. For each subsequence, the system uses a linear fitting algorithm to fit the trend of data change, generates a trend fitting line, calculates the deviation value between each data point in the subsequence and the trend fitting line, and takes the root mean square of all deviation values as the trend fitting deviation for that collection period. If the deviation value exceeds the preset drift judgment threshold, the water content detection data for that collection period is marked as drift data, and the start and end times of the collection period and the deviation value are recorded.
[0022] Step 133: Combine the flight environment data corresponding to the observation dataset to perform environmental matching status analysis, and include data whose probe working parameters and concurrent environmental parameters have a matching degree lower than the set matching value into the data to be investigated.
[0023] The system extracts flight environment data from the observation dataset that is contemporaneous with the data collected by the Nevzorov probe. This data includes parameter sequences such as atmospheric temperature, relative humidity, atmospheric pressure, and atmospheric particulate matter concentration collected by airborne meteorological sensors, with each parameter sequence's acquisition time corresponding one-to-one with the Nevzorov probe data acquisition time. The system retrieves a correlation model between the Nevzorov probe's normal operating parameters and environmental parameters from its operating parameter library. This model is built based on the probe's physical operating principle and includes the normal ranges for probe voltage signals and moisture content detection values under different environmental parameters. For each acquisition time, the system calculates the matching degree between the Nevzorov probe's operating parameters and the contemporaneous environmental parameters. Specifically, the probe's operating parameters are substituted into the correlation model to obtain the theoretical range of the probe's operating parameters under that environmental parameter, and the degree of overlap between the actual operating parameters and the theoretical range is calculated. The higher the overlap, the higher the matching degree. The system presets a matching threshold and includes Nevzorov probe data corresponding to acquisition times with a matching degree below this threshold in the data to be investigated, while simultaneously recording the acquisition time, probe operating parameter values, and corresponding environmental parameter values for that data.
[0024] Step 134: Compare the data to be investigated with reference to the physical working threshold of the Nevzorov probe to determine whether the data to be investigated conforms to the actual working rules, and simultaneously perform abrupt jump detection to identify data mutation points and mark the dataset corresponding to the mutation period.
[0025] The system reads the physical operating thresholds of the Nevzorov probe from its device parameter library, including the upper and lower limits of TWC detection, LWC detection, maximum and minimum voltage signal values, etc. The system compares the TWC, LWC, and voltage signal values in the data to be investigated with their corresponding physical operating thresholds. If a detection value or signal value exceeds the threshold range, the data is determined to be inconsistent with actual operating patterns and marked as physically abnormal data. Simultaneously, the system uses a sliding window method to perform abrupt change detection on the Nevzorov probe data in the observation dataset. A fixed-length sliding window is set, and the mean and standard deviation of the data within each window are calculated. The mean of the current window is compared with the mean of the previous window. If the difference exceeds a preset abrupt change threshold, a data mutation point is determined within that window, and the corresponding acquisition period is marked as a mutation period. All data within this period are included in the mutation dataset, and the start and end times of the mutation period and the magnitude of data change are recorded.
[0026] Step 135: Perform an integrity scan on the observation dataset to mark missing data, calculate the percentage of missing data in each acquisition period, and record the location of missing data.
[0027] The system first reads the acquisition times of all data in the observation dataset, generating a complete acquisition time sequence. This sequence is generated based on the preset sampling frequency of the airborne equipment and includes all theoretically possible acquisition times. The system compares the acquisition times of the actual observation data with the complete acquisition time sequence, identifies missing acquisition times, and marks these times as missing data points. The system divides the complete acquisition time sequence into multiple acquisition periods of fixed duration, and calculates the ratio of the number of missing data points in each acquisition period to the total number of acquisition times in that period, obtaining the missing data percentage for each acquisition period. The system presets a missing data percentage threshold. If the missing data percentage of a certain acquisition period exceeds this threshold, the acquisition period is marked as a missing data segment, and the start and end times of the segment, the specific time of the missing data points, and the missing data percentage are recorded.
[0028] Step 136: Combining the information obtained from signal stability screening, noise identification, drift detection, environmental matching status analysis, physical rationality comparison, sudden jump detection, and missing data labeling, generate abnormal state classification conditions, classify the datasets that meet the abnormal state classification conditions as abnormal data subsets and remove them from the observation dataset, and obtain the quality control dataset after multi-dimensional quality control processing.
[0029] The system integrates potential abnormal signal markers obtained from signal stability screening, noise data markers obtained from noise identification, drift data markers obtained from drift detection, data to be investigated markers obtained from environmental matching status analysis, physical abnormal data markers obtained from physical rationality comparison, mutation dataset markers obtained from sudden jump detection, and missing data segment markers obtained from missing data markers to generate abnormal state classification conditions. These conditions include: noise data marked as potential abnormal signals, data collected during the same period marked as drift data, data marked as data to be investigated and physically anomalous, data marked as mutation datasets, and data marked as missing data segments. The system traverses the observation dataset, filters out data that meets the above abnormal state classification conditions, classifies it into an abnormal data subset, and then removes this abnormal data subset from the observation dataset. The remaining data is the quality control dataset after multi-dimensional quality control processing. The data in this dataset all meet the normal state requirements of each quality control dimension and can be used for subsequent cloud analysis and data correction processing.
[0030] Step 140: Extract the environmental and location feature parameters from the quality control dataset and perform a multi-condition joint analysis of cloud entry. Select the collected data during the cloud entry period as valid observation data through multi-condition collaborative matching.
[0031] Step 141: Extract temperature parameters, humidity parameters, air pressure parameters, and atmospheric particulate matter concentration parameters from the quality control dataset as environmental feature parameters, and simultaneously extract flight longitude parameters, flight latitude parameters, and flight altitude parameters as location feature parameters.
[0032] The system extracts atmospheric temperature, relative humidity, atmospheric pressure, and particulate matter concentration parameter sequences from the airborne meteorological sensor data module in the quality control dataset. All parameter sequences are arranged in chronological order of acquisition time, with each parameter value bearing a corresponding acquisition timestamp. The system also extracts flight longitude, latitude, and altitude parameter sequences from the airborne navigation system data module. The acquisition time of each parameter sequence corresponds one-to-one with the acquisition time of the meteorological sensor data, ensuring a complete temporal match between environmental and location characteristic parameters. The system stores these extracted parameter sequences in a unified feature parameter library, providing a data foundation for subsequent multi-condition joint analysis in the cloud. Each parameter sequence retains its original equipment identification and acquisition flight information.
[0033] Step 142: Perform temperature and humidity gradient analysis on the temperature and humidity parameters in the environmental characteristic parameters, calculate the temperature and humidity change rate at different altitudes, and generate temperature and humidity gradient curves.
[0034] The system pre-defines flight altitude layer division criteria, dividing the flight altitude range into multiple consecutive altitude layers at fixed altitude intervals. The system groups temperature and humidity parameter sequences according to the altitude layers corresponding to the flight altitude parameters, with each altitude layer corresponding to a subset of temperature and humidity data. For each temperature data subset within an altitude layer, the system calculates the ratio of the temperature difference to the time difference between adjacent acquisition times to obtain the temperature change rate. The average temperature change rate of all temperature change rates within the same altitude layer is taken as the average temperature change rate for that altitude layer. The average humidity change rate for each altitude layer is calculated using the same method. The system generates temperature gradient curves and humidity gradient curves with flight altitude as the x-axis and the average temperature and humidity change rates as the y-axis, respectively. These two curves together form the temperature and humidity gradient curves, reflecting the changing trends of temperature and humidity at different altitude layers.
[0035] Step 143: Obtain the optical scattering trend based on the optical detection data in the quality control dataset, and determine the atmospheric medium homogeneity by the temporal variation characteristics of the optical scattering intensity.
[0036] The system extracts an optical scattering intensity parameter sequence from the CDP or CIP probe data modules in the quality control dataset. This sequence contains the optical scattering intensity value corresponding to each acquisition moment, and the magnitude of the value reflects the degree of scattering of light by particles in the atmosphere. The system performs time-series analysis on the optical scattering intensity parameter sequence, calculating the ratio of the difference in optical scattering intensity between adjacent acquisition moments to the time difference, thus obtaining a scattering intensity change rate sequence. The system presets an atmospheric medium homogeneity judgment threshold. If the values in the scattering intensity change rate sequence exceeding a preset proportion exceed this threshold, the atmospheric medium is determined to be non-uniform during that period, with a high probability of cloud particles present. If the change rate values are all within the threshold range, the atmospheric medium is determined to be homogeneous, indicating a clear sky. The system associates and stores the optical scattering trend analysis results with the corresponding acquisition time periods, marking the atmospheric medium homogeneity status of each time period.
[0037] Step 144: Associate the flight altitude parameter in the location feature parameters with the preset altitude layer division standard to determine the altitude layer category corresponding to each collection time; perform geographic coordinate gridding on the flight longitude parameter and flight latitude parameter, divide the collection area into grid cells of preset size and mark the grid position corresponding to each collection data.
[0038] The system reads a preset flight altitude classification standard, which is set based on atmospheric circulation characteristics and cloud distribution patterns, and includes multiple continuous altitude intervals and corresponding altitude category labels. The system traverses the flight altitude parameter sequence, comparing the flight altitude value at each acquisition time with the altitude classification standard to determine the altitude interval to which the altitude value belongs. The system then assigns the corresponding altitude category label to the data at that acquisition time, completing the flight altitude association operation. The system records the altitude category corresponding to each acquisition time. Simultaneously, the system presets a geographic coordinate grid size. Based on the maximum and minimum values of flight longitude and latitude in the quality control dataset, the acquisition area is divided into multiple square grid cells of preset sizes, each grid cell corresponding to a unique grid number. The system maps the flight longitude and latitude values at each acquisition time to the corresponding grid cell, marking the grid location number of the data, completing the geographic coordinate gridding operation. The system records the grid location number corresponding to each acquisition time.
[0039] Step 145: Based on the results obtained from temperature and humidity gradient analysis, optical scattering trend analysis, flight altitude layer correlation and geographic coordinate gridding, create a logical combination rule for cloud entry judgment, set the weight coefficients corresponding to each analysis result and construct a dynamic criterion adjustment strategy, and adjust the criterion threshold according to the differences in climate characteristics of different flight areas.
[0040] The system integrates temperature and humidity gradient analysis results, optical scattering trend analysis results, flight altitude layer correlation results, and geographic coordinate gridding results to create a logical combination rule for cloud entry judgment. This rule includes the coordinated matching of multiple conditions: when the temperature and humidity gradient curve shows that the rate of change of temperature and humidity at a certain altitude layer exceeds the normal range under clear sky conditions, the optical scattering trend analysis determines that the atmospheric medium is inhomogeneous, the flight altitude is in an altitude layer category where cloud systems frequently occur, and the geographic coordinates are located in a grid cell with high cloud incidence, the system determines that the collection time is within a cloud entry period. The system assigns a corresponding weight coefficient to each analysis result. The weight coefficient is determined based on the degree of influence of each factor on the cloud entry status. For example, the weight coefficient of the optical scattering trend analysis result is the highest, followed by the weight coefficient of the geographic coordinate gridding result. The system constructs a dynamic criterion adjustment strategy. This strategy includes the mapping relationship between climate feature labels of different flight areas and corresponding criterion thresholds. The system reads the climate feature labels of the collection area from the regional climate database of the airborne navigation system and adjusts the thresholds of each condition in the logical combination rule for cloud entry judgment according to the labels. For example, in areas with high cloud incidence, the judgment threshold for the rate of change of optical scattering intensity is reduced to improve the sensitivity of cloud entry status recognition.
[0041] Step 146: Integrate the cloud boundary recognition model with the logical combination rules using the cloud boundary location information output by the cloud boundary recognition model, and verify the judgment result of the logical combination rules using the cloud boundary location information output by the cloud boundary recognition model.
[0042] The system calls a pre-trained cloud boundary recognition model, a time-series classification model built on a Long Short-Term Memory (LSTM) network. The inputs are optical scattering intensity sequences, temperature sequences, humidity sequences, and flight altitude sequences. The output is the cloud boundary state label corresponding to the collection time, with labels including "inside the cloud," "cloud boundary," and "clear sky." The system inputs the corresponding parameter sequences from the quality control dataset into the cloud boundary recognition model to obtain the cloud boundary state label for each collection time, thus acquiring the cloud boundary location information, i.e., the collection time marked as "cloud boundary." The system compares the cloud entry time determined by the logical combination rules with the cloud boundary location information output by the cloud boundary recognition model, calculating the overlap. If the overlap is lower than a preset verification threshold, the weight coefficients of each analysis result in the logical combination rules are adjusted, and the cloud entry judgment is re-executed until the overlap reaches the verification threshold, completing the cloud boundary model integration and improving the accuracy of the cloud entry judgment results.
[0043] Step 147: Based on the integrated judgment model, perform confidence screening and mark the collected data that simultaneously meet the temperature and humidity gradient conditions, optical scattering trend conditions, flight altitude conditions, and geographic coordinate conditions as candidate cloud entry data.
[0044] The system uses an integrated judgment model to evaluate data at each acquisition time in the quality control dataset. It calculates the confidence level of the data's fulfillment of temperature and humidity gradient conditions, optical scattering trend conditions, flight altitude conditions, and geographic coordinate conditions. The confidence level is obtained by weighted summation of the degree of fulfillment of each condition and its corresponding weight coefficient. The system presets a confidence threshold, marking quality control data at acquisition times with confidence levels higher than this threshold as candidate cloud entry data. Candidate cloud entry data includes Nevzorov probe data at that acquisition time, environmental characteristic parameters, location characteristic parameters, and the numerical values indicating the degree of fulfillment of each condition.
[0045] Step 148: Optimize the candidate cloud entry data based on spatiotemporal continuity, delete discrete data points with discontinuous spatiotemporal distribution, determine the collected data corresponding to the cloud entry period and use it as valid observation data.
[0046] The system sorts candidate cloud ingress data by acquisition time, constructing a candidate data time series. The system uses a sliding window method to perform spatiotemporal continuity analysis on this series. A fixed-duration sliding window is set, and the proportion of candidate data within each window to the total number of acquisition times within that window is calculated. If the proportion is higher than a preset continuity threshold, the data within that window is considered continuous cloud ingress data; if the proportion is lower than the threshold, discrete data points are considered to exist within the window. The system marks acquisition times within the window that are not marked as candidate cloud ingress data as clear-sky times. If multiple times before and after a candidate cloud ingress data point are clear-sky times, the data is considered a spatiotemporally discontinuous discrete data point and is removed from the candidate cloud ingress dataset. The system merges the remaining candidate cloud ingress data by acquisition time period, determining the acquisition time period corresponding to continuous candidate cloud ingress data as the cloud ingress time period. All acquisition data within this time period are considered valid observation data, which are stored in the valid dataset for water content baseline correction processing.
[0047] Step 150: In response to the water content baseline correction task, extract the baseline reference features from the effective observation data to construct the correction benchmark conditions, correct the deviation of the water content data in the effective observation data through the correction benchmark conditions to obtain the corrected observation data, use the corrected observation data to perform particle spectrum number analysis according to the differences in particle spectrum features to obtain the grade division results, and extract particle spectrum distribution information based on the grade division results.
[0048] Step 1511: In response to the water content baseline correction task, the effective observation data is divided into time periods, and the time periods during which the data collection process did not enter the cloud and the atmospheric environment was stable are identified. The corresponding observation data are then used as clean air segment data for clean air segment extraction.
[0049] The system first reads the configuration information for the water content baseline correction task, which includes the criteria for determining the clean air segment. These criteria are based on cloud ingress analysis results and environmental parameter stability settings. The system sorts the valid observation data by acquisition time, constructing a complete observation time series. Combining the clear-sky period markers from the cloud ingress analysis results, the system filters out the acquisition periods marked as clear-sky from the observation time series. For each clear-sky period, the system analyzes the stability of its environmental parameters, calculating the standard deviations of atmospheric temperature, relative humidity, and atmospheric pressure within that period. If the standard deviations of all parameters are lower than the preset stability threshold, the period is determined to be a clean air segment—a period where no clouds have entered the system and the atmospheric environment is stable. The system extracts all observation data within this period as clean air segment data, including water content detection data from the Nevzorov probe, probe operating signal data, concurrent environmental parameter data, and location parameter data. This clean air segment data is stored in the baseline reference dataset for constructing the correction benchmark conditions.
[0050] Step 1512: Perform statistical analysis on the moisture content data in the clean air segment data, calculate the average value of the moisture content data in the clean air segment as the moisture content mean benchmark, and calculate the standard deviation of the moisture content data as a reference for the fluctuation range.
[0051] The system extracts moisture content detection data collected by the Nevzorov probe from the clean air segment data, including TWC and LWC data sequences. Each sequence contains moisture content values at all acquisition times within the corresponding clean air segment. The system performs statistical analysis on the TWC data sequences, calculating the arithmetic mean of all values to obtain the TWC mean baseline. It also calculates the square root of the average of the squares of the differences between all values and the mean baseline to obtain the TWC standard deviation, which serves as a reference for the TWC data fluctuation range. The same method is used to calculate the LWC mean baseline and LWC standard deviation for the LWC data sequences, serving as a reference for the LWC data fluctuation range. The system stores the moisture content mean baseline and fluctuation range reference in a calibration baseline parameter library for constructing initial calibration baseline conditions.
[0052] Step 1513: Based on the mean water content benchmark and the fluctuation range reference, construct initial correction benchmark conditions, compare the water content data in the effective observation data with the initial correction benchmark conditions, and determine the deviation data that exceeds the fluctuation range reference.
[0053] The system constructs initial calibration benchmark conditions based on the mean water content baseline and fluctuation range reference. These conditions stipulate that the water content data in the valid observation data should fall within the mean baseline plus or minus three standard deviations. Data exceeding this range is considered biased. The system iterates through the water content data at each acquisition time in the valid observation data, including TWC and LWC data, comparing each data with the corresponding initial calibration benchmark conditions. If TWC data exceeds the TWC mean baseline plus or minus three TWC standard deviations, it is marked as TWC biased data; similarly, if LWC data exceeds the LWC mean baseline plus or minus three LWC standard deviations, it is marked as LWC biased data. The system stores all marked biased data in a biased dataset, recording the acquisition time, water content value, and degree of bias for each data point—the difference between the value and the benchmark range boundary.
[0054] Step 1514: Group the deviation data according to the acquisition height and acquisition time period, perform piecewise linear correction on the data of different groups, obtain the correction coefficient of each group through linear fitting, and perform preliminary correction on the deviation data.
[0055] The system extracts flight altitude parameters and acquisition time information corresponding to deviation data from valid observation data. It divides the deviation data into multiple altitude groups based on flight altitude, each corresponding to a fixed altitude range; similarly, it divides the deviation data into multiple time period groups based on acquisition time, each corresponding to a fixed acquisition time period. For the deviation data within each altitude group, the system uses the least squares method to perform linear fitting with the acquisition time as the x-axis and water content data as the y-axis, obtaining a linear fitting line for that altitude group. The slope and intercept of the line are the correction coefficients for that group. The system substitutes each deviation data point within this group into the linear correction formula to obtain the initially corrected water content data. The correction formula is that the corrected data equals a linear combination of the correction coefficient and the original data. The same method is used to perform piecewise linear correction on the deviation data within each time period group, obtaining the initially corrected water content data. The system stores the initially corrected data in the initially corrected dataset, retaining the acquisition time and grouping information of the original data.
[0056] Step 1515: Perform nonlinear deviation analysis on the initially corrected water content data, identify the nonlinear deviation characteristics and perform nonlinear deviation adjustment, construct the nonlinear deviation correction function and substitute the data to complete the secondary correction.
[0057] The system performs nonlinear deviation analysis on the initially corrected moisture content data, comparing it with the mean benchmark of the clean air data. It calculates the deviation value for each data point and uses a polynomial fitting algorithm to fit the trend of the deviation values. If the fitting result shows that the trend of the deviation values does not conform to a linear law, nonlinear deviation characteristics are identified. The system constructs a nonlinear deviation correction function based on polynomial fitting, taking the initially corrected moisture content data as input and outputting the corresponding correction amount. The system substitutes the initially corrected moisture content data into the nonlinear deviation correction function to obtain the second-corrected moisture content data, which eliminates the influence of nonlinear deviation. The system stores the second-corrected data in a second-correction dataset, recording the acquisition time and correction amount for each data point.
[0058] Step 1516: Incorporate the temperature, humidity and air pressure parameters from the environmental characteristic parameters into the environmental parameter adaptive model, establish the correlation mapping relationship between environmental parameters and correction coefficients, and dynamically update the correction coefficients according to real-time changes in environmental parameters.
[0059] The system calls a pre-trained adaptive environmental parameter model, a regression model trained using a gradient boosting tree algorithm. The inputs are environmental parameters such as atmospheric temperature, relative humidity, and atmospheric pressure, and the output is the corresponding water content correction coefficient. The system inputs temperature, humidity, and air pressure parameters from the effective observation data into this model to obtain the correction coefficients for each data collection time, establishing a mapping relationship between environmental parameters and correction coefficients. The system dynamically adjusts the water content data after secondary correction using these correction coefficients. The adjustment formula is: the dynamically corrected data equals the secondary corrected data multiplied by the corresponding correction coefficient. The system monitors changes in environmental parameters in real time. If the rate of change of environmental parameters exceeds a preset update threshold, the system re-inputs the model to obtain new correction coefficients and performs dynamic updates to ensure the correction results adapt to changes in different environmental conditions. The system stores the dynamically corrected data in a dynamic correction dataset, retaining the original environmental parameter and correction coefficient information.
[0060] Step 1517: Perform residual analysis on the moisture content data after linear and nonlinear corrections, calculate the residuals between the corrected data and the mean of the clean air segment data, and evaluate the residual distribution characteristics.
[0061] The system extracts water content data from the dynamically corrected dataset after linear, nonlinear, and environmental parameter dynamic corrections, including TWC-corrected and LWC-corrected data sequences. The system extracts corresponding water content mean benchmarks from the clean air segment dataset, including TWC and LWC mean benchmarks. The difference between each data point in the TWC-corrected data sequence and the TWC mean benchmark is calculated to obtain the TWC residual sequence; the LWC residual sequence is calculated using the same method. The system performs statistical analysis on the residual sequences, calculating the mean, median, standard deviation, skewness, and kurtosis coefficients of the residuals, and assesses the distribution characteristics of the residuals, including the central location, dispersion, and distribution pattern. If the standard deviation of the residuals exceeds the preset allowable range, the correction effect is deemed insufficient.
[0062] Step 1518: Based on the residual analysis results, perform a correction effect evaluation. If the residual ratio exceeds the preset allowable range, trigger the baseline condition reconstruction process, re-extract the clean air segment data and update the mean moisture content baseline and fluctuation range reference; repeat the correction process until the residual distribution meets the preset requirements, and obtain the corrected observation data after deviation correction.
[0063] The system evaluates the correction effect based on the residual analysis results, calculating the proportion of data points in the residual sequence that exceed the preset allowable residual range, i.e., the residual percentage. If the residual percentage exceeds the preset allowable range, a baseline condition reconstruction process is triggered. The system re-extracts clean air segment data from the valid observation data, updates the mean water content baseline and fluctuation range reference, and reconstructs the correction baseline conditions based on the updated baseline parameters. This process repeats the steps of determining deviation data, piecewise linear correction, nonlinear deviation adjustment, dynamic correction of environmental parameters, and residual analysis until the residual percentage is below the preset allowable range and the residual distribution conforms to the preset normal distribution characteristics. The system integrates the finally corrected water content data with the corresponding environmental parameters, location parameters, and acquisition time information into corrected observation data, which is stored in the corrected dataset for particle spectral number analysis processing.
[0064] Step 1521: Obtain particle detection data from the corrected observation data, and extract particle size parameters, particle number concentration parameters, and particle velocity parameters from the particle detection data as key feature parameters of the particle spectrum.
[0065] The system extracts particle detection data from the CDP or CIP probe data module of the calibrated observation data. This data contains particle measurement information corresponding to each acquisition time, including particle size measurement, particle number concentration per unit volume, and particle velocity. The system extracts particle size parameter sequences, particle number concentration parameter sequences, and particle velocity parameter sequences from the particle detection data. All sequences are arranged in order of acquisition time, and each parameter value carries a corresponding acquisition timestamp and device identification information. The system stores the three extracted parameter sequences uniformly in a particle spectrum feature parameter library as key feature parameters of the particle spectrum.
[0066] Step 1522: Perform particle size distribution analysis on the particle size parameter, count the proportion of particles in different size ranges and generate a particle size distribution histogram, and eliminate data fluctuation interference through histogram smoothing; perform time series analysis on the particle number concentration parameter, identify concentration change characteristics and mark the concentration peak period and concentration trough period, calculate the concentration change rate in different periods and establish a concentration change trend model; based on the results obtained from particle size distribution analysis and concentration change characteristic analysis, determine the key difference points of particle spectrum characteristics, and identify the characteristic boundary points of different particle groups through a key difference point detection algorithm.
[0067] The system analyzes the particle size distribution of the particle size parameter sequence. It presets multiple continuous particle size intervals, calculates the proportion of particles in each interval to the total number of particles, and generates a particle size distribution histogram. The horizontal axis of this histogram represents the particle size interval, and the vertical axis represents the proportion of particles. The system uses a moving average algorithm to smooth the particle size distribution histogram. Centered on each particle size interval, it calculates the average proportion of particles in multiple adjacent intervals, using this as the smoothed proportion for that interval, eliminating interference from random data fluctuations. Simultaneously, the system performs time-series analysis on the particle number concentration parameter sequence, calculating the ratio of the concentration difference to the time difference between adjacent sampling times to obtain the concentration change rate sequence. Based on the sign and magnitude of the change rate, it identifies the characteristics of concentration increases, decreases, and stable changes. The moment when the concentration change rate changes from positive to negative is marked as the concentration peak moment, and the moment when it changes from negative to positive is marked as the concentration trough moment. The time interval between consecutive peak moments and trough moments is defined as the concentration peak period and concentration trough period. The system uses an exponential smoothing algorithm to fit the concentration parameter sequence, establishing a concentration change trend model to reflect the long-term concentration change trend. The system combines particle size distribution analysis to obtain particle size distribution characteristics with concentration change analysis to obtain concentration change trends, identifying key differences in particle spectrum characteristics, such as changes in the peak position of particle size distribution and abrupt changes in concentration change rates. The system employs a density-based key difference point detection algorithm to analyze key feature parameters of the particle spectrum, identifying characteristic boundary points for different particle populations. These boundary points represent abrupt changes in particle size, concentration, and other characteristics of different particle populations, and the corresponding particle size and concentration ranges are marked at these locations.
[0068] Step 1523: Introduce a clustering algorithm to cluster the key feature parameters of the particle spectrum, and cluster particle data with similar particle size and consistent concentration change trends into the same category to obtain the initial clustering results.
[0069] The system employs a density clustering algorithm to cluster key feature parameters of the particle spectrum. This algorithm clusters based on the density distribution of these key feature parameters and can identify clusters of arbitrary shapes. The system uses particle size, particle number concentration, and particle velocity as feature dimensions for clustering, constructing a feature vector matrix where each row represents a key feature vector of the particle spectrum at a given acquisition time. The system sets parameters for the clustering algorithm, including the neighborhood radius and minimum density threshold. The neighborhood radius determines the granularity of the clusters, while the minimum density threshold determines the minimum size of each cluster. The system runs the density clustering algorithm, grouping feature vectors whose distance is less than the neighborhood radius and whose density is greater than the minimum density threshold into the same category. This results in multiple clusters, each corresponding to a group of particles with similar sizes and consistent concentration trends. This result is the initial clustering result. The system stores the feature parameter range and acquisition time range corresponding to each category in the initial clustering result database.
[0070] Step 1524: Construct a regularized level setting template based on the preset particle spectrum level division rules. Compare the initial clustering results obtained from clustering with the regularized level setting template, adjust the clustering category boundaries to conform to the level division labels, extract the spectral width parameter from the particle spectrum data of the corrected observation data, calculate the particle size distribution width of different particle categories and record the extreme values of spectral width, perform peak information capture on the particle size distribution histogram through a peak detection algorithm, determine the particle size peak and concentration peak of each particle category and mark the corresponding feature information, and generate a level association logic integration model based on the regularized level setting results, spectral width parameter extraction and peak information capture, and determine the particle size connection relationship, concentration relationship and spectral width progression relationship between each level based on the level association logic integration model.
[0071] The system reads preset particle spectrum level division rules, which are set based on meteorological research and operational application needs. These rules include information such as the number of levels, reference particle size ranges for each level, and reference concentration ranges. Based on these rules, the system constructs a standardized level setting template, which includes the label, reference particle size range, and reference concentration range for each level. The system compares the initial clustering results with the standardized level setting template, calculating the matching degree between each cluster category and each level. The matching degree is determined by the degree of overlap between the feature range of the cluster category and the reference range of the level. The system adjusts the cluster category boundaries based on the matching degree, aligning the cluster categories with the corresponding level label ranges to obtain the standardized level setting results. The system extracts the spectral width parameter from the particle spectrum data of the corrected observation data. The spectral width parameter represents the standard deviation of the particle size distribution corresponding to each particle category, reflecting the dispersion of particle size. The system calculates the particle size distribution width for different particle categories and records the maximum and minimum spectral width values as spectral width extrema. The system employs a peak detection algorithm to analyze the particle size distribution histogram, capturing peak points to determine the peak particle size for each particle category (i.e., the particle size with the highest distribution percentage). Simultaneously, it determines the corresponding peak concentration (i.e., the particle number concentration value corresponding to that particle size), and marks the peak particle size, peak concentration, and corresponding acquisition time for each particle category. Based on the results of regularized gear setting, spectral width parameter extraction, and peak information capture, the system generates a gear-level correlation logic integration model. This model includes the particle size connection relationship between gears (i.e., the continuity of particle size ranges between adjacent gears), the concentration correlation relationship (i.e., the correlation of concentration change trends between adjacent gears), and the spectral width progression relationship (i.e., the progression of spectral width changes between adjacent gears).
[0072] Step 1525: Combine the particle size correlation, concentration correlation and spectral width progression to perform structured processing of the grade data and generate grade division results.
[0073] The system integrates particle size correlation, concentration correlation, and spectral width progression relationships in the level-level association logic model. It then performs structured processing on the data for each level in the regularized level-level setting results. Specifically, it organizes the particle size range, concentration range, spectral width range, peak characteristics, and other information for each level according to a preset structured format. Each level corresponds to a structured data unit, which includes level label, particle size distribution characteristics, concentration distribution characteristics, spectral width characteristics, peak characteristics, and the range of acquisition time. The system arranges all structured data units in ascending order of particle size, generating level-level division results. This result clearly records the detailed information of each level in the particle spectrum, divided according to characteristic differences.
[0074] Step 1526: Retrospectively verify the gear division results using historical particle spectrum data to compare the adaptation coefficients of the division results with historical standard gears. Cross-validate the results using corrected observation data from different flight sorties to adjust the gear division boundaries to adapt to different observation scenarios. Verify the division results using a particle spectrum theoretical model to ensure that the division results conform to the physical laws of particle spectrum formation. Based on the gear division results verified by multi-level division, and combining the feature boundary points and the gear division results, extract the particle size distribution characteristics, concentration distribution characteristics, spectral width characteristics, and peak characteristics of each gear and integrate them in gear order to obtain particle spectrum distribution information.
[0075] The system first performs retrospective verification of the gear classification results using historical particle spectrum data. It reads standard gear classification results corresponding to particle spectrum data from the same flight area and similar flight time periods from the historical observation database, calculates the fit coefficient between the current gear classification result and the historical standard gears. The fit coefficient is determined by the degree of overlap in the number of gears, particle size range, and concentration range between the two. If the fit coefficient is lower than a preset retrospective verification threshold, the gear classification boundary is adjusted. The system then performs cross-validation using corrected observation data from different flight sorties. Particle spectrum data is extracted from corrected observation data from other flight sorties, and the same gear classification method is used to obtain comparative gear classification results. The current gear classification result is compared with the comparative results, and the gear classification boundary is adjusted to adapt to the differences in particle spectrum characteristics in different observation scenarios. Finally, the system performs theoretical verification of the classification results using a particle spectrum theoretical model. This model is based on cloud physics theory and includes the physical processes and laws of particle spectrum formation. The gear classification results are substituted into the model to determine whether they conform to the physical laws of particle generation, growth, and sedimentation. If they do not conform, the gear classification boundary is adjusted. Based on the grade division results verified by multi-level division, the system combines the previously identified feature boundary points to extract the particle size distribution features, concentration distribution features, spectral width features, and peak features of each grade. Each feature contains the corresponding parameter range and statistical values. The features of all grades are integrated into unified particle spectrum distribution information in grade order. This information comprehensively reflects the distribution law of the particle spectrum.
[0076] Step 160: Construct an inversion model based on the correlation between particle spectrum distribution and liquid water content. Import the particle spectrum distribution information into the inversion model to perform liquid water content inversion optimization. Adjust the parameters of the inversion model based on the feature matching results during the inversion optimization process. Combine the inversion calculation results with the parameter adjustment records to generate the inversion optimization results.
[0077] Step 161: Determine the mathematical mapping relationship between particle size, particle concentration and liquid water content based on particle physical property information, and establish a correlation model between particle spectrum distribution and liquid water content. Construct an initial inversion model through the correlation model.
[0078] The system reads particle physics property information from a particle physics database, including formulas relating the density, volume, and mass of liquid water particles. Based on this information, the system determines the mathematical mapping relationship between particle size, particle concentration, and liquid water content. Specifically, the mass of a single particle equals the particle density multiplied by the particle volume, where the particle volume is the volume of a sphere calculated based on particle size, and the liquid water content is the sum of the masses of all particles per unit volume, which is equal to the integral of the product of particle concentration and the mass of a single particle across all particle sizes. Based on this mathematical mapping relationship, the system establishes a correlation model between particle spectral distribution and liquid water content. The input to this model is particle spectral distribution information, including particle size distribution, concentration distribution, spectral width, and peak characteristics, and the output is the liquid water content value. The system constructs an initial inversion model using this correlation model, which includes the calculation logic for the mapping relationship, input / output interfaces, and parameter configuration modules.
[0079] Step 162: Import the particle size distribution data, concentration distribution data, spectral width data and peak data from the particle spectral distribution information into the initial inversion model, start the iterative inversion operation, and calculate the liquid water content result of the initial inversion.
[0080] The system extracts particle size distribution data, concentration distribution data, spectral width data, and peak value data from a particle spectral distribution information database. This data is then converted into the input format required by the initial inversion model, i.e., a sequence of feature parameters arranged in order of range, with each range containing the corresponding particle size range, concentration range, spectral width value, and peak particle size and concentration value. The system imports the converted input data into the input interface of the initial inversion model and initiates iterative inversion calculations. The initial parameters for the iterations are the model's preset default parameters, including the weighting coefficients of each feature parameter and mathematical model coefficients. During each iteration, the system calculates the liquid water content value based on the correlation model, obtaining the initial inversion liquid water content result, which is a sequence of liquid water content values arranged in order of collection time.
[0081] Step 163: Perform residual distribution matching between the initial inverted liquid water content result and the auxiliary observation data in the corrected observation data, calculate the residual between the initial inverted liquid water content result and the auxiliary observation data, generate a residual distribution curve, determine the distribution shape and central tendency of the residual through the residual distribution curve, and obtain the residual distribution matching result.
[0082] Step 1631: Extract auxiliary observation data related to liquid water content inversion from the corrected observation data and perform time-series alignment processing. Perform preprocessing on the time-series aligned auxiliary observation data. Based on the flight altitude parameters in the quality control dataset, perform hierarchical labeling on the preprocessed auxiliary observation data to divide the auxiliary observation data subsets corresponding to different flight altitude layers.
[0083] The system first filters out auxiliary observation data directly related to the liquid water content inversion from the corrected observation data, including LWC data corrected by the Nevzorov probe and liquid water content estimates calculated based on cloud droplet size by the CDP probe, ensuring the homology of physical quantities between the auxiliary data and the airborne in-situ observations. The system precisely aligns the acquisition time series of the auxiliary observation data with the acquisition time series of the initial liquid water content inversion results, using linear interpolation to fill in missing data points, ensuring complete temporal matching. After alignment, the system performs standardization preprocessing on the auxiliary observation data, converting the data into dimensionless standard values through Z-score transformation to eliminate dimensional differences between data acquired from different devices. Subsequently, the system extracts flight altitude parameters from the quality control dataset and divides the preprocessed auxiliary observation data into subsets corresponding to multiple altitude layers according to a preset atmospheric vertical stratification standard, with each subset corresponding to a continuous altitude interval.
[0084] Step 1632: Based on the preset residual calculation model, perform point-by-point difference calculation between the initial inverted liquid water content result and the corresponding time series auxiliary observation data to obtain the residual value at each acquisition time. The residual value is calculated by subtracting the standardized value of the corresponding auxiliary observation data from the initial inverted liquid water content result value.
[0085] The system invokes a pre-defined residual calculation model, which is built upon error propagation theory and explicitly defines the calculation logic and data alignment rules for the residuals. The system iterates through the initial inversion results and the standardized values of the corresponding time-series auxiliary observation data at each acquisition time, performing point-by-point interpolation to obtain the residual value at each time point. This residual value directly reflects the degree of deviation between the initial inversion results and the auxiliary observation data. The system stores all residual values as a residual sequence in the order of acquisition time to ensure the temporal continuity of subsequent analysis. Simultaneously, it records the flight altitude layer information corresponding to each residual value, providing data support for error analysis at different altitude layers.
[0086] Step 1633: Sort the calculated residual values according to the acquisition time order to construct a residual time series. Based on the residual time series, use a polynomial fitting algorithm to generate a residual distribution curve. During the curve generation process, optimize the curve smoothness by adjusting the fitting order.
[0087] The system rearranges all residual values in chronological order of acquisition time to construct a continuous residual time series, ensuring that the series accurately reflects the trend of residual changes over time. The system employs a polynomial fitting algorithm to perform curve fitting on the residual time series, dynamically adjusting the fitting order based on the residual fluctuation characteristics: if the residual fluctuations are drastic, the fitting order is increased to retain more details; if the residual changes are stable, the fitting order is decreased to ensure curve smoothness. After fitting, a residual distribution curve is generated, which visually displays the temporal variation pattern of the residuals, including the peak values, trough values, and overall fluctuation trend.
[0088] Step 1634: Extract features from the residual distribution curve, identify peak points, valley points and inflection points in the curve, count the residual values and collection times corresponding to each feature point, and analyze the feature differences of the residual distribution curves under different altitude layers by combining the flight altitude layer information marked by the layer.
[0089] The system performs feature extraction on the residual distribution curve. It identifies peaks and troughs by points where the first derivative is zero, and inflection points by points where the second derivative is zero. These feature points correspond to key moments and extreme values in residual changes. The system statistically analyzes the residual value, acquisition time, and corresponding flight altitude information for each feature point, storing the feature points and altitude information in association. Subsequently, the system compares the residual distribution curve characteristics at different altitudes, analyzing the differences in fluctuation amplitude, extreme values, and frequency of change of the residuals at each altitude, thereby identifying the distribution patterns of inversion errors under different atmospheric vertical stratifications.
[0090] Step 1635: Use statistical analysis methods to determine the distribution pattern of the residual values, calculate the skewness coefficient and kurtosis coefficient of the residual values, use the skewness coefficient to determine the symmetry of the residual distribution, use the kurtosis coefficient to determine the steepness of the residual distribution, and determine whether the residual distribution conforms to the characteristics of a normal distribution.
[0091] The system employs a statistical analysis method based on normality testing to determine the distribution shape of the residual values. First, it calculates the skewness coefficient of the residual sequence: if the skewness coefficient is close to zero, the residual distribution is symmetrical; if the skewness coefficient is greater than zero, the residual distribution is right-skewed with a higher proportion of positive residuals; if the skewness coefficient is less than zero, the residual distribution is left-skewed with a higher proportion of negative residuals. Simultaneously, the system calculates the kurtosis coefficient of the residual sequence: if the kurtosis coefficient is close to 3, the residual distribution has the same steepness as a normal distribution; if the kurtosis coefficient is greater than 3, the residual distribution is steeper than a normal distribution with a higher proportion of extreme residuals; if the kurtosis coefficient is less than 3, the residual distribution is flatter than a normal distribution with a lower proportion of extreme residuals. The system combines the results of the skewness and kurtosis coefficients to determine whether the residual distribution conforms to the characteristics of a normal distribution. If it does, the inversion error is mainly random error; if it does not, there is systematic error or outlier interference.
[0092] Step 1636: Calculate the arithmetic mean, median, and mode of the residuals. The arithmetic mean reflects the central position of the residuals, the median eliminates the influence of extreme residuals on the judgment of central tendency, and the mode represents the residual with the highest frequency.
[0093] The system calculates the arithmetic mean of the residual sequence, which reflects the overall central position of the residuals. If the mean is not zero, it indicates a systematic bias in the initial inversion results. To eliminate the interference of extreme residual values on the judgment of central tendency, the system calculates the median of the residual sequence. This value more robustly reflects the central position of the residuals and is not affected by individual extremely large or small residual values. Simultaneously, the system uses the most frequent value in the residual sequence as the mode. This value characterizes the most common error magnitude in the initial inversion results and reflects the typical characteristics of inversion errors. The system combines the arithmetic mean, median, and mode for comprehensive analysis to fully understand the central tendency and typical error characteristics of the residuals.
[0094] Step 1637: Combining the characteristics of the residual distribution curve, the results of the distribution pattern judgment, and the statistical data of central tendency, and combining the residual difference information of different flight altitude layers, generate residual distribution matching results that include the residual time series change pattern, distribution type, and center position parameters.
[0095] The system integrates residual distribution curve characteristics, distribution pattern judgment results, central tendency statistics, and residual difference information at different altitude levels to generate complete residual distribution matching results. These results include the temporal variation patterns of the residuals, such as the fluctuation period and extreme value occurrence periods; the distribution type of the residuals, such as normal distribution and right-skewed distribution; the central location parameters of the residuals, including the arithmetic mean, median, and mode; and the residual difference characteristics at different flight altitude levels, such as the residual fluctuation amplitude and extreme value magnitude at each altitude level. The system stores the residual distribution matching results in a residual analysis results database.
[0096] Step 164: Based on the residual distribution matching results, perform correlation index evaluation, calculate the correlation coefficient between the inversion results and each characteristic parameter in the particle spectrum distribution information, determine the key characteristic parameters affecting the inversion accuracy, and obtain the correlation index evaluation results.
[0097] The system performs correlation index evaluation based on residual distribution matching results. First, it extracts the initial inverted liquid water content result sequence from the inversion result database and extracts various feature parameter sequences from the particle spectral distribution information database, including the mean sequence of particle size distribution, the mean sequence of concentration distribution, the spectral width sequence, and the peak particle size sequence. The system uses the Pearson correlation coefficient algorithm to calculate the correlation coefficient between the liquid water content result sequence and each feature parameter sequence. The correlation coefficient ranges from -1 to +1, with a larger absolute value indicating a stronger correlation. The system sorts the feature parameters according to the absolute value of the correlation coefficient. Feature parameters with an absolute correlation coefficient greater than a preset core threshold are identified as core influencing feature parameters; those between a preset secondary threshold and the core threshold are secondary influencing feature parameters; and those less than the secondary threshold are weakly influencing feature parameters. The system compiles the correlation ranking results of the feature parameters, the determination results of the key feature parameters, and the corresponding correlation coefficient values into a correlation index evaluation result, which is then stored in the correlation evaluation result database.
[0098] Step 165: Generate a dynamic adjustment strategy for model parameters based on the correlation index evaluation results, and use the residual size as a feedback signal for parameter adjustment to adjust the weight coefficients and mathematical model coefficients corresponding to each feature parameter in the initial inversion model.
[0099] Step 1651: Quantitatively analyze the evaluation results of the correlation index, extract the correlation coefficient values between each feature parameter and the inversion result, sort the feature parameters in descending order of the absolute value of the correlation coefficient, and divide them into three levels: core influence feature parameters, secondary influence feature parameters, and weak influence feature parameters. Core influence feature parameters are feature parameters whose absolute value of the correlation coefficient is greater than a preset core threshold, secondary influence feature parameters are feature parameters whose absolute value of the correlation coefficient is between a preset secondary threshold and a core threshold, and weak influence feature parameters are feature parameters whose absolute value of the correlation coefficient is less than a preset secondary threshold.
[0100] The system quantifies and analyzes the correlation index evaluation results. First, it extracts the correlation coefficient between each particle spectral characteristic parameter and the liquid water content inversion result. This value reflects the degree of influence of the characteristic parameter on the inversion result. The system sorts the characteristic parameters from largest to smallest based on the absolute value of the correlation coefficient. According to preset core and secondary thresholds, the characteristic parameters are divided into three levels: core-influence characteristic parameters have a correlation coefficient absolute value greater than the core threshold, indicating the most significant impact on the inversion result; secondary-influence characteristic parameters have a correlation coefficient absolute value between the secondary and core thresholds, indicating some influence on the inversion result; and weak-influence characteristic parameters have a correlation coefficient absolute value less than the secondary threshold, indicating negligible influence on the inversion result. The system stores the classification results in a feature hierarchy library, determining the importance level of each characteristic parameter.
[0101] Step 1652: Based on the hierarchical division results, construct parameter adjustment priority rules, set the adjustment priority of the weight coefficients corresponding to core influencing feature parameters to be higher than that of secondary influencing feature parameters, and the adjustment priority of secondary influencing feature parameters to be higher than that of weakly influencing feature parameters. At the same time, establish a linear correlation mapping between the residual size and the parameter adjustment magnitude, and construct a residual feedback adjustment function. The input of the residual feedback adjustment function is the absolute value of the residual value, and the output is the corresponding parameter adjustment amount. The larger the residual value, the larger the output parameter adjustment amount.
[0102] The system constructs parameter adjustment priority rules based on the hierarchical division of feature parameters, clearly defining that the weight coefficients of core influencing feature parameters have the highest adjustment priority, followed by secondary influencing feature parameters, and weakly influencing feature parameters have the lowest priority. This ensures that model parameter adjustments prioritize the features with the greatest impact on the inversion results. Simultaneously, the system establishes a linear correlation between the absolute value of the residuals and the magnitude of parameter adjustments, constructing a residual feedback adjustment function: when the absolute value of the residuals is large, a larger parameter adjustment amount is output to accelerate model convergence; when the absolute value of the residuals is small, a smaller parameter adjustment amount is output to avoid model parameter oscillations. The residual feedback adjustment function ensures that the magnitude of parameter adjustments matches the severity of the inversion error, achieving precise adjustment of model parameters.
[0103] Step 1653: For the weight coefficients in the initial inversion model, the gradient descent algorithm is used to iteratively adjust the weight coefficients corresponding to the core influencing feature parameters. The objective function is to minimize the residual. In each iteration, the adjustment direction and magnitude of the weight coefficients are determined based on the residual feedback adjustment function, and the adjustment trajectory of the weight coefficients is recorded synchronously.
[0104] The system employs a gradient descent algorithm to iteratively adjust the weight coefficients corresponding to the core influencing feature parameters in the initial inversion model, with the objective function being the minimization of the mean squared error of the residuals. During each iteration, the system calculates the gradient of the objective function with respect to the weight coefficients to determine the adjustment direction: if the gradient is positive, the weight coefficient is decreased; if the gradient is negative, the weight coefficient is increased. Simultaneously, the system determines the adjustment magnitude of the weight coefficients based on the residual feedback adjustment function, ensuring that the adjustment magnitude matches the current residual size. The system synchronously records the time, pre-adjustment value, post-adjustment value, and corresponding residual value for each weight coefficient adjustment, generating a weight coefficient adjustment trajectory to visually demonstrate the optimization process.
[0105] Step 1654: Adopt a local fine-tuning strategy for the weight coefficients corresponding to the secondary influencing feature parameters, and adjust the fine-tuning range in combination with the differences in the observation environment corresponding to the flight sorties, so that the weight coefficients can adapt to the feature response requirements under different observation scenarios.
[0106] The system employs a local fine-tuning strategy for the weight coefficients corresponding to secondary influencing feature parameters to avoid model instability caused by large adjustments. The system adjusts the fine-tuning amplitude based on the differences in the observation environment corresponding to each flight, such as the climate type and cloud type of the flight area: in observation scenarios with complex cloud systems and large differences in particle spectrum characteristics, the fine-tuning amplitude is appropriately increased to allow the weight coefficients to quickly adapt to feature changes; in observation scenarios with stable cloud systems and small differences in particle spectrum characteristics, the fine-tuning amplitude is appropriately decreased to maintain model stability. This local fine-tuning strategy ensures that the weight coefficients of secondary influencing feature parameters can adapt to the feature response requirements of different observation scenarios, further improving the model's versatility.
[0107] Step 1655: For the mathematical model coefficients, based on the theory of the correlation between particle spectrum distribution and liquid water content, and combined with the distribution morphology information in the residual distribution matching results, determine the adjustment boundary of the mathematical model coefficients.
[0108] The system determines the theoretical adjustment boundaries of the mathematical model coefficients in the initial inversion model based on the theoretical relationship between particle spectral distribution and liquid water content, such as the physical correlation between particle volume and mass and the integral calculation logic of liquid water content. This ensures that the adjusted coefficients conform to the basic laws of cloud physics. Simultaneously, the system combines the distribution pattern information from the residual distribution matching results, such as the skewness and kurtosis of the residuals, to further adjust the practical adjustment boundaries of the mathematical model coefficients: if the residual distribution is right-skewed, it indicates that the model is not well-fitted for high particle concentration scenarios, and the upward adjustment boundary of the mathematical model coefficients is appropriately expanded; if the residual distribution is left-skewed, it indicates that the model is not well-fitted for low particle concentration scenarios, and the downward adjustment boundary of the mathematical model coefficients is appropriately expanded. Through the combination of theory and practice, the system determines the reasonable adjustment boundaries of the mathematical model coefficients, ensuring the scientific and rational nature of the model parameter adjustments.
[0109] Step 1656: Construct parameter adjustment constraints, using the operating parameter range of the Nevzorov probe and the atmospheric physical property parameter range as the constraint basis. Perform real-time verification of the weighting coefficients and mathematical model coefficients during the adjustment process. If the parameters exceed the constraint range after adjustment, trigger the backtracking mechanism to restore the parameter values to the values before adjustment and redetermine the adjustment scheme.
[0110] The system constructs parameter adjustment constraints, using the operating parameter ranges of the Nevzorov probe, such as the probe's voltage signal range and water content detection range, and the ranges of atmospheric physical characteristic parameters, such as atmospheric pressure and temperature range, as constraints to ensure that the adjusted model parameters conform to the equipment's operating characteristics and atmospheric physical laws. During parameter adjustment, the system verifies in real time whether the adjusted weight coefficients and mathematical model coefficients are within the constraints. If the parameters exceed the constraints, a backtracking mechanism is triggered to restore the parameters to their pre-adjustment values, and the adjustment scheme is re-determined based on the residual distribution matching results, avoiding unreasonable model output results due to parameters exceeding physical boundaries.
[0111] Step 1657: Use a portion of the data in the quality control dataset as the cross-validation dataset. After each parameter adjustment, import the validation dataset into the initial inversion model for inversion calculation. Calculate the error value between the cross-validation inversion result and the actual observed data in the validation dataset to obtain the cross-validation result.
[0112] The system randomly selects a portion of data from the quality control dataset as the cross-validation dataset. This dataset is independent of the training dataset and is used to validate the performance of the model after parameter tuning. After each parameter tuning, the system imports the cross-validation dataset into the initial inversion model for inversion calculation, obtaining the cross-validation inversion result. Subsequently, the system calculates the error values between the cross-validation inversion result and the actual observed data in the validation dataset, such as mean squared error and mean absolute error, to obtain the cross-validation result. The cross-validation result is used to evaluate the generalization ability of the model after parameter tuning and to avoid overfitting.
[0113] Step 1658: Based on the parameter adjustment priority rules, residual feedback adjustment function, parameter adjustment trajectory record, parameter adjustment constraints and cross-validation results, generate a dynamic adjustment strategy for model parameters. The dynamic adjustment strategy includes adjustment rules for feature parameters at each level, correlation standards between residuals and adjustment magnitudes, parameter adjustment constraint range and adjustment effect verification process.
[0114] The system integrates parameter adjustment priority rules, residual feedback adjustment functions, parameter adjustment trajectory records, parameter adjustment constraints, and cross-validation results to generate a complete dynamic model parameter adjustment strategy. This strategy includes adjustment rules for feature parameters at each level, such as using gradient descent iterative adjustment for core influencing feature parameters and local fine-tuning for secondary influencing feature parameters; correlation criteria between residuals and adjustment magnitude, such as larger residual values resulting in larger adjustment magnitudes; parameter adjustment constraint ranges, such as the physical boundaries of weight coefficients and mathematical model coefficients; and the adjustment effect verification process, such as cross-validation methods and error evaluation metrics.
[0115] Step 166: Set the convergence criteria for the inversion operation, monitor the trend of residual changes during the iterative inversion operation, and determine that the inversion operation has reached the convergence state when the residual change is less than the preset convergence threshold and the number of consecutive iterations reaches the preset value.
[0116] The system sets convergence criteria for the inversion operation, which includes two conditions: the first condition is that the change in residual between two consecutive iterations is less than a preset convergence threshold, where the change in residual is the absolute value of the difference between the mean residuals of the two iterations; the second condition is that the number of iterations that consecutively satisfy the first condition reaches a preset convergence threshold. During the iterative inversion operation, the system monitors the change in residual and the number of consecutive conditions satisfied in real time after each iteration. When both conditions are satisfied simultaneously, the inversion operation is considered to have converged, and the inversion result at this point is the optimal result. If either condition is not satisfied, the inversion operation is considered not to have converged, and iterative operations and parameter adjustments need to continue.
[0117] Step 167: If convergence is not achieved, trigger the error feedback loop, feed back the current residual analysis results to the model parameter adjustment stage, readjust the model parameters and continue to perform iterative inversion calculation.
[0118] When the system detects that the inversion operation has not reached convergence, it triggers an error feedback loop, feeding back the current residual distribution matching results, correlation index evaluation results, and residual changes to the model parameter adjustment stage. Based on the feedback residual analysis results, the system re-executes the dynamic model parameter adjustment strategy, adjusting the weight coefficients and mathematical model coefficients in the initial inversion model. The direction and magnitude of the adjustment are determined based on the residual changes and correlation indices. If the residual changes are large and the core influencing feature parameters have high correlation, the adjustment magnitude of the weight coefficients corresponding to the core influencing feature parameters is increased; if the residual changes are small and the secondary influencing feature parameters have high correlation, the weight coefficients corresponding to the secondary influencing feature parameters are locally fine-tuned. After parameter adjustment, the system continues to perform iterative inversion operations, calculates new inversion results, and re-performs residual analysis and convergence checks until the inversion operation reaches convergence.
[0119] Step 168: During the inversion optimization process, record the adjustment value, adjustment time, adjustment basis, and corresponding inversion results for each model parameter adjustment, and generate optimization records.
[0120] During the inversion optimization process, the system establishes an optimization record module to record detailed information about each model parameter adjustment. Each time a model parameter adjustment is performed, the system records the adjusted weight coefficients and mathematical model coefficients, including the values before and after adjustment, the specific timestamp of the adjustment, the basis for the adjustment (i.e., the corresponding residual analysis results and correlation index evaluation results), and the results of the iterative inversion calculation after adjustment. The system stores the above information in the optimization record module in chronological order of adjustment, generating an optimization record. This record fully reflects the trajectory of parameter adjustments and the corresponding changes in inversion results during the inversion optimization process, and can be used for subsequent verification of inversion results and model improvement.
[0121] Step 169: After the inversion operation reaches convergence, perform feature matching optimization on the inversion results before output to ensure that the feature matching degree between the inversion results and the particle spectrum distribution information meets the preset requirements.
[0122] After determining that the inversion operation has reached convergence, the system extracts the final inversion result sequence from the inversion result database and the corresponding particle spectral distribution feature parameter sequence from the particle spectral distribution information database. The system calculates the feature matching degree between the inversion result sequence and the particle spectral distribution feature parameter sequence. The matching degree is determined by indicators such as the consistency of their changing trends and peak positions. If the matching degree is lower than a preset feature matching threshold, a feature matching optimization process is triggered. The system adjusts the inversion result based on the feature parameters of the particle spectral distribution information. The adjustment method is to correct outliers in the inversion result according to the concentration distribution characteristics and particle size distribution characteristics of the particle spectrum, so that the changing trend of the inversion result is consistent with the changing trend of the particle spectral distribution characteristics. The system repeats the calculation of the matching degree and adjustment of the inversion result until the matching degree reaches the preset requirement, thus obtaining the inversion result after feature matching optimization.
[0123] Step 1610: Combine the final inversion calculation results with the optimization records to generate an inversion optimization result that includes inversion values, parameter adjustment trajectories, residual analysis results, and correlation indicators.
[0124] The system integrates the final inversion calculation result after feature matching optimization with information such as parameter adjustment trajectory, residual analysis results, and correlation indicators from the optimization record to generate the inversion optimization result. This result consists of three parts: the first part is the inversion numerical sequence, i.e., the liquid water content inversion results arranged in chronological order of acquisition time; the second part is the parameter adjustment trajectory, i.e., a detailed record of the time, adjustment value, and basis for each model parameter adjustment; and the third part is the analysis result set, including residual distribution matching results, correlation indicator evaluation results, convergence judgment results, and feature matching optimization results.
[0125] Optionally, the method further includes: Step 210: Perform spatiotemporal correlation mapping based on the inversion optimization results and flight trajectory data. Bind the liquid water content inversion values in the inversion optimization results with the flight longitude parameters, flight latitude parameters, and flight altitude parameters at the corresponding acquisition time in terms of time and spatial dimensions to obtain a spatiotemporal correlation dataset.
[0126] The system extracts liquid water content inversion numerical sequences from the inversion optimization result library, which contain inversion values corresponding to each acquisition time. It also extracts flight longitude, latitude, and altitude parameter sequences from the location feature parameter module of the quality control dataset, with each parameter sequence's acquisition time corresponding one-to-one with the acquisition time of the inversion numerical sequence. The system binds the liquid water content inversion values at each acquisition time with the corresponding flight longitude, latitude, and altitude parameters to construct spatiotemporal correlated data units. Each unit contains information such as acquisition time, liquid water content inversion value, flight longitude, latitude, and altitude. The system arranges all spatiotemporal correlated data units in the order of acquisition time to obtain a spatiotemporal correlated dataset. This dataset achieves spatiotemporal dimensional binding between the liquid water content inversion results and flight trajectory data, reflecting the distribution of liquid water content at different spatiotemporal locations. The system stores this dataset in the spatiotemporal correlated dataset library for subsequent spatiotemporal clustering analysis.
[0127] Step 220: Perform spatiotemporal clustering analysis on the spatiotemporal correlated dataset. Use the spatiotemporal density clustering algorithm to mine the distribution and clustering characteristics of liquid water content in different regions and at different heights, and generate spatiotemporal clusters.
[0128] The system employs a spatiotemporal density clustering algorithm to analyze spatiotemporally correlated datasets. This algorithm considers both the temporal and spatial dimensions of the data, enabling the identification of dense regions in both dimensions. The system sets parameters for the spatiotemporal density clustering algorithm, including a temporal neighborhood threshold, a spatial neighborhood threshold, and a minimum density threshold. The temporal neighborhood threshold is the maximum time interval between two data collection points; the spatial neighborhood threshold is the maximum distance between two flight positions; and the minimum density threshold is the minimum number of data units required to form a cluster. During operation, the system clusters spatiotemporally correlated data units whose time interval and spatial distance are both less than the temporal neighborhood threshold into the same cluster. A cluster is considered valid when the number of data units in it exceeds the minimum density threshold. The system generates multiple spatiotemporal clusters, each corresponding to a region of concentrated liquid water content within a spatiotemporal area. Each cluster includes information such as the temporal range, spatial range, average liquid water content, and number of data units within that region. The system stores these spatiotemporal clusters in a spatiotemporal clustering result database for subsequent atmospheric dynamics feature correlation analysis.
[0129] Step 230: Correlate spatiotemporal clusters with atmospheric dynamics feature data, and establish a mapping relationship between cluster features and atmospheric circulation features and airflow motion features through feature correlation analysis.
[0130] The system reads atmospheric dynamic characteristic data contemporaneous with the spatiotemporal clusters from an airborne meteorological database. This data includes atmospheric circulation characteristic parameters and airflow motion characteristic parameters. Atmospheric circulation characteristic parameters include meridional wind speed, zonal wind speed, and vertical wind speed, while airflow motion characteristic parameters include airflow vorticity and divergence. The acquisition time of all parameters corresponds one-to-one with the acquisition time of the spatiotemporal correlation dataset. For each spatiotemporal cluster, the system extracts the statistical characteristics of its corresponding atmospheric dynamic characteristic data, including mean, standard deviation, and trend. Simultaneously, it extracts the characteristic parameters of the spatiotemporal clusters, including the mean liquid water content, the spatial extent of the clustered area, and the duration of the cluster. The system uses a feature correlation analysis algorithm to calculate the correlation coefficient between the cluster characteristic parameters and the atmospheric dynamic statistical characteristics, determining the degree of correlation between the two. The system establishes a mapping relationship between cluster characteristics and atmospheric circulation and airflow characteristics based on the magnitude of the correlation coefficient. For example, when the zonal wind speed is large, the spatial range of the liquid water content accumulation area is large; when the vertical wind speed is upward, the average liquid water content is higher. The system stores this mapping relationship in the association mapping model library for subsequent construction of dynamic models of atmospheric liquid water distribution.
[0131] Step 240: Construct a dynamic model of atmospheric liquid water distribution based on the mapping relationship. Adjust the spatiotemporal response coefficient of the model by adjusting the parameters in the inversion optimization results, so that the parameter-adjusted trajectory calibration model can characterize the variation law of liquid water content under different spatiotemporal conditions.
[0132] The system constructs a dynamic model of atmospheric liquid water distribution based on the mapping relationship between cluster features and atmospheric dynamic features. This model is a time-series prediction model built using a deep learning algorithm. The input is a sequence of atmospheric dynamic feature parameters, and the output is the predicted spatiotemporal distribution of liquid water content. The system extracts parameter adjustment trajectory data from the inversion optimization result library. This data includes the time, adjustment value, and corresponding changes in the inversion results for each model parameter adjustment. The system employs a parameter calibration algorithm, inputting the parameter adjustment trajectory data into the dynamic model of atmospheric liquid water distribution to adjust the model's spatiotemporal response coefficients. This ensures that the model's output trends are consistent with the inversion optimization results, improving the model's ability to represent the variation patterns of liquid water content under different spatiotemporal conditions. Through multiple iterative calibrations, the system ensures that the error between the model's prediction results and the inversion optimization results is less than a preset calibration error threshold, resulting in a calibrated dynamic model of atmospheric liquid water distribution. This calibrated model accurately reflects the correlation between atmospheric dynamic features and liquid water content distribution, and the system stores it in the atmospheric liquid water model library.
[0133] Step 250: The dynamic model of atmospheric liquid water distribution is updated iteratively over time using the spatiotemporal correlation dataset. The feature weights of the model are corrected using the inversion optimization results of the new data collection flights. The model characterization results are output to reflect the spatiotemporal evolution of atmospheric liquid water in the region.
[0134] The system divides the spatiotemporal correlation dataset into training and validation datasets. The training dataset is used to iteratively update the dynamic model of atmospheric liquid water distribution. Each iteration inputs a sequence of atmospheric dynamic characteristic parameters for a given time period and outputs the corresponding predicted liquid water content distribution. The predicted results are compared with the actual liquid water content inversion results, and the error value is calculated. The model's feature weights are adjusted based on the error value. The system uses the inversion optimization results from new data collection flights as new training data, repeatedly executing the iterative update process to enable the model to adapt to dynamic changes in the atmospheric environment. The system uses the validation dataset to validate the updated model, calculating the prediction error of the validation dataset. If the error is less than a preset validation error threshold, the model update is considered complete. The system applies the updated dynamic model of atmospheric liquid water distribution to characterize the regional atmospheric liquid water distribution. By inputting the atmospheric dynamic characteristic data of the region, the system obtains the spatiotemporal distribution prediction results of liquid water content. This result serves as the model characterization of the spatiotemporal evolution of regional atmospheric liquid water. The system outputs this result to a meteorological data application platform, providing data support for meteorological research and weather modification operations.
[0135] This application realizes a closed-loop scientific management and control of airborne in-situ observation data from raw acquisition to accurate inversion, breaking through the technical bottlenecks of isolated links, simple correction logic, and insufficient inversion accuracy in traditional airborne Nevzorov probe data processing. Specifically, by associating and classifying data collection flights and merging flight time thresholds, the organic integration of observation data from multiple time periods was achieved. Multi-dimensional quality control processing, centered on the Nevzorov probe data quality control task, eliminated invalid data from abnormal operating conditions at the source, fundamentally ensuring data reliability. Multi-condition joint analysis of cloud entry based on environmental and location characteristic parameters, through the collaborative matching of multi-dimensional features, accurately located cloud entry time periods, improving the targeting and accuracy of effective observation data selection. Combined with dynamic correction of baseline reference features and graded particle spectrum number analysis, not only was the baseline offset of water content caused by environmental background interference eliminated, but also the precise characterization of particle spectrum distribution information was achieved through refined particle spectrum grade division. Based on the inversion model constructed according to the correlation between particle spectrum and liquid water content, the parameters were dynamically adjusted through feature matching during the inversion optimization process, further strengthening the fit between the liquid water content inversion results and the actual particle spectrum characteristics.
[0136] This application improves the overall quality and application value of observation data from the airborne Nevzorov probe through the synergistic effect of technical means and data flow in various stages, providing more accurate and reliable observation data support for cloud microphysics research.
[0137] Based on the same inventive concept, this application also provides a Nevzorov water content probe data quality control and processing system. (See also...) Figure 2 As shown, this is a schematic diagram of a possible Nevzorov water content probe data quality control processing system provided in an embodiment of this application. Figure 2 The Nevzorov water content probe data quality control processing system 200 includes a processor 210 and a memory 220. The processor 210 and memory 220 are interconnected via a communication bus. The memory 220 stores computer programs executable by the processor 210. By executing the instructions stored in the memory 220, the processor 210 can perform the steps of the aforementioned Nevzorov water content probe data quality control processing method based on airborne in-situ observations.
[0138] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium including a computer program. When the computer program is run on the Nevzorov water content probe data quality control processing system, the computer program is used to cause the Nevzorov water content probe data quality control processing system to perform the steps of the above-described Nevzorov water content probe data quality control processing method based on airborne in-situ observation.
Claims
1. A method for quality control processing of Nevzorov water content probe data based on airborne in-situ observation, characterized in that, The method includes: Acquire airborne in-situ observation data, and perform data classification processing based on the acquisition flight association information of the airborne in-situ observation data to obtain an initial data group; The initial data set is merged based on flight time threshold conditions to generate an observation dataset; The observation dataset is subjected to quality control processing according to the Nevzorov probe data quality control task. Abnormal data subsets that are in abnormal working state are identified from the observation dataset and deleted to obtain the quality control dataset. Environmental and location feature parameters are extracted from the quality control dataset and a multi-condition joint analysis of cloud entry is performed. The collected data during the cloud entry period is selected as valid observation data through the collaborative matching of multiple conditions. In response to the baseline correction task for water content, baseline reference features are extracted from the effective observation data to construct correction benchmark conditions. The water content data in the effective observation data is corrected for deviation using the correction benchmark conditions to obtain corrected observation data. The corrected observation data is then used to perform particle spectrum number analysis according to the differences in particle spectrum features to obtain the grade division results. Particle spectrum distribution information is extracted based on the grade division results. An inversion model is constructed by correlating the particle spectrum distribution with the liquid water content. The particle spectrum distribution information is then imported into the inversion model to perform liquid water content inversion optimization. The parameters of the inversion model are adjusted based on the feature matching results during the inversion optimization process. The inversion optimization results are generated by combining the inversion calculation results with the parameter adjustment records.
2. The method as described in claim 1, characterized in that, The process of performing quality control processing on the observation dataset according to the Nevzorov probe data quality control task involves identifying and deleting abnormal data subsets that are in abnormal working states from the observation dataset to obtain a quality control dataset, including: Based on the Nevzorov probe data quality control task, the quality control dimensions and the analysis priorities of each dimension are determined. Probe working signal data are extracted from the observation dataset to perform signal stability screening. Adaptive threshold judgment conditions for signal fluctuations are constructed through time series pattern analysis. Signal data exceeding the adaptive threshold judgment conditions are marked as potential abnormal signals. The potential abnormal signals are identified by noise identification. The noise signal components are separated by comparing the signal frequency characteristics with the standard working frequency range. At the same time, drift detection is performed on the water content detection data in the observation dataset to track the changing trend of the detection data in different collection periods and calculate the trend fitting deviation. By combining the flight environment data corresponding to the observation dataset, an environmental matching status analysis is performed, and data whose matching degree between the probe operating parameters and the environmental parameters of the same period is lower than the set matching value are included in the data to be investigated. The data to be investigated is physically compared with the physical working threshold of the Nevzorov probe to determine whether the data to be investigated conforms to the actual working rules. Sudden jump detection is performed simultaneously to identify data mutation points and mark the dataset corresponding to the mutation period. Perform an integrity scan on the observation dataset to mark missing data, calculate the percentage of missing data in each collection period, and record the location of missing data. By combining information obtained from signal stability screening, noise identification, drift detection, environmental matching status analysis, physical rationality comparison, sudden jump detection, and missing data labeling, abnormal state classification conditions are generated. Data sets that meet the abnormal state classification conditions are classified as abnormal data subsets and removed from the observation dataset, resulting in a quality control dataset that has undergone multi-dimensional quality control processing.
3. The method as described in claim 1, characterized in that, The process involves extracting environmental and location feature parameters from the quality control dataset and performing multi-condition joint analysis of cloud entry. Through collaborative matching of multiple conditions, data collected during the cloud entry period is selected as valid observation data, including: Temperature, humidity, air pressure, and atmospheric particulate matter concentration parameters are extracted from the quality control dataset as environmental feature parameters, while flight longitude, flight latitude, and flight altitude parameters are extracted as location feature parameters. Perform temperature and humidity gradient analysis on the temperature and humidity parameters in the environmental characteristic parameters, calculate the rate of change of temperature and humidity at different altitudes, and generate temperature and humidity gradient curves. The optical scattering trend is obtained based on the optical detection data in the quality control dataset, and the atmospheric medium homogeneity is judged by the temporal variation characteristics of the optical scattering intensity. The flight altitude parameter in the location feature parameters is associated with the preset altitude layer division standard to determine the altitude layer category corresponding to each collection time. The flight longitude and flight latitude parameters are subjected to geographic coordinate gridding, the collection area is divided into grid cells of preset size and the grid position corresponding to each collected data is marked; Based on the results obtained from temperature and humidity gradient analysis, optical scattering trend analysis, flight altitude layer correlation and geographic coordinate gridding, a logical combination rule for cloud entry judgment is created. The weight coefficients corresponding to each analysis result are set and a dynamic criterion adjustment strategy is constructed. The criterion threshold is adjusted according to the differences in climate characteristics of different flight areas. The cloud boundary recognition model is integrated with the logical combination rules, and the judgment result of the logical combination rules is verified by the cloud boundary location information output by the cloud boundary recognition model. Based on the integrated judgment model, confidence screening is performed, and the collected data that simultaneously meets the conditions of temperature and humidity gradient, optical scattering trend, flight altitude, and geographic coordinates are marked as candidate cloud entry data. The candidate cloud entry data is optimized based on spatiotemporal continuity. Discrete data points with discontinuous spatiotemporal distribution are deleted, and the collected data corresponding to the cloud entry period is determined and used as valid observation data.
4. The method according to any one of claims 1-3, characterized in that, In response to the water content baseline correction task, baseline reference features are extracted from the effective observation data to construct correction benchmark conditions. The water content data in the effective observation data is then corrected using these benchmark conditions to obtain corrected observation data, including: In response to the water content baseline correction task, the effective observation data is divided into time periods, and the time periods during the collection process during which the cloud body did not enter the atmospheric environment and the corresponding observation data are used as clean air segment data for clean air segment extraction. Statistical analysis was performed on the moisture content data in the clean air section data, and the average value of the moisture content data in the clean air section was calculated as the moisture content mean benchmark. At the same time, the standard deviation of the moisture content data was calculated as a reference for the fluctuation range. Based on the mean water content benchmark and the fluctuation range reference, an initial correction benchmark condition is constructed. The water content data in the effective observation data is compared with the initial correction benchmark condition to determine the deviation data that exceeds the fluctuation range reference. The deviation data is grouped according to the acquisition height and acquisition time period. Piecewise linear correction is performed on the data of different groups. The correction coefficient of each group is obtained by linear fitting and the deviation data is initially corrected. Nonlinear deviation analysis was performed on the initially corrected water content data to identify nonlinear deviation characteristics and perform nonlinear deviation adjustment. A nonlinear deviation correction function was constructed and the data was substituted to complete the secondary correction. The temperature, humidity, and air pressure parameters among the environmental characteristic parameters are incorporated into the environmental parameter adaptive model to establish a correlation mapping relationship between environmental parameters and correction coefficients. The correction coefficients are dynamically updated based on real-time changes in environmental parameters. Residual analysis was performed on the moisture content data after linear and nonlinear corrections to calculate the residuals between the corrected data and the mean of the clean air data and to evaluate the residual distribution characteristics. Based on the residual analysis results, the correction effect is evaluated. If the residual ratio exceeds the preset allowable range, the baseline condition reconstruction process is triggered to re-extract the clean air segment data and update the mean moisture content baseline and fluctuation range reference. The correction process is repeated until the residual distribution meets the preset requirements, and the corrected observation data after deviation correction is obtained.
5. The method as described in claim 4, characterized in that, The process involves using the corrected observation data to perform particle spectral number analysis based on differences in particle spectral characteristics to obtain tier classification results, and extracting particle spectral distribution information based on the tier classification results, including: Obtain particle detection data from the corrected observation data, and extract particle size parameters, particle number concentration parameters, and particle velocity parameters from the particle detection data as key feature parameters of the particle spectrum. Particle size distribution analysis is performed on the particle size parameter to statistically analyze the proportion of particles in different size ranges and generate a particle size distribution histogram. Histogram smoothing is then used to eliminate data fluctuation interference. Time series analysis is performed on the particle number concentration parameter to identify concentration change characteristics and mark the concentration peak and valley periods. The concentration change rate in different periods is calculated and a concentration change trend model is established. Based on the results obtained from particle size distribution analysis and concentration change characteristic analysis, key difference points in particle spectrum characteristics are determined, and a key difference point detection algorithm is used to identify the characteristic boundary points of different particle populations. A clustering algorithm is introduced to cluster the key feature parameters of the particle spectrum, and particle data with similar particle size and consistent concentration change trends are clustered into the same category to obtain the initial clustering results; Based on the preset particle spectrum level division rules, a regularized level setting template is constructed. The initial clustering results obtained by clustering are compared with the regularized level setting template. The clustering category boundaries are adjusted to conform to the level division labels. The spectral width parameter is extracted from the particle spectrum data of the corrected observation data. The particle size distribution width of different particle categories is calculated and the extreme values of spectral width are recorded. Peak information is captured by the particle size distribution histogram through a peak detection algorithm. The particle size peak and concentration peak of each particle category are determined and the corresponding feature information is marked. Based on the regularized level setting results, the content obtained by spectral width parameter extraction and peak information capture, a level association logic integration model is generated. Based on the level association logic integration model, the particle size connection relationship, concentration relationship and spectral width progression relationship between each level are determined. By combining the particle size correlation, the concentration correlation, and the spectral width progression, structured processing of the grade data is performed to generate grade division results; The grading results are retrospectively verified using historical particle spectrum data to compare the adaptation coefficients between the grading results and historical standard gradings. Cross-validation is performed using calibrated observation data from different flight sorties to adjust the grading boundaries to adapt to different observation scenarios. The grading results are theoretically verified using a particle spectrum theoretical model to ensure that the grading results conform to the physical laws of particle spectrum formation. Based on the grading results verified by multi-level grading, and combined with the aforementioned feature boundary points and grading results, particle size distribution characteristics, concentration distribution characteristics, spectral width characteristics, and peak characteristics of each grading are extracted and integrated in grading order to obtain particle spectrum distribution information.
6. The method as described in claim 1, characterized in that, The process involves constructing an inversion model based on the correlation between particle spectral distribution and liquid water content. The particle spectral distribution information is then imported into the inversion model to perform liquid water content inversion optimization. The inversion model parameters are adjusted based on feature matching results during the optimization process. Finally, the inversion optimization results are generated by combining the inversion calculation results with parameter adjustment records, including: Based on particle physics characteristics, the mathematical mapping relationship between particle size, particle concentration and liquid water content is determined, and a correlation model between particle spectrum distribution and liquid water content is established. An initial inversion model is constructed through the correlation model. The particle size distribution data, concentration distribution data, spectral width data, and peak data from the particle spectral distribution information are imported into the initial inversion model, and the iterative inversion operation is started to calculate the liquid water content result of the initial inversion. The residual distribution of the initially inverted liquid water content result is matched with the auxiliary observation data in the corrected observation data. The residual between the initially inverted liquid water content result and the auxiliary observation data is calculated and a residual distribution curve is generated. The distribution shape and central tendency of the residual are determined by the residual distribution curve to obtain the residual distribution matching result. Based on the residual distribution matching results, a correlation index evaluation is performed, the correlation coefficient between the inversion results and each characteristic parameter in the particle spectrum distribution information is calculated, the key characteristic parameters affecting the inversion accuracy are determined, and the correlation index evaluation results are obtained. Based on the correlation index evaluation results, a dynamic adjustment strategy for model parameters is generated, and the residual size is used as a feedback signal for parameter adjustment to adjust the weight coefficients and mathematical model coefficients corresponding to each feature parameter in the initial inversion model. Set convergence criteria for the inversion operation, monitor the trend of residual changes during the iterative inversion operation, and determine that the inversion operation has reached convergence when the residual change is less than the preset convergence threshold and the number of consecutive iterations reaches the preset value. If convergence is not achieved, an error feedback loop is triggered, feeding back the current residual analysis results to the model parameter adjustment stage, readjusting the model parameters and continuing to perform iterative inversion calculations. During the inversion optimization process, the adjustment value, adjustment time, adjustment basis, and corresponding inversion results of each model parameter adjustment are recorded to generate an optimization record; Once the inversion operation reaches convergence, feature matching optimization is performed on the inversion results before output to ensure that the feature matching degree between the inversion results and the particle spectrum distribution information meets the preset requirements. By combining the final inversion calculation results with the optimization records, an inversion optimization result containing inversion values, parameter adjustment trajectories, residual analysis results, and correlation indicators is generated.
7. The method as described in claim 6, characterized in that, The process involves performing residual distribution matching between the initially inverted liquid water content result and the auxiliary observation data in the corrected observation data, calculating the residual between the initially inverted liquid water content result and the auxiliary observation data, generating a residual distribution curve, and determining the distribution shape and central tendency of the residuals through the residual distribution curve to obtain the residual distribution matching result, including: Auxiliary observation data related to liquid water content inversion is extracted from the corrected observation data and time-series aligned. Preprocessing is performed on the time-series aligned auxiliary observation data. Based on the flight altitude parameters in the quality control dataset, the preprocessed auxiliary observation data is hierarchically labeled to divide the auxiliary observation data subsets corresponding to different flight altitude layers. Based on the preset residual calculation model, the initial inverted liquid water content result is compared with the corresponding time series auxiliary observation data by point-by-point difference calculation to obtain the residual value at each acquisition time. The residual value is calculated by subtracting the standardized value of the corresponding auxiliary observation data from the initial inverted liquid water content result. The calculated residual values are sorted according to the acquisition time order to construct a residual time series. Based on the residual time series, a polynomial fitting algorithm is used to generate a residual distribution curve. During the curve generation process, the curve smoothness is optimized by adjusting the fitting order. Feature extraction is performed on the residual distribution curve to identify peak points, valley points and inflection points in the curve, and the residual values and collection times corresponding to each feature point are statistically analyzed. Combined with the flight altitude layer information of the layered label, the characteristic differences of the residual distribution curves under different altitude layers are analyzed. Statistical analysis methods are used to determine the distribution pattern of the residuals, and the skewness coefficient and kurtosis coefficient of the residuals are calculated. The skewness coefficient is used to determine the symmetry of the residual distribution, and the kurtosis coefficient is used to determine the steepness of the residual distribution, so as to determine whether the residual distribution conforms to the characteristics of a normal distribution. Calculate the arithmetic mean, median and mode of the residuals. The arithmetic mean reflects the central position of the residuals, the median eliminates the influence of extreme residuals on the judgment of central tendency, and the mode represents the residuals with the highest frequency. By combining the characteristics of the residual distribution curve, the results of the distribution pattern judgment, and the statistical data of central tendency, and by combining the residual difference information of different flight altitude layers, a residual distribution matching result containing the residual time series change pattern, distribution type, and center position parameter is generated.
8. The method as described in claim 6, characterized in that, The step of generating a dynamic adjustment strategy for model parameters based on the correlation index evaluation results, using the residual magnitude as a feedback signal for parameter adjustment, and adjusting the weight coefficients and mathematical model coefficients corresponding to each feature parameter in the initial inversion model, includes: The correlation index evaluation results are quantitatively analyzed, and the correlation coefficient values between each feature parameter and the inversion results are extracted. The feature parameters are sorted in descending order of the absolute value of the correlation coefficient, and divided into three levels: core influence feature parameters, secondary influence feature parameters, and weak influence feature parameters. Core influence feature parameters are feature parameters whose absolute value of the correlation coefficient is greater than a preset core threshold. Secondary influence feature parameters are feature parameters whose absolute value of the correlation coefficient is between a preset secondary threshold and a core threshold. Weak influence feature parameters are feature parameters whose absolute value of the correlation coefficient is less than a preset secondary threshold. Based on the hierarchical division results, a parameter adjustment priority rule is constructed. The adjustment priority of the weight coefficients corresponding to the core influencing feature parameters is set higher than that of the secondary influencing feature parameters, and the adjustment priority of the secondary influencing feature parameters is higher than that of the weakly influencing feature parameters. At the same time, a linear correlation mapping is established between the residual size and the parameter adjustment magnitude to construct a residual feedback adjustment function. The input of the residual feedback adjustment function is the absolute value of the residual value, and the output is the corresponding parameter adjustment amount. The larger the residual value, the larger the output parameter adjustment amount. For the weight coefficients in the initial inversion model, the gradient descent algorithm is used to iteratively adjust the weight coefficients corresponding to the core influencing feature parameters. With the goal of minimizing the residual, the adjustment direction and magnitude of the weight coefficients are determined based on the residual feedback adjustment function in each iteration, and the adjustment trajectory of the weight coefficients is recorded synchronously. A local fine-tuning strategy is adopted for the weight coefficients corresponding to the secondary influencing feature parameters. The fine-tuning amplitude is adjusted in combination with the differences in the observation environment corresponding to the flight sorties, so that the weight coefficients can adapt to the feature response requirements under different observation scenarios. For the mathematical model coefficients, the adjustment boundary of the mathematical model coefficients is determined based on the correlation theory between particle spectrum distribution and liquid water content, combined with the distribution morphology information in the residual distribution matching results. Parameter adjustment constraints are constructed, and the working parameter range and atmospheric physical property parameter range of the Nevzorov probe are used as the constraint basis. The weight coefficients and mathematical model coefficients are verified in real time during the adjustment process. If the parameters exceed the constraint range after adjustment, a backtracking mechanism is triggered to restore the parameter values before adjustment and redetermine the adjustment plan. A portion of the data in the quality control dataset is used as the cross-validation dataset. After each parameter adjustment, the validation dataset is imported into the initial inversion model for inversion calculation. The error value between the cross-validation inversion result and the actual observed data in the validation dataset is calculated to obtain the cross-validation result. Based on parameter adjustment priority rules, residual feedback adjustment functions, parameter adjustment trajectory records, parameter adjustment constraints, and cross-validation results, a dynamic adjustment strategy for model parameters is generated. The dynamic adjustment strategy includes adjustment rules for feature parameters at each level, correlation standards between residuals and adjustment magnitudes, parameter adjustment constraint ranges, and adjustment effect verification procedures.
9. A Nevzorov water content probe data quality control and processing system, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the data quality control processing method for the Nevzorov water content probe based on airborne in-situ observations as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program that, when run on the Nevzorov water content probe data quality control processing system, causes the Nevzorov water content probe data quality control processing system to perform the steps of the Nevzorov water content probe data quality control processing method based on airborne in-situ observations as described in any of claims 1 to 8.