An intelligent red tide occurrence probability prediction method based on neural network and key factor identification

By identifying high-risk areas for red tides and constructing a probability forecasting network, the problem of insufficient risk differentiation in existing red tide prediction methods has been solved, enabling accurate prediction and efficient early warning of red tide occurrence probability.

CN122022071BActive Publication Date: 2026-06-26自然资源部天津海洋中心(自然资源部天津海洋预报台)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
自然资源部天津海洋中心(自然资源部天津海洋预报台)
Filing Date
2026-04-08
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing red tide prediction methods fail to effectively distinguish risk differences between different sea areas, lack identification of causal relationships, resulting in limited stability and generalization ability of prediction results, and lack of refined expression based on probability, making it difficult to meet the needs of hierarchical decision-making in practical applications.

Method used

By acquiring historical red tide event data and multi-source monitoring data, and combining spatial clustering and risk characterization functions to identify high-risk areas for red tides, multi-scale correlation analysis, sensitivity analysis, and causal correlation identification are conducted to construct a probability forecasting network and output the probability results of red tide occurrence.

Benefits of technology

It has improved the accuracy and stability of red tide prediction, realized the transformation from qualitative judgment to quantitative probability expression, enhanced the continuity and operability of early warning results, and can better serve marine environmental monitoring and prevention and control decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122022071B_ABST
    Figure CN122022071B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent prediction, in particular to a red tide occurrence probability intelligent prediction method based on neural network and key factor identification, comprising S1: obtaining historical red tide event data, multi-station marine environment monitoring data, red tide emergency monitoring data and station continuous hydrological and meteorological observation data of a target sea area, and processing to obtain a standardized monitoring data set; identifying a red tide prediction key area, extracting a multi-period monitoring sequence to form a key area time sequence sample set; S2: performing multi-scale correlation analysis, sensitivity analysis and causal correlation identification to obtain a key factor sorting result; extracting a target key factor affecting red tide occurrence, determining a threshold boundary range of each target key factor, and generating a key factor threshold representation set; S3: inputting a probability prediction network and outputting a red tide occurrence probability result; and generating red tide occurrence early warning information of the target sea area within a prediction period. The present application improves the accuracy and stability of red tide prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent forecasting technology, and in particular to an intelligent forecasting method for the probability of red tide occurrence based on neural networks and key factor identification. Background Technology

[0002] Red tides are a typical harmful algal bloom in marine ecosystems. Their occurrence is influenced by multiple environmental factors, including water temperature, salinity, nutrients, light conditions, and hydrodynamic processes, exhibiting significant multi-factor coupling and temporal evolution characteristics. With the increasing human activities in nearshore waters, the frequency of red tides is on the rise, posing a serious threat to fisheries production, marine ecological security, and public health. Therefore, establishing a high-precision, early-warning method for predicting the probability of red tide occurrence has become an important research direction in the field of marine environmental monitoring and control. In existing technologies, empirical threshold models are usually built based on single or a few environmental factors, or statistical regression, machine learning, and other methods are used to predict red tide occurrence. However, most methods rely on modeling with data from the entire region, failing to effectively distinguish the risk differences between different sea areas. Furthermore, feature selection is often based solely on correlation analysis, lacking the identification of causal relationships, and easily introducing redundant or spurious factors, thus affecting the accuracy of model prediction.

[0003] Existing red tide prediction methods still have shortcomings in feature representation and model construction. On the one hand, most methods directly use raw monitoring data or simple normalized data as input, lacking a structured expression of the characteristics of suitable ranges for environmental factors, and making it difficult to reflect the role mechanism of key environmental conditions in the formation of red tides. On the other hand, traditional models mostly adopt static or weak time-series modeling methods, which are difficult to effectively capture the dynamic evolution of environmental factors in the time dimension and their lag effects. They also lack the ability to adaptively allocate weights for the combination of multiple factors, resulting in limited stability and generalization ability of prediction results. In addition, existing early warning results are mostly based on qualitative levels or simple threshold judgments, lacking a refined expression based on probability, which is not conducive to hierarchical decision-making in practical applications. Summary of the Invention

[0004] This invention provides an intelligent forecasting method for red tide occurrence probability based on neural networks and key factor identification. It can integrate key area screening, key factor identification, threshold feature representation and time series probability modeling to improve prediction accuracy and early warning reliability.

[0005] A smart forecasting method for red tide occurrence probability based on neural networks and key factor identification includes the following steps:

[0006] S1: Acquire historical red tide event data, multi-station marine environmental monitoring data, red tide emergency monitoring data, and continuous hydrological and meteorological observation data of the target sea area. Perform anomaly removal, missing data repair, time alignment, and unified archiving on all types of data to obtain a standardized monitoring dataset. Based on the standardized monitoring dataset and historical red tide event data, calculate the red tide risk characterization results for each sea area unit, and identify key red tide forecasting areas based on the red tide risk characterization results. Extract multi-time period monitoring sequences corresponding to the key red tide forecasting areas from the standardized monitoring dataset to form a time series sample set for key areas.

[0007] S2: Perform multi-scale correlation analysis, sensitivity analysis, and causal association identification on the time-series sample set of the key areas to obtain the ranking results of key factors; based on the ranking results of key factors, extract the target key factors affecting the occurrence of red tides, and combine the monitoring changes before, during, and after the occurrence of red tides to determine the threshold boundary range of each target key factor; according to the target key factors and the threshold boundary range, perform threshold state encoding on the time-series sample set of the key areas to generate a key factor threshold characterization set;

[0008] S3: Input the key factor threshold representation set into the probability prediction network. The probability prediction network performs collaborative learning on the combination relationship and temporal evolution relationship of the key factors and outputs the red tide occurrence probability result. Based on the red tide occurrence probability result, generate red tide occurrence early warning information for the target sea area within the forecast period.

[0009] Optionally, S1 specifically includes:

[0010] S11: Collect historical red tide event records, multi-station marine environmental monitoring data, red tide emergency monitoring data, and continuous hydro-meteorological observation data from coastal stations in the target sea area;

[0011] S12: Preprocess various types of data to obtain a standardized monitoring dataset;

[0012] S13: Based on the standardized monitoring dataset and the historical red tide event data, a spatiotemporal clustering algorithm is used to divide the sea area into units, and the historical red tide occurrence frequency and environmental factor comprehensive index of each sea area unit are calculated to obtain the red tide risk characterization results; according to the preset risk threshold, key areas with high red tide risk are identified.

[0013] S14: Extract continuous monitoring sequences corresponding to the key areas from the standardized monitoring dataset, and extract subsequences for multiple time periods before, during, and after the red tide event, based on the red tide event occurrence time, to construct a time series sample set for the key areas.

[0014] Optionally, S12 specifically includes:

[0015] Box plots were used to identify and remove outliers that were outside the normal range.

[0016] Missing data are filled in using linear interpolation based on time series data or spatial interpolation methods based on adjacent stations.

[0017] Data from different sources is resampled to the same temporal resolution, and time delays caused by differences in sampling frequency are corrected.

[0018] The processed data is converted into a standardized format and a database is established to obtain a standardized monitoring dataset.

[0019] Optionally, the box plot method specifically includes sorting the time series data of any monitoring variable, calculating the first quartile, the third quartile and the corresponding interquartile range, using the first quartile minus a preset multiple of the interquartile range as the lower threshold, and the third quartile plus a preset multiple of the interquartile range as the upper threshold, and judging the monitoring data below the lower threshold or above the upper threshold as outliers and removing them.

[0020] Optionally, S2 specifically includes:

[0021] S21: Wavelet coherence analysis is used on the time series sample set of the key area to extract the time-frequency correlation features between environmental factors and red tide occurrence at different time scales, and to identify environmental factors that are significantly related to red tide events.

[0022] S22: Based on the random forest feature importance assessment or mutual information method, calculate the contribution of each environmental factor to the occurrence of red tide, sort them according to the size of the contribution, and select the environmental factors with the highest number of sorted environmental factors as sensitive environmental factors.

[0023] S23: For the screened highly sensitive environmental factors, Granger causality test is used to identify the causal relationship between the highly sensitive environmental factors and the occurrence of red tides, eliminate pseudo-correlation factors that are only correlated, and obtain the ranking results of key factors.

[0024] S24: Based on the ranking results of the key factors, select several environmental factors that rank highly as target key factors; using the time of occurrence of historical red tide events as a benchmark, statistically analyze the range of numerical changes of each target key factor before, during and after the occurrence of red tide, and use the quantile method to determine the threshold boundary range of each target key factor.

[0025] S25: Based on the threshold boundary range, perform state encoding on each target key factor in the time series sample set of the key region, use fuzzy membership function to calculate the membership value of the key factor value relative to the threshold range, and generate a key factor threshold representation set including time series state information.

[0026] Optionally, the Granger causality test constructs a baseline prediction model based solely on historical data of red tide occurrence characterization sequences, and an extended prediction model that incorporates historical sequences of environmental factors into the baseline prediction model. By comparing the prediction errors or fitting effects of the baseline prediction model and the extended prediction model, it is determined whether the historical information of environmental factors can significantly improve the prediction ability of red tide occurrence. When the model's prediction ability is significantly improved after incorporating the historical sequences of environmental factors, it is determined that the environmental factors have a Granger causal relationship with red tide occurrence and are retained as key target factors; otherwise, they are determined to be pseudo-correlation factors and are removed.

[0027] Optionally, the fuzzy membership function calculates the corresponding membership value based on the deviation between the actual value of the target key factor and the optimal threshold at each time point, so as to reflect the suitability contribution of the target key factor to the occurrence of red tide at the corresponding time point; when the value of the target key factor is closer to the optimal threshold, the membership value is higher, and when it deviates further from the optimal threshold, the membership value is lower.

[0028] Optionally, S3 specifically includes:

[0029] S31: Construct a probabilistic prediction network including an input layer, a hidden layer, and an output layer;

[0030] S32: Using the actual occurrence of historical red tide events as labels, the probability prediction network is trained in a supervised manner using the cross-entropy loss function, and the network parameters are optimized through the backpropagation algorithm until the model converges.

[0031] S33: Input the key factor threshold representation set obtained after processing the real-time monitoring data through steps S1-S2 into the trained probability forecasting network to obtain the probability value of red tide occurrence in the target sea area within the preset forecast period.

[0032] Optionally, the input layer includes the key factor threshold representation set, which includes the threshold state encoding sequence of each target key factor at multiple time steps; the hidden layer uses a temporal convolutional network structure to extract temporal features from the input sequence, and learns the combination relationship and weight allocation between different key factors through an attention mechanism; the output layer uses the Sigmoid function to map the deep features learned by the temporal convolutional network into red tide occurrence probability values ​​and outputs the probability results.

[0033] Optionally, S33 further includes:

[0034] Based on the red tide occurrence probability value and combined with the preset multi-level warning thresholds, a corresponding red tide warning level is generated;

[0035] The obtained warning level information will be distributed to relevant management departments and public users through visual map display, SMS push or API interface.

[0036] The beneficial effects of this invention are:

[0037] This invention identifies high-risk key areas for red tides based on historical red tide events and multi-source monitoring data, combined with spatial clustering and risk characterization functions. This avoids indiscriminate modeling of data from the entire sea area, reduces data redundancy, and improves analytical efficiency. Through a layer-by-layer screening mechanism using multi-scale correlation analysis, sensitivity analysis, and Granger causality tests, environmental factors are jointly identified from three dimensions: correlation, contribution, and causality. This effectively eliminates pseudo-correlation factors with only statistical correlation, improving the reliability of key factor identification results. Based on the quantile method, threshold boundary ranges for target key factors are constructed, and fuzzy membership functions are used to transform the original monitoring data into a continuous representation reflecting the suitability of red tides, realizing a mapping from the original numerical space to the risk semantic space. This ensures that the input features not only have high information density but also reflect the environmental suitability mechanism in the red tide formation process, enhancing the effectiveness and interpretability of the model input.

[0038] This invention utilizes a temporal convolutional network to model the threshold representation set of key factors, effectively extracting the changing trends, lag responses, and local dynamic features of various environmental factors over time. By introducing an attention mechanism, the feature representations of different key factors are adaptively weighted, dynamically focusing on the combination of key factors that play a dominant role in red tide occurrence, thus enhancing the expressive power of multi-factor coupling relationships. In the output layer, the probability of red tide occurrence is directly output through a probability mapping function, and risk classification is achieved by combining multi-level warning thresholds. This transforms the forecast results from traditional qualitative judgments to quantitative probability expressions, not only improving the accuracy and stability of red tide prediction but also enhancing the continuity and operability of warning results, thus better serving marine environmental monitoring and red tide prevention and control decisions. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention;

[0041] Figure 2 This is a schematic diagram of the logic framework of an embodiment of the present invention. Detailed Implementation

[0042] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. Those skilled in the art may employ other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0043] like Figures 1-2 As shown, a smart forecasting method for red tide occurrence probability based on neural networks and key factor identification includes the following steps:

[0044] S1: Acquire historical red tide event data, multi-station marine environmental monitoring data, red tide emergency monitoring data, and continuous hydrological and meteorological observation data of the target sea area. Perform anomaly removal, missing data repair, time alignment, and unified archiving on all types of data to obtain a standardized monitoring dataset. Based on the standardized monitoring dataset and historical red tide event data, calculate the red tide risk characterization results for each sea area unit, and identify key areas for red tide forecasting based on the red tide risk characterization results. Extract multi-time period monitoring sequences corresponding to key areas for red tide forecasting from the standardized monitoring dataset to form a time series sample set for key areas.

[0045] S1 specifically includes:

[0046] S11: Collect historical red tide event records, multi-station marine environmental monitoring data, red tide emergency monitoring data, and continuous hydrological and meteorological observation data from coastal stations in the target sea area to form a raw multi-source monitoring dataset.

[0047] S12: Perform the following processing operations on various types of raw data, including:

[0048] S121, Outlier Removal: Using box plots, based on quartiles , and interquartile range The anomaly detection range is:

[0049] ;

[0050] in, For a certain monitoring variable at time 10:00 The observed values, , The first and third quartiles, The interquartile range (IQR) is used for outlier identification. Box plots are a method based on the location characteristics of data distribution. The core idea is to use the median and upper and lower quartile intervals to characterize the normal data range and identify data deviating from this range as outliers. When an observation is lower than the first quartile minus 1.5 times the IQR, it indicates that it is significantly lower than the lower bound of the normal data distribution; when an observation is higher than the third quartile plus 1.5 times the IQR, it indicates that it is significantly higher than the upper bound of the normal data distribution.

[0051] S122, Missing Data Repair: Data is filled using different interpolation methods based on its temporal and spatial attributes, including:

[0052] For missing data in time series data, linear interpolation is used for completion, and for missing time points... The effective observations before and after it are respectively , The interpolation result is expressed as:

[0053] ;

[0054] For spatially distributed data, the inverse distance weighted interpolation method for adjacent stations is used, expressed as:

[0055] ;

[0056] in, For missing moments, Indicates the position to be interpolated The estimated value at that location, Indicates the first The measured values ​​of environmental factors at each site. For the point to be interpolated and the first The distance between stations This is the distance weighting index, used to control the degree to which distance affects weight decay. Its value ranges from 1 to 3. When the value is small, the influence of distance on the weight is weak, and distant stations still have some influence, making it suitable for scenarios with relatively gentle spatial changes. When the value is large, the weight decays faster with distance, and the interpolation result depends more on neighboring stations, making it suitable for scenarios with large spatial gradient changes. Indicates the first Each site has an interpolation location Weighting coefficients;

[0057] For missing data in time series, linear interpolation is used for repair. Between two adjacent known observations in time, it is assumed that the data change is continuous and approximately linear. When data is missing at a certain moment, the two most recent valid observations before and after that moment can be used to construct a straight line connecting the two points to estimate the missing location. The closer the missing value is to the previous moment, the closer its value is to the observation value at the previous moment; the closer it is to the next moment, the closer its value is to the observation value at the next moment, thus achieving a smooth transition weighted by time.

[0058] For spatially distributed data, an inverse distance weighted interpolation method is used for completion. This method is based on the principle of spatial proximity, meaning that observation points that are spatially closer have a greater impact on the target location, while observation points that are farther away have a smaller impact. In actual calculations, several known observation stations are selected around the location to be interpolated. Different weights are assigned to these stations based on their distance from the target location, with higher weights for closer stations and lower weights for farther stations. The observation values ​​of each station are then weighted and averaged according to their corresponding weights to obtain an estimate of the target location.

[0059] S123, Time Alignment: Unifying data from different sources to the same time resolution. Through the resampling function The original time series is processed and represented as follows:

[0060] ;

[0061] ;

[0062] in, This is the resampling function. Condition 1 indicates that the original sampling frequency is higher than the target time resolution, and condition 2 indicates that the original sampling frequency is lower than the target time resolution. Original time The observed values, For the target time The corresponding resampling results, To reach the target time The set of original sampling points within the corresponding time window, For time window The number of original sampling points within; , To reach the target time Two adjacent initial sampling times, , For the observations corresponding to the original sampling time, For the target time The interpolation results are as follows: When the original sampling frequency is higher than the target time resolution, the resampling function adopts the downsampling aggregation method, that is, within each target time window, multiple original observations falling into the window are statistically aggregated to obtain the resampling value corresponding to the window. The aggregation method can be selected as the mean, maximum, minimum or cumulative value according to the physical meaning of the monitored variable; when the original sampling frequency is lower than the target time resolution, the upsampling interpolation scheme is adopted, that is, a new sampling time is introduced on the target time axis, and the value of the target time is estimated according to the change relationship between adjacent original observation points.

[0063] Simultaneously, time delay correction is performed on different data sequences based on the cross-correlation function, expressed as:

[0064] ;

[0065] in, This is the time lag. This is the corrected time lag.

[0066] The cross-correlation function measures the similarity between two sequences under different translation conditions. When the sequences to be corrected are shifted forward or backward by a certain amount of time, if the numerical changes of the two sequences at the corresponding time points become more consistent, the cross-correlation value will increase; when the two sequences are poorly aligned, the cross-correlation value will be smaller. Therefore, by iterating through different candidate time lags, the optimal time lag can be determined by finding the time offset that maximizes the cross-correlation function. The entire sequence to be corrected is then shifted along the time axis by this optimal time lag, thereby achieving time lag correction.

[0067] S124, Unified Archiving: Convert the processed data into a unified format and build a standardized database to form a standardized monitoring dataset. ;

[0068] S13: Based on standardized monitoring datasets Based on historical red tide event data, the target sea area is spatially divided into units, and clustering functions are used. Obtain the set of sea area units ;

[0069] The DBSCAN clustering algorithm is used to cluster the marine units in the standardized monitoring dataset. For any marine unit, if the number of neighboring marine units within its preset neighborhood radius is not less than the minimum sample size threshold, it is identified as a core unit, and the clustering region is expanded to its density-reachable area based on the core unit. Marine units located in the neighborhood of a core unit but not meeting the core unit identification criteria are merged into boundary units. Marine units that do not belong to any core unit neighborhood and do not meet the core unit identification criteria are marked as noise units. The set of marine units is obtained based on the above clustering results.

[0070] For each sea area unit Calculate the historical frequency of red tide occurrences, expressed as:

[0071] ;

[0072] Simultaneously, a comprehensive index of environmental factors is constructed, expressed as:

[0073] ;

[0074] Based on the above results, a red tide risk characterization function is constructed:

[0075] ;

[0076] Based on preset risk thresholds Screening for high-risk units that meet the criteria:

[0077] ;

[0078] in, For the first Each sea area unit For unit Number of red tide events in Chinese history To calculate the duration of the time, The frequency of red tide occurrence, For unit Inner The average value of each environmental factor These are the environmental factor weighting coefficients. It is a comprehensive index of environmental factors. This is a red tide risk indicator value. These are the weighting coefficients. Historical red tide frequency reflects the cumulative characteristics of red tide occurrence in a target sea area over a longer time scale. It can directly characterize whether the area belongs to a high-incidence area of ​​red tides and has strong historical indicative significance and regional stability. Therefore, in the key area identification stage, historical red tide frequency should usually occupy a relatively high weight to highlight the fundamental role of historically high-incidence sea areas in risk identification. The weight coefficient corresponding to historical red tide frequency... The value range is 0.4-0.7; the comprehensive environmental factor index reflects the degree of support of the current or stage-specific environmental state of a marine unit for red tide formation. It can reflect the comprehensive influence of environmental factors such as water temperature, salinity, chlorophyll, and nutrients on the conditions for red tide formation. This index is relatively sensitive to short-term fluctuations and can enhance the responsiveness of risk assessment to real-time environmental changes. Therefore, it also needs to be assigned a certain weight. However, since environmental factors are easily affected by short-term disturbances, local observation fluctuations, and monitoring noise, their stability is usually weaker than the historical red tide occurrence frequency. Therefore, its weight should generally not be much higher than the historical red tide occurrence frequency. The weight coefficient corresponding to the comprehensive environmental factor index is... The value range is 0.3-0.6; The risk threshold is adaptively determined based on the distribution of red tide risk characterization results across all marine units. The high-quantile values ​​of the red tide risk characterization results are taken as the risk threshold, so that marine units with high-level risk characterization values ​​are identified as key high-risk areas for red tide. The high-quantile values ​​are the 70th to 90th percentiles. This is a set of key regions.

[0079] S14: From standardized monitoring datasets Extract key region set The corresponding continuous monitoring sequence for each red tide event Let the time of its occurrence be Constructing a time window The multivariate monitoring data sequences within this time window are extracted as samples to form a key area time series sample set, represented as:

[0080] ;

[0081] ;

[0082] in, For the first Second red tide event The time of the red tide event. The length of the time window before the event. The length of the time window after the event. For a single time series sample, This is a time series sample set for key regions. This represents the number of samples.

[0083] S2: Perform multi-scale correlation analysis, sensitivity analysis, and causal relationship identification on the time series sample set of key areas to obtain the ranking results of key factors; based on the ranking results of key factors, extract the target key factors affecting the occurrence of red tides, and combine the monitoring changes before, during and after the occurrence of red tides to determine the threshold boundary range of each target key factor; based on the target key factors and the threshold boundary range, perform threshold state encoding on the time series sample set of key areas to generate a key factor threshold characterization set.

[0084] S2 specifically includes:

[0085] S21: Time series sample set for key regions Multi-scale correlation analysis was performed between the time series of various environmental factors and the characterization sequence of red tide occurrence; wavelet coherence analysis was used to calculate the frequency domain correlation between the two at different time scales for any environmental factor sequence. Characteristic sequence of red tide occurrence The wavelet coherence coefficient is expressed as:

[0086] ;

[0087] According to different scales The wavelet coherence coefficient distribution under the model was used to identify environmental factors that exhibited high coherence across multiple scales, serving as a set of candidate environmental factors significantly associated with red tide occurrence.

[0088] in, For the first Time series of environmental factors This is a characterization sequence for red tide occurrence. , The wavelet transform results are for the corresponding sequences. for The complex conjugate, For scale parameters, To facilitate smoothing, the product term and energy term of the wavelet coefficients are weighted and averaged within a certain time window and scale window. The smoothing process includes two aspects: first, smoothing along the time direction, that is, selecting a time window of a certain length near the current time and weighting and averaging the wavelet coefficients within the window to reduce the impact of short-term random fluctuations; second, smoothing along the scale direction, that is, averaging the wavelet coefficients within adjacent scale ranges to enhance the continuity and consistency between different time scales. By jointly smoothing in both time and scale dimensions, the stability and interpretability of the coherence coefficients can be effectively improved.

[0089] S22: Based on the candidate environmental factor set, a random forest model is used to calculate the contribution of each environmental factor to the occurrence of red tide. A random forest model is constructed with environmental factors as input and red tide occurrence characteristics as output. The contribution of each factor is obtained based on the feature importance assessment mechanism, expressed as:

[0090] ;

[0091] Based on the contribution of each environmental factor The environmental factors are sorted and selected as the set of highly sensitive environmental factors, with the top 4 to 6 contributing factors being preferred. If too many environmental factors are selected, a large number of environmental factors with limited contribution or only weak correlation will be retained, which will not only increase the computational burden of subsequent Granger causality tests, but also easily weaken the convergence and discriminative power of the key factor ranking results. Therefore, controlling the selection range to the top 5 is more conducive to achieving a balance between retaining effective candidate factors and suppressing the introduction of redundant factors.

[0092] in, For the first The contribution of each environmental factor In the first The first decision tree The importance gain brought by each feature denoted as the number of decision trees in the random forest.

[0093] S23: For a set of highly sensitive environmental factors, the Granger causality test is used to identify the causal relationship between each environmental factor and the occurrence of red tides. For any environmental factor sequence... Red tide characterization sequence We construct a model that uses only historical information about red tides themselves and a model that incorporates historical information about environmental factors, represented as follows:

[0094] ;

[0095] ;

[0096] in, The lag order is related to the data's temporal resolution and the environmental process's response period. When the data is on an hourly scale, the lag order can range from several hours to tens of hours; when the data is on a daily scale, the lag order can range from 1 to 7 days. , These are the parameters of the autoregressive model. , This is the error term;

[0097] By comparing the prediction errors or fitting effects of the two models, it can be determined whether environmental factors have a causal impact on red tide occurrence; if the model prediction error significantly decreases after introducing environmental factor sequences, then it can be determined that... right If a Granger causal relationship exists, the environmental factor is retained; otherwise, it is judged as a pseudo-correlation factor and removed, and the final ranking result of the key factors is obtained. The prediction error is significantly reduced, which means that after introducing the historical sequence of the environmental factor, the error of the extended prediction model relative to the baseline prediction model decreases by a preset proportion threshold, which is 10%-30%.

[0098] S24: Based on the ranking results of the key factors, select the top-ranked interference factors as the target key factor set. At the time of each red tide event Based on this, the value sequences of each target key factor before, during and after the event are extracted, and their numerical distribution range is statistically analyzed.

[0099] The threshold boundary range of each target key factor is determined based on the quantile method. For the ... The threshold for each key objective factor is expressed as follows:

[0100] ;

[0101] in, Indicates the first The low threshold of each key target factor is used to characterize the lower boundary of the factor's potential role in the occurrence of red tides. Indicates the first The optimal threshold for each key target factor is used to characterize the typical value level of that factor that is most conducive to the occurrence of red tides. Indicates the first The high threshold of each key target factor is used to characterize the upper bound of the factor's potential role in the occurrence of red tides. For the set of key factors of the target, For the first The factor of the first Quantiles The median, For the first Quantiles The lower quantile parameter, ranging from 0.1 to 0.3, is used to characterize lower-level intervals. This represents the high quantile parameter, with a value range of 0.7 to 0.9, and is used to characterize a higher level range.

[0102] S25: Based on the threshold boundary range, perform state encoding on each target key factor in the time series samples of the key region. For any target key factor... The continuous encoding using fuzzy membership functions is as follows:

[0103] ;

[0104] in, For the first Membership values ​​of key objective factors. To control the parameters of the membership function width, when The closer to the optimal threshold When the value of a membership degree is closer to 1, its membership degree value gradually decreases as it deviates further from the optimal threshold.

[0105] The membership results of each key factor at each time point are combined in chronological order to form a key factor threshold representation set, which is represented as follows:

[0106] ;

[0107] in, This is a key factor threshold representation set. The number of key factors for the target.

[0108] S25 employs a membership function to quantify the contribution of a factor to the suitability of red tide occurrence based on the closeness of its value to the optimal threshold. For each target key factor, the optimal threshold is used as the central reference point, and the deviation between the actual observed value of the key factor and the optimal threshold is used as the metric. When the factor value at a certain moment is closer to the optimal threshold, it indicates that the environmental conditions are more favorable for red tide occurrence, and its corresponding membership value is higher; conversely, when the factor value deviates far from the optimal threshold, it indicates that the promoting effect of the environmental conditions on red tide occurrence is weakened, and its membership value decreases accordingly. The membership function uses an exponential decay function to describe this closeness relationship, making the membership... The membership degree shows a continuous and smooth decreasing trend as the degree of deviation increases. The function has the characteristic of being symmetrical about the optimal threshold, that is, regardless of whether the factor value is higher or lower than the optimal threshold, as long as the degree of deviation is the same, its influence on the red tide suitability is considered consistent, thus reflecting a concentrated characterization of the optimal range. By introducing a scale parameter to adjust the change amplitude of the function, the sensitivity of the membership degree to the degree of deviation can be controlled. When the scale parameter is small, the membership degree function changes more steeply near the optimal threshold, indicating that it has high suitability only within a narrow range. When the scale parameter is large, the function changes more gently, indicating that a wider range of factor values ​​is allowed to still have a certain degree of suitability.

[0109] S3: Input the key factor threshold representation set into the probability forecasting network. The probability forecasting network performs collaborative learning on the combination relationship and temporal evolution relationship of the key factors and outputs the red tide occurrence probability result. Based on the red tide occurrence probability result, generate red tide occurrence warning information for the target sea area during the forecast period.

[0110] S3 specifically includes:

[0111] S31: Construct a probabilistic prediction network for predicting the probability of red tide occurrence. The probabilistic prediction network includes an input layer, hidden layers, and an output layer.

[0112] S311, Input Layer: The obtained key factor threshold set As input layer data, the key factor threshold set consists of the membership sequence of each target key factor at multiple time steps, and its input form can be expressed as follows:

[0113] ;

[0114] in, For the first The key factors of the target in the first Membership value at each time step The number of key factors for the target. To calculate the duration of the time, For dimension The time-series feature matrix;

[0115] S312, Hidden Layer: The hidden layer uses a temporal convolutional network to extract features from the input sequence. For the first... Layered convolution, the convolution operation adopts a causal convolution structure to ensure that the output depends only on the current and historical time step information, and its output is represented as:

[0116] ;

[0117] in, For the first Hidden layer output, , For the first Layer convolution kernel parameters and bias terms, For convolution operators, For activation functions;

[0118] After extracting temporal features, an attention mechanism is introduced to learn the weighted contributions of different key factors. The attention weights are calculated as follows:

[0119] ;

[0120] The weighted feature is represented as follows:

[0121] ;

[0122] in, For the first Feature representation vectors of key factors For attention weights, For attention query vectors, Representing vectors transpose, The feature representation after attention weighting. The result is the linear transformation of the output layer. For the first Attention scores for key factors, exponential function Used to map attention scores to positive values ​​and enhance the difference between larger scores;

[0123] S313, Output Layer: The output layer uses the Sigmoid function to map the features extracted by the hidden layer to red tide occurrence probability values ​​between 0 and 1, expressed as:

[0124] ;

[0125] ;

[0126] in, This indicates the probability of a red tide occurring in the target sea area during the forecast period. The result is a linear change in the output layer. These are the output layer weight parameters, with values ​​ranging from -0.05 to 0.05. This is the output layer bias parameter, with a value range of -0.1 to 0.1;

[0127] S32: Using the actual occurrence of historical red tide events as labels, a cross-entropy loss function is constructed to conduct supervised training of the probability prediction network, expressed as:

[0128] ;

[0129] in, The label represents the actual occurrence of red tide; 1 indicates occurrence, and 0 indicates no occurrence. The loss function;

[0130] To measure the difference between the model's predictions and the actual situation, a cross-entropy loss function is introduced. The cross-entropy loss function measures the degree of difference between the model's predictions and the actual labels. When the actual label is 1, the model outputs the probability of red tide occurrence. The closer the value is to 1, the smaller the corresponding loss value; when the actual label is 0, the model outputs the probability of red tide occurrence. The closer it is to 0, the smaller the corresponding loss value; when the deviation between the model output probability and the actual label increases, the cross-entropy loss function value increases accordingly.

[0131] The backpropagation algorithm is used to obtain the prediction result and corresponding loss through forward computation. Then, starting from the output layer, the error signal generated by the loss is propagated forward layer by layer to the hidden layer and the input layer. In each layer, according to the computational structure of that layer, the contribution of the layer's parameters to the final loss is calculated, i.e., the gradient information of each parameter.

[0132] After obtaining the gradients of each parameter, the network parameters are updated using the gradient descent method, as follows:

[0133] ;

[0134] Repeat the above training process until the loss function converges or the preset number of training rounds is reached;

[0135] in, For network parameter set, Represents the loss function For network parameter set The gradient is used to characterize the sensitivity of each network parameter to changes in the loss function; its value reflects the rate of change of the loss function when each parameter undergoes a small change in a certain direction under the current parameter state. To determine the learning rate, the preset number of training rounds is generally set to 50-200 rounds;

[0136] The convergence of the loss function is determined in the following way:

[0137] (1) Based on the change in loss: when the change in the loss function is less than the preset change threshold in several consecutive iterations At this point, the model is considered to have converged:

[0138] ;

[0139] in, Indicates the first The loss function value calculated at the end of the iteration. For the training iteration count index, The preset threshold for the range of change, The range of values ​​is ;

[0140] (2) Based on validation set performance: Stop training when the model’s loss on the validation set no longer decreases or the prediction accuracy no longer improves.

[0141] S33: The key factor threshold representation set obtained after processing the real-time monitoring data through steps S1-S2. Input the trained probability prediction network to obtain the probability of red tide occurrence in the target sea area during the prediction period;

[0142] The trained probability prediction network inputs the key factor threshold representation set into a temporal convolutional network, extracts the evolution features of each target key factor at continuous time steps, learns the relative contribution of different target key factors to the occurrence of red tide through an attention mechanism, and performs weighted fusion of the extracted multi-factor temporal features. Finally, the output layer maps the probability value of red tide occurrence in the target sea area within the preset forecast period.

[0143] Based on the probability value of red tide occurrence The risk level is determined by combining preset multi-level early warning thresholds, and is expressed as follows:

[0144] ;

[0145] Finally, red tide warning information is generated and distributed through visual map display, SMS push or API interface;

[0146] in, A set of key factor threshold representations for real-time input. To predict the probability of red tide occurrence, The first warning threshold, set at 0.3, is used to identify the transition zone from low to medium risk. When the probability of red tide occurrence is low, it indicates that the target sea area is less likely to experience a red tide during the forecast period, and strong warning measures are usually not required. However, if the probability has reached a certain level, it indicates that environmental conditions have a certain tendency for red tide formation, and attention needs to be paid to this situation. If the value is too low, it is easy to misjudge a large number of weak fluctuation samples as medium risk, resulting in too frequent warnings. If the value is too high, it may miss the risk signals in the early stages of red tide occurrence, which is not conducive to early warning. The second warning threshold, set at 0.6, is used to identify the transition zone from medium to high risk. When the probability of red tide occurrence reaches a high level, it indicates that multiple key environmental factors have shown strong synergistic suitability, and the possibility of red tide occurring in the target sea area during the forecast period has significantly increased. At this time, a high-risk warning should be issued to prompt management departments to take timely measures such as increased monitoring, enhanced patrols, or emergency preparedness. If the value is too low, the high-risk judgment will be too lenient, which may result in too many high-level warnings. If the value is too high, some samples that already show a clear trend of red tide formation may remain at the medium-risk level, thus affecting the timeliness of the warning. It is at the warning level.

[0147] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0148] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for intelligent forecasting of red tide occurrence probability based on neural networks and key factor identification, characterized in that, Includes the following steps: S1: Acquire historical red tide event data, multi-station marine environmental monitoring data, red tide emergency monitoring data, and continuous hydrological and meteorological observation data of the target sea area. Perform anomaly removal, missing data repair, time alignment, and unified archiving on all types of data to obtain a standardized monitoring dataset. Based on the standardized monitoring dataset and historical red tide event data, calculate the red tide risk characterization results for each sea area unit, and identify key areas for red tide forecasting based on the red tide risk characterization results. Multi-time monitoring sequences corresponding to the key areas of red tide forecasting are extracted from the standardized monitoring dataset to form a time series sample set of key areas. S2: Perform multi-scale correlation analysis, sensitivity analysis, and causal association identification on the time-series sample set of the key areas to obtain the ranking results of key factors; based on the ranking results of key factors, extract the target key factors affecting the occurrence of red tides, and determine the threshold boundary range of each target key factor by combining the monitoring changes before, during, and after the occurrence of red tides; according to the target key factors and the threshold boundary range, perform threshold state encoding on the time-series sample set of the key areas to generate a key factor threshold characterization set; specifically including: S21: Wavelet coherence analysis is used on the time series sample set of the key area to extract the time-frequency correlation features between environmental factors and red tide occurrence at different time scales, and to identify environmental factors that are significantly related to red tide events. S22: Based on the random forest feature importance assessment or mutual information method, calculate the contribution of each environmental factor to the occurrence of red tide, sort them according to the size of the contribution, and select the environmental factors with the highest number of sorted environmental factors as sensitive environmental factors. S23: For the screened highly sensitive environmental factors, Granger causality test is used to identify the causal relationship between the highly sensitive environmental factors and the occurrence of red tides, eliminate pseudo-correlation factors that are only correlated, and obtain the ranking results of key factors. S24: Based on the ranking results of the key factors, select several environmental factors that rank highly as target key factors; using the time of occurrence of historical red tide events as a benchmark, statistically analyze the range of numerical changes of each target key factor before, during and after the occurrence of red tide, and use the quantile method to determine the threshold boundary range of each target key factor. S25: Based on the threshold boundary range, perform state encoding on each target key factor in the time series sample set of the key region, use fuzzy membership function to calculate the membership value of the key factor value relative to the threshold range, and generate a key factor threshold representation set including time series state information. S3: Input the key factor threshold representation set into the probability prediction network. The probability prediction network performs collaborative learning on the combination relationship and temporal evolution relationship of the key factors and outputs the red tide occurrence probability result. Based on the red tide occurrence probability result, generate red tide occurrence early warning information for the target sea area within the forecast period.

2. The intelligent forecasting method for red tide occurrence probability based on neural network and key factor identification according to claim 1, characterized in that, S1 specifically includes: S11: Collect historical red tide event records, multi-station marine environmental monitoring data, red tide emergency monitoring data, and continuous hydro-meteorological observation data from coastal stations in the target sea area; S12: Preprocess various types of data to obtain a standardized monitoring dataset; S13: Based on the standardized monitoring dataset and the historical red tide event data, a spatiotemporal clustering algorithm is used to divide the sea area into units, and the historical red tide occurrence frequency and environmental factor comprehensive index of each sea area unit are calculated to obtain the red tide risk characterization results; according to the preset risk threshold, key areas with high red tide risk are identified. S14: Extract continuous monitoring sequences corresponding to the key areas from the standardized monitoring dataset, and extract subsequences for multiple time periods before, during, and after the red tide event, based on the red tide event occurrence time, to construct a time series sample set for the key areas.

3. The intelligent forecasting method for red tide occurrence probability based on neural network and key factor identification according to claim 2, characterized in that, S12 specifically includes: Box plots were used to identify and remove outliers that were outside the normal range. Missing data are filled in using linear interpolation based on time series data or spatial interpolation methods based on adjacent stations. Data from different sources is resampled to the same time resolution, and time delays caused by differences in sampling frequency are corrected. The processed data is converted into a standardized format and a database is established to obtain a standardized monitoring dataset.

4. The intelligent forecasting method for red tide occurrence probability based on neural network and key factor identification according to claim 3, characterized in that, The box plot method specifically includes sorting the time series data of any monitoring variable, calculating the first quartile, the third quartile, and the corresponding interquartile range, using the first quartile minus a preset multiple of the interquartile range as the lower threshold, and the third quartile plus a preset multiple of the interquartile range as the upper threshold, and identifying monitoring data below the lower threshold or above the upper threshold as outliers and removing them.

5. The intelligent forecasting method for red tide occurrence probability based on neural network and key factor identification according to claim 1, characterized in that, The Granger causality test constructs a baseline prediction model based solely on historical data of red tide occurrence characterization sequences, and an extended prediction model that incorporates historical sequences of environmental factors into the baseline prediction model. By comparing the prediction errors or fitting effects of the baseline and extended prediction models, it is determined whether the historical information of environmental factors can significantly improve the prediction ability of red tide occurrence. When the model's prediction ability is significantly improved after incorporating the historical sequences of environmental factors, it is determined that the environmental factors have a Granger causal relationship with red tide occurrence and are retained as key target factors; otherwise, they are determined to be pseudo-correlated factors and are removed.

6. The intelligent forecasting method for red tide occurrence probability based on neural network and key factor identification according to claim 1, characterized in that, The fuzzy membership function calculates the corresponding membership value based on the deviation between the actual value of the target key factor and the optimal threshold at each time point, so as to reflect the suitability contribution of the target key factor to the occurrence of red tide at the corresponding time point. The closer the target key factor is to the optimal threshold, the higher the membership value; the further it deviates from the optimal threshold, the lower the membership value.

7. The intelligent forecasting method for red tide occurrence probability based on neural network and key factor identification according to claim 1, characterized in that, S3 specifically includes: S31: Construct a probabilistic prediction network including an input layer, a hidden layer, and an output layer; S32: Using the actual occurrence of historical red tide events as labels, the probability prediction network is trained in a supervised manner using the cross-entropy loss function, and the network parameters are optimized through the backpropagation algorithm until the model converges. S33: Input the key factor threshold representation set obtained after processing the real-time monitoring data through steps S1-S2 into the trained probability forecasting network to obtain the probability value of red tide occurrence in the target sea area within the preset forecast period.

8. The intelligent forecasting method for red tide occurrence probability based on neural network and key factor identification according to claim 7, characterized in that, The input layer includes the key factor threshold representation set, which includes the threshold state encoding sequence of each target key factor at multiple time steps; the hidden layer uses a temporal convolutional network structure to extract temporal features from the input sequence, and learns the combination relationship and weight allocation between different key factors through an attention mechanism; the output layer uses the sigmoid function to map the deep features learned by the temporal convolutional network into red tide occurrence probability values ​​and outputs the probability results.

9. The intelligent forecasting method for red tide occurrence probability based on neural network and key factor identification according to claim 7, characterized in that, S33 further includes: Based on the red tide occurrence probability value and combined with the preset multi-level warning thresholds, a corresponding red tide warning level is generated; The obtained warning level information will be distributed to relevant management departments and public users through visual map display, SMS push or API interface.

Citation Information

Patent Citations

  • CN119623766A

  • CN121602355A