An annual green tide prediction method and system fusing space scene and prediction nodes
Patent Information
- Application Number
- CN202611319213.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明的目的在于提供一种融合空间场景与预测节点的年度绿潮预测方法及系统,以解决固定空间范围、固定统计时段导致预测信息利用不足,以及模型评价与目标年度信息边界不清的问题
第一,种源边界等权、历史可达频次和高频显现斑块采用不同生成规则,避免把多个空间情景误解为同一范围的机械外扩,并能够比较不同空间信息对年度预测的贡献。
Smart Images

Figure CN122839307A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of marine ecological disaster prediction technology, specifically involving an annual green tide prediction method and system that integrates spatial scenarios and prediction nodes. Background Technology
[0002] The annual scale of the green tide of *Ulva prolifera* in the Yellow Sea is influenced by a combination of factors, including early seed source conditions, nutrient supply, light and heat conditions, wind and wave transport, and meteorological disturbances. Daily satellite remote sensing can identify the distribution of the green tide, but disaster prevention operations require assessing the overall spatial scale that may be affected that year before the outbreak season or at the beginning of the outbreak in order to pre-allocate monitoring and response resources.
[0003] Existing annual or medium- to long-term forecasting methods typically pre-fix the study area and environmental statistical period, and then establish regression or machine learning models based on regional average environmental factors. However, the interannual forecast information carried by different spatial extents is not the same; the same environmental variable may also possess both short-term and long-term memories. If the spatial region or forecast month is fixed in advance, it is difficult to simultaneously consider cross-year stability and forecast lead time.
[0004] For example, CN114611800B discloses a method for predicting the medium- and long-term trend of the Yellow Sea green tide, which establishes a prediction relationship using a preset region, fixed factors, and fixed time periods. This type of scheme does not use scene weights generated based on different spatial information and multiple candidate prediction nodes as a joint configuration for evaluation under strict forward time boundaries, nor does it use prediction performance, error stability, process feature reproducibility, and lead time to jointly determine the target configuration.
[0005] Therefore, there is a need for an annual green tide forecasting method that can compare different spatial information aspects, automatically identify reliable forecast nodes, and does not use future labels and post-node information in the target year. Summary of the Invention
[0006] The purpose of this invention is to provide an annual green tide prediction method and system that integrates spatial scenarios and prediction nodes, so as to solve the problems of insufficient utilization of prediction information due to fixed spatial range and fixed statistical period, as well as unclear boundaries between model evaluation and target annual information.
[0007] The first objective of this invention is to provide an annual green tide prediction method that integrates spatial scenes and prediction nodes, comprising: S1. Obtain the daily green tide detection results of the outbreak season in historical years, spatially combine the green tide detection grids of the valid observation days in the same year, calculate the annual independent coverage area, and form a historical annual scale label. S2. Based on different spatial information or weight generation rules that distinguish them from each other, construct at least two candidate spatial scene weights; S3. Acquire environmental data, perform spatial weighting according to the weight of each candidate spatial scenario, and form the environmental time series corresponding to each candidate spatial scenario. S4. Set up candidate prediction nodes, and construct ecological memory features and process features based on the environmental time series within the information boundary of each candidate prediction node. S5. Combine candidate spatial scenarios and candidate prediction nodes into candidate configurations; for any year to be evaluated, use only the scale labels before that year and the environmental data up to the candidate nodes to fit a prediction model that can output the process feature action state and obtain the prediction value; summarize the prediction results and process feature action states of multiple years to be evaluated to obtain the prediction performance, error stability, feature selection and action direction records of each candidate configuration. S6. Based on the prediction performance, error stability, process feature reproducibility determined by process feature selection and action direction recording, and minimum prediction lead, candidate configurations are screened to determine the target spatial scene and target prediction node. S7. For target years for which the true scale label has not yet been obtained, construct spatial weights according to the target spatial scenario rules, use historical scale labels and environmental data up to the previous year, redetermine the target year memory window and fit the prediction model, and output the target year green tide scale prediction value using only the target year environmental data up to the target prediction node.
[0008] Preferably, the candidate spatial scene weights include at least two of the following categories: The seed source category scene weight is obtained by rasterizing the candidate seed source boundary. The weight of each effective grid within the candidate range is 1, and the weight outside the candidate range is 0. The weights for transportation scenarios are obtained by accumulating the annual occurrence counts of each grid cell based on the green tide annual grid cells in previous historical years before the year to be evaluated or the target year, and then normalizing and spatially smoothing them. The weights for display-type scenes are obtained by performing high-frequency filtering and connected patch identification based on the annual occurrence frequency, and by assigning weights to the retained patches according to the annual occurrence frequency. Among them, the spatial ranges of different spatial scenarios are allowed to overlap; the seed source scenario is generated based on the candidate boundary, and the transport and manifestation scenarios are allowed to use the historical annual occupancy information, but they adopt different weight conversion rules and are not formed by mechanically cutting, expanding or copying the seed source boundary.
[0009] Preferably, the annual independent coverage area is the union area of the grid spaces detected as green tides on each valid observation day of the outbreak season in the same year. The same grid is counted only once, regardless of how many times it is detected in the same year. The annual occupancy status is determined only when the grid meets the preset annual valid observation conditions. Grids that do not meet the preset annual valid observation conditions are marked as missing measurements and are not considered as non-occurring grids.
[0010] Preferably, the ecological memory features are obtained by setting a lag length and a cumulative length for the environmental time series prior to the candidate prediction node; the process features include at least one of nutrient supply, light and heat conditions, wind transport, wave persistence and disturbance forcing, and their interaction features.
[0011] Preferably, the prediction performance includes at least two of the following: coefficient of determination, root mean square error, and mean absolute relative error; the process feature reproducibility is determined based on the selection frequency and consistency of the direction of action of the process feature in several effective year-by-year forward evaluation rounds; the candidate configuration is marked as a valid configuration only when the prediction performance, error stability, process feature reproducibility, and minimum prediction lead all meet preset conditions.
[0012] Preferably, the prediction model is a sparse regression model that can output the state of feature action, that is, a regression model that shrinks some feature coefficients to zero through sparse constraints, thereby recording the feature selection state and the direction of non-zero coefficients; when performing a logarithmic transformation on the annual independent coverage area, the inverse transformation correction factor is determined based on the out-of-sample residuals before the target year to obtain the area-scale prediction value.
[0013] Preferably, in the forward testing implementation of the frozen configuration structure, the labeled historical data is divided into a basic training period, a configuration screening and freezing period, and a forward testing period. Candidate configurations are compared only during the configuration screening and freezing period. After entering the forward testing period, the spatial scene type and generation rules, target prediction nodes, feature identities and construction rules, memory window search rules, and model types are frozen. The configuration structure is not reselected based on the test year results. The standardized parameters, memory window, and model parameters for each test year are only re-determined using the real labels formed before that test year.
[0014] Preferably, it also includes: S8, after the true scale label of the target year is formed, calculate and archive the prediction error of that year; review the current configuration according to the same forward rule as the historical evaluation; if the current configuration still meets the preset conditions, maintain the configuration structure and add new labels to refit the model parameters; if the current configuration degrades, re-evaluate the candidate configuration, and redetermine the target configuration only if the new candidate configuration meets the preset improvement conditions.
[0015] Preferably, after obtaining the target year's real-scale label, the target spatial scene type, target prediction node, and feature construction rules are kept unchanged, and the health of the current configuration is re-evaluated according to the historical forward evaluation rules; if the current configuration has not degraded, new labels are added to refit the model parameters; the candidate configuration evaluation is only re-executed when the current configuration degrades, and the current configuration is replaced only when the candidate new configuration meets the preset improvement conditions relative to the current configuration.
[0016] A second objective of this invention is to provide an annual green tide prediction system that integrates spatial scenarios and prediction nodes, comprising: The tag building module is used to generate historical year-scale tags based on the daily green tide detection results of historical years; The scene weight construction module is used to construct candidate spatial scene weights based on different spatial information or weight generation rules that distinguish them from each other. The environmental sequence extraction module is used to spatially weight environmental data according to the candidate spatial scene weights to form an environmental time series. The memory and process feature module is used to construct ecological memory features and process features within the information boundaries of candidate prediction nodes; Configure a joint evaluation module to combine candidate spatial scenes and candidate prediction nodes into candidate configurations and perform year-by-year forward evaluation; The target configuration determination module is used to determine the target spatial scene and target prediction nodes based on prediction performance, error stability, process feature reproducibility and minimum prediction lead. The annual forecast module is used to output the target annual green tide scale forecast value using only the target annual environmental data before the target forecast node.
[0017] Preferably, it also includes an update evaluation module, which archives the prediction error and reviews the current configuration after the target year's true scale label is formed, and updates the target configuration when the current configuration degrades and the candidate new configuration meets the preset improvement conditions.
[0018] The beneficial effects of this invention are: First, different generation rules are used for seed source boundary equal weight, historical reachability frequency and high-frequency manifestation patches to avoid misinterpreting multiple spatial scenarios as mechanical expansion of the same range, and to compare the contribution of different spatial information to annual forecasts.
[0019] Second, in this invention, the spatial scene and prediction node are jointly evaluated, rather than a specific source region or month is pre-specified, which enables the formation of an executable selection rule between prediction accuracy and prediction lead time.
[0020] Third, this invention will strictly limit the label cutoff year and node data cutoff time for the year-by-year forward evaluation to prevent information from subsequent years from entering the early model selection.
[0021] Fourth, this invention reduces the risk of selecting the accidental optimal configuration by relying solely on a single accuracy index through joint gating of predictive performance, error stability, and process characteristic reproducibility.
[0022] Fifth, this invention separates the target year prediction from the historical configuration evaluation, which can be directly used for future year predictions where the actual annual area has not yet been obtained, and can form a traceable record containing spatial scenes, prediction nodes and historical cut-off years. Attached Figure Description
[0023] Figure 1 A flowchart of a preferred embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the construction relationship of differentiated spatial scene weights in a preferred embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the relationship between the enhanced implementation method and the basic prediction method in a preferred embodiment of the present invention; Figure 4 This is a block diagram of the annual green tide forecasting system and its operational phases in a preferred embodiment of the present invention; Figure 5 This is a schematic diagram illustrating the relationship between the memory window, the year-by-year forward evaluation, and the target year prediction in a preferred embodiment of the present invention. Figure 6 This is a flowchart of the target year label-free prediction in a preferred embodiment of the present invention; Figure 7 This is a flowchart illustrating the evaluation and controlled update process after the actual tag arrives in a preferred embodiment of the present invention. Detailed Implementation
[0024] The present invention will be further described below with reference to the accompanying drawings. The embodiments described are for illustrative purposes only and are not intended to limit the scope of protection. Without changing the technical chain of "constructing spatial scenes with different rules—multiple prediction nodes—forward temporal evaluation—reliability gating—unlabeled prediction for the target year—controlled update after prediction," the grid resolution, environmental variables, candidate nodes, threshold, and regressor can all be adjusted according to the data conditions.
[0025] This invention performs annual spatial union of daily green tide monitoring results during the outbreak season to form an annual independent coverage area label. Based on candidate source boundaries, historical reachability frequencies, and historical high-frequency manifestation patterns, spatial scene weights with different rules are generated. Marine and meteorological environmental time series are extracted according to spatial weights, and ecological memory and process characteristics are constructed within the information boundaries of candidate prediction nodes. Spatial scenes and prediction nodes are combined into candidate configurations, and forward evaluation is used year by year to obtain prediction performance, error stability, and process characteristic reproducibility. Target spatial scenes and target prediction nodes are determined by combining the minimum prediction lead. For the target year, label-free prediction is performed using only environmental data prior to the nodes and historical labels prior to that year. After the true label for the target year is formed, prediction errors are archived and the prediction framework is updated in a controlled manner. This invention avoids pre-fixing spatial areas and prediction months, and can improve the stability, lead time, and traceability of annual green tide scale prediction.
[0026] I. Symbols and Time Boundaries: To avoid mixing up historical modeling years, annual evaluation years, and actual business target years, this embodiment uses the following symbols uniformly: y represents a historical year with a known actual annual scale label; u represents a year temporarily considered unknown during forward evaluation; T represents the target prediction year in actual business. In business forecasting, u = T. p represents the spatial grid, d represents the outbreak season date, t represents the environmental data date, v represents environmental variables, s represents the spatial scenario, k represents candidate prediction nodes, and j represents the candidate configuration number. This indicates the number of days lags behind the prediction node, a represents the cumulative window length, i represents the sequence number of the year-by-year forward evaluation, n represents the number of valid evaluation years participating in the calculation of the configuration index, and g represents the process feature sequence number.
[0027] For example, when forecasting the annual size for 2027 in January 2027, T=2027 and u=T; the actual annual size labels that can be used are 2008-2026, i.e., y=2008, ..., 2026; the actual area for 2027 and environmental data after the target forecast node for 2027 must not be used.
[0028] II. Basic Prediction Implementation Methods: Please see Figure 1 A core prediction process for an annual green tide prediction method integrating spatial scenes and prediction nodes includes S1-S7, used to form a predicted value for the annual green tide scale when the true label of the target year is unknown; in a preferred post-prediction update implementation, it also includes S8, which is executed after the outbreak season of the target year ends and the true label is formed, used to archive prediction errors and determine whether to update model parameters or re-evaluate prediction configuration. Specifically: S1. Construct historical annual scale labels: Obtain daily green tide detection results for multiple historical annual outbreak seasons, spatially combine the green tide detection grids of valid observation days within the same year, calculate the annual independent coverage area, and form historical annual scale labels. S2. Construct differentiated spatial scene weights: Based on different spatial information or weight generation rules that distinguish them from each other, construct multiple candidate spatial scene weights; S3. Extract the time series of scene environment: Obtain marine environment data and meteorological environment data, and perform spatial weighting according to the weight of each candidate spatial scene to form the environmental time series corresponding to each candidate spatial scene; S4. Constructing ecological memory and process features: Set up multiple candidate prediction nodes, and construct ecological memory features and process features based on the environmental time series within the information boundary of each candidate prediction node. S5. Joint Evaluation of Spatial Scenarios and Prediction Nodes: Combine candidate spatial scenarios and candidate prediction nodes into candidate configurations; for any year to be evaluated in the year-by-year forward evaluation, use only the historical year scale labels before that year and the environmental data up to the corresponding candidate prediction node to fit the prediction model and obtain the prediction value for that year; summarize the prediction results of multiple years to be evaluated and the status of process characteristics in each round to obtain the prediction performance, error stability index, and process characteristic selection and direction of action records for each candidate configuration. S6. Determine the target spatial scene and target prediction node: Based on the prediction performance, cross-year error stability, process feature reproducibility determined by the process feature selection and action direction record, and minimum prediction lead, the candidate configurations are screened to determine the target spatial scene and target prediction node. S7. Target Year Unlabeled Prediction: For target years that have not yet obtained a true annual scale label, construct the spatial weight corresponding to the target year according to the generation rules of the target spatial scene, use the historical annual scale label up to the year before the target year and the corresponding historical environmental data to redetermine the target year memory window and fit the prediction model, and only use the target year environmental data up to the target prediction node to output the target year green tide scale prediction value. S8. Post-prediction evaluation and controlled update: After the true scale label for the target year is formed, calculate and archive the prediction error for that year; review the current configuration according to the same forward rules as the historical evaluation; if the current configuration still meets the preset conditions, maintain the configuration structure and add new labels to refit the model parameters; if the current configuration degrades, re-evaluate the candidate configurations, and redetermine the target configuration only if the new candidate configuration meets the preset improvement conditions.
[0029] To better understand the technical solution of the present invention, non-limiting examples are provided below for explanation: S1. Construct historical annual scale tags: S1.1 Sources of remote sensing and environmental data: This embodiment reconstructs the daily green tide distribution using MODIS Terra and Aqua Level-1A imagery from 2008 onwards. Atmospheric correction is performed using SeaDAS to generate a 1 km Rayleigh corrected reflectance product. Floating seaweed is then identified using the phytoplankton index (FAI) and a unified threshold. Environmental data utilizes 5-day reanalysis data from the European Centre for Medium-Range Weather Forecasts (ERA) and global marine biogeochemical history simulation products from the Copernicus Marine Service. All data are subsequently unified to a 0.25° daily grid and a consistent effective marine mask. Other sensors or products are used as alternative data sources only when they can provide equivalent daily green tide detection status or corresponding environmental variables.
[0030] S1.2 Daily monitoring and determination of effective observations: In this embodiment, May 1st to August 31st of each year is designated as the green tide outbreak season. B(y,d,p) represents whether a green tide was detected on date d of historical year y at grid p, with a value of 1 if detected and 0 if not detected. M(y,d,p) represents whether the grid on that day is a valid observation, with a value of 1 if valid and 0 if invalid. For grid days affected by cloud cover, land pollution, sensor missing data, or quality control failures, M(y,d,p) is set to 0, and the grid day is not assigned the "no green tide detected" B=0 state, nor is it included in the numerator calculation of the proportion of annual non-detection or valid observations.
[0031] When the effective coverage ratio of a certain remote sensing date in the research sea area reaches a preset requirement, that date is included in the annual composite; the example uses 60%. For a specific grid, if a green tide is detected on any effective observation day, its annual occupancy status can be judged as 1; if no green tide is detected, its annual occupancy status can only be judged as 0 if the number of effective observations or the effective observation ratio of the grid during the outbreak season reaches the preset annual observation requirement; otherwise, it is marked as missing data for the year. One feasible approach is to use the number of effective remote sensing dates included in the composite for that year as the denominator, requiring that the effective observation date ratio of the grid is not less than 60%. The "historical year that meets the annual effective observation conditions" referred to in this invention is the year in which the corresponding grid can form an effective O(y,p) value according to the above rules.
[0032] S1.3 Annual Spatial Union and Area Labels: The annual occupancy status O(y,p) and annual independent coverage area A(y) are calculated using the following formula: (1) It should be noted that: only when the grid... It is only calculated when the annual valid observation conditions are met. If the conditions for valid annual observation are not met, the observation will be marked as missing and will not be considered as such. No mesh occurred.
[0033] In the formula, For the year A set of dates within the outbreak season; For the year A set of grids that meet the conditions for effective annual observation; The function is an indicator function; it takes the value 1 if the condition is true, and 0 otherwise. a(p) represents the actual area represented by grid p. O(y,p)=1 indicates that grid p detected a green tide on at least one valid observation day in year y; regardless of how many times the same grid detected it in the same year, it is only counted once. The unit of A(y) is 10^- ... 4Square kilometers. Equation (1) only determines O(y,p) for grids that meet the annual effective observation conditions; grids with missing measurements in a year are not included as O(y,p)=0. If the proportion of effective observation grids in a year is insufficient to support the annual area calculation, the label for that year will not be included in the subsequent model evaluation.
[0034] S1.4 Correction for differences in observation opportunities: The annual occupancy frequency statistics simultaneously record the number of historical years with valid annual occupancy status for each grid, and set a minimum number of valid historical years. The main implementation method constructs Path and Bloom patterns based on the frequency of occurrences within valid years. When there are significant differences in the number of valid historical years across different grids, the frequency of occurrences can be divided by the number of valid observation years for that grid to form an annual occupancy ratio as an alternative weighting basis. This approach aims to correct for uneven observation opportunities: a grid with fewer occurrences may have fewer actual occurrences, or it may simply have fewer available remote sensing years; using the occupancy ratio can reduce the systematic underestimation caused by the latter. Whether this alternative method is enabled and its threshold are determined before modeling and maintained consistently within the same evaluation round.
[0035] This example demonstrates the following: the outbreak season is from May 1st to August 31st; the effective coverage threshold for the daily-scale research sea area is 60%; and the annual area unit is 10. 4 Square kilometers; the above values are executable implementation parameters and do not restrict the use of other predetermined data quality thresholds that are consistently implemented in each year.
[0036] S2. Construct differentiated spatial scene weights: Please see Figure 2 In this embodiment, four candidate spatial scenes are constructed: Source_A, Source_B, Path, and Bloom. These four scenes are not the same polygon that expands from the inside out, but rather parallel candidate scenes formed based on different spatial information or different transformation rules, and their spatial ranges may overlap.
[0037] S2.1 Binary weighted scenario for seed source class: Source_A primarily focuses on the candidate provenance range of the Subei Shoal; Source_B expands upon Source_A to the adjacent waters of the Yangtze River Estuary to examine whether the expanded range brings incremental prediction information. After rasterization of the candidate boundaries, all valid grids within the range have a weight of 1, while those outside the range have a weight of 0. The provenance class weights do not use historical occurrence frequency, are not adjusted based on distance from the boundary, and are not Gaussian smoothed.
[0038] S2.2 Path historical reachability frequency scenario: For each forward evaluation year u, the number of times the grid year occurs is first calculated using only the historical occupancy status of years prior to u: (2) F(u,p) represents the relationship between grid p and u from 2008. The number of years in which the effective annual occupied grid is determined to be affected by the green tide is calculated. Here, "effective year" follows the definition from step S1.2: the corresponding grid meets the annual observation requirements and has a valid O(y,p) value in that year. The same grid is counted only once in the same year; if the annual observation is invalid, that year is skipped and not counted as 0. Therefore, F(u,p) is not the daily occurrence count, the number of effective occurrence days, or the probability of occurrence. The path weight is obtained by normalizing and spatially smoothing F(u,p) to emphasize the historically accessible space and its relative occurrence intensity.
[0039] One directly implementable method for path generation is as follows: First, calculate P(u,p)=F(u,p) / max[F(u,q)]. Here, P(u,p) represents the initial normalized weight of the path in grid p under the year u to be evaluated, without spatial smoothing, and its value ranges from 0 to 1; q is used to traverse all effective grids in the research area. Then, convolve P(u,p) using a two-dimensional Gaussian kernel of a preset scale, and normalize the result to 0-1 to obtain W(Path,u,p). W(Path,u,p) represents the final spatial scene weight of the path after Gaussian smoothing, effective domain constraint, and normalization. Step S3 uses this weight to extract the environmental time series. To prevent the smoothing weights from extending to grids that have never appeared in history, an effective domain mask with F(u,p)>0 can be reapplied after smoothing. In this embodiment, the Gaussian smoothing scale is 2 grids; if max[F(u,q)]=0, then the path scene for that year has no historical construction basis and is directly marked as invalid, not replaced by a matrix of all 1s.
[0040] S2.3 Bloom high-frequency connectivity display scenario: Bloom uses the same F(u,p) but employs a different transformation rule than Path. First, candidate grids are filtered based on a high-frequency threshold. Then, connected patches are identified using eight-neighbor relationships, and isolated patches smaller than a preset size are removed. Next, the original annual occurrence frequency differences within each qualified patch are preserved, and the occurrence frequencies of the retained grids are normalized to 0-1 and spatially smoothed to obtain continuous Bloom weights. To strictly limit the weight range, a qualified patch mask can be reapplied after smoothing and normalized again. If no high-frequency patches meet the basic size requirements, the Bloom scene for that year is marked as invalid, rather than being replaced by arbitrary small patches. This ensures that Bloom represents historically high-frequency and spatially continuous display areas, rather than a small number of isolated pixels.
[0041] S2.4 Scene Registration and Annual Updates: Source, Path, and Bloom use the same grid size, latitude and longitude orientation, and effective land and sea area. After generation, they are registered separately for scene type, fixed boundary or dynamic frequency attribute, historical cutoff year, number of non-zero grids, weight value range, and parameter version. Source is generated from candidate boundaries; Path and Bloom can share the historical occurrence count, but they use continuous frequency weighting rules and high-frequency connected patch screening weighting rules respectively, and do not use specific Source boundaries for mechanical clipping, expansion, or duplication.
[0042] For Path and Bloom, the historical cutoff year is expanded to the next year for each forward-looking or evaluation year. 1. Weights are regenerated; what is frozen is the scene type and generation rules, not the old matrix formed in a certain historical year. Seed source scenes are fixed boundary scenes, and are re-rasterized according to the predetermined boundaries when the unified grid changes. For any year u to be evaluated, only information before u is used to generate the spatial weight W(s,u,p) for this round; in this round of model, the environmental features with the same name in the training year and the year to be evaluated are all given this weight and the same extraction rules, and the occupancy information of year u or subsequent years cannot be used to generate weights.
[0043] In this example, Bloom uses an annual occurrence threshold of 8, a minimum number of connected pixels of 20, and a Gaussian smoothing scale of 2 grids; Path uses a Gaussian smoothing scale of 2 grids. All of these parameters are registered with the scene construction rules to facilitate reconstruction using the same caliber in subsequent years.
[0044] S3. Extract the time series of the scene environment: S3.1 Original environmental variables and their corresponding processes: This embodiment uses ten daily-scale raw environmental variables. In Copernicus Marine Service products, nitrate concentration (NO3), phosphate concentration (PO4), and silicate concentration (Si) correspond to nutrient supply or seed source memory; surface solar radiation (SSR) and surface temperature (SKT) in ERA5 correspond to photothermal conditions; 10-meter zonal wind (U10) and meridional wind (V10) correspond to wind transport; mean wave period (MWP) corresponds to wave persistence; and total precipitation (TP) and mean sea level pressure (MSL) correspond to disturbance forcing.
[0045] The above process categories are ecological semantic classifications of variables, used for constructing process features in S4 and explaining feature reproducibility in S6. They do not treat original variables within the same category as a common "standardized group." Each process feature that ultimately enters the model calculates its standardized parameters independently using only the current training set, avoiding the inclusion of information from different dimensions or the year to be evaluated in the standardization process.
[0046] S3.2 Spatial Weighted Extraction: Marine biogeochemical products, marine reanalysis data, and meteorological reanalysis data are unified into a spatial grid consistent with the scene weights. E(v,t,p) represents the environmental value of variable v at date t and grid p, and Q(v,t,p) indicates whether the environmental value is valid. The scene-weighted environmental value X(s,u,v,t) is calculated using the following formula: (3) In the formula, the numerator is the weighted sum of "weight × environmental value" for each effective grid, and the denominator is the total number of weights with effective environmental data on that day. Dividing the two yields the weighted average within the scene, rather than the total weight that varies with the number of effective grids. This ensures that even if the number of effective grids varies slightly across different dates, the resulting time series maintains the same physical dimensions. W(s,u,p) represents the spatial scene weights constructed for year u; within the same evaluation round, these weights are used simultaneously to extract environmental sequences for both the training year and year u, maintaining consistency in the spatial definition of features with the same name.
[0047] S3.3 Missing Measurement Handling and Node Boundaries: If the weighting ratio of valid environmental data coverage on a given day is lower than a preset value, that day will be recorded as missing data. When supplementary data is needed, only data from the corresponding prediction node and earlier days are allowed for unidirectional supplementation or other causal processing; any interpolation, smoothing, and resampling must not use data after the prediction node. If processing cannot be done based solely on data before the node, that day will not participate in window statistics.
[0048] In this example, the ERA5 and Copernicus environmental variables are unified to a 0.25° daily scale grid, and the effective environmental weight ratio threshold is set to 80%. The output is a daily scale weighted environmental time series corresponding to each evaluation year u, spatial scene s, and environmental variable v.
[0049] S4. Constructing ecological memory and process characteristics: S4.1 Candidate prediction nodes and memory windows: Please see Figure 5 The candidate prediction nodes are set to September 1st, October 1st, November 1st, and December 1st of the previous year, and January 1st, February 1st, March 1st, April 1st, and May 1st of the current year, for a total of 9 nodes. For any scenario s, the year to be evaluated u, the environmental variable v, and the candidate node k, a delay length is set. Let the cumulative length be a; let the candidate node date be t(u,k), then the memory feature Z(s,u,v,k); a) is: (4) Therefore, the end date of the cumulative window is no later than the candidate prediction node, and the window selection for the year u to be evaluated can only use the true labels before u.
[0050] S4.2 Single-Variable Window Search and Invalid Window Handling: For each environment variable, the memory window is searched using only data from the current training year set in each training round. This is done for candidate lag lengths. Given the candidate cumulative length 'a', first extract the environmental statistical features corresponding to the window within each training year. Only when the proportion of valid dates within the window in that year meets the preset requirements, consecutive missing data does not exceed the allowable range, and all data used are located before the corresponding prediction node, is that year counted as a valid year for the candidate window. Then, using the environmental feature sequence formed by each candidate window within the valid year, calculate the window determination coefficient R²(win) with the corresponding year area label, and select the one with the largest R²(win). a) The combination serves as a single-channel memory window for this training round and this environment variable.
[0051] When the difference in R²(win) among multiple candidate windows does not exceed the preset numerical precision, the windows with the most valid years, the shortest cumulative length 'a', and the shortest lag length are selected in sequence. Smaller combinations. The number of valid years refers to the number of years within the current training year set where the candidate window satisfies the requirements of data integrity, consecutive missing tests, and pre-node information boundaries, and successfully generates window features.
[0052] If the proportion of valid dates for a candidate window is less than 80%, has more than 3 consecutive days of missing data, cannot be processed based solely on data prior to the node, or its statistic remains unchanged between training years, then it will be removed from the candidate window set for that environmental variable in this round and will not participate in the R² (win) comparison. If all candidate windows for an environmental variable are invalid, then that variable will not be included in the feature construction for this round; if the remaining valid variables are insufficient to form the input for the preset model, then the corresponding "candidate configuration - year to be evaluated" will be marked as invalid, and windows will not be forcibly filled. Window search must not use data from the year to be evaluated or subsequent years, and the windows obtained in each round are only used for out-of-sample prediction in that round.
[0053] S4.3 Process Feature Merging Each original environmental variable first determines its own hysteresis-cumulative window according to step S4.2. Different variables can have different hysteresis lengths and cumulative lengths. For the current spatial scenario... Candidate prediction nodes and any training year Calculate the variable-level memory values separately for each variable according to the selected window. For example, NO3, PO4, and S i Formed separately , and Although the three memory values come from different time windows, they all represent an annual feature value of the corresponding variable formed in the same year, the same spatial scenario, and the same prediction node. Therefore, they can be merged along the annual dimension.
[0054] In this embodiment, the unstandardized value of nutrient provenance memory is... The same year's NO3, PO4 and S i The memory values are obtained by taking the arithmetic mean. This means that the window calculations for the three variables are performed separately first, and then the average of the memory values for the three years is calculated. It is not done by directly averaging the daily data for the three variables or by setting a common window for the three variables. Year to be evaluated Calculate using the same rules based on the three variable windows determined during the training phase. If invalid memory values exist among the constituent variables, they will only be merged when the number of valid constituent variables reaches a preset lower limit; if the lower limit is not reached, the nutritional germplasm memory for that year will be marked as missing.
[0055] The remaining basic process features are constructed using the same method of "first forming variable-level memory values, then merging them according to process relationships." Among these, the unstandardized value of photoactivation... SSR memory values were used, and the temperature background was not standardized. Using SKT memory values, wind transport unstandardized values The vector modulus is calculated from the U10 and V10 memory values, and the wave persistence is calculated using unstandardized values. Using MWP memory values, perturbation-forced unstandardized values The values are synthesized from TP and MSL memory values according to preset rules. For component variables with different dimensions that need to be synthesized with equal weights, they can be first made dimensionless using the current training set, and then merged. The merging method used is consistent throughout all forward evaluation rounds of the same candidate configuration.
[0056] After merging the six basic process features, three interactive features—seed source-light coupling, seed source-temperature coupling, and seed source-wind field coupling—were further constructed, each derived from... , and This yields nine unstandardized process characteristics, including nutrient source memory, light activation, temperature background, wind transport, wave persistence, disturbance forcing, and three interaction terms.
[0057] In each forward training cycle, the mean and standard deviation of the nine process features mentioned above are calculated using the current training year set, and each feature is standardized individually. Taking nutrient origin memory as an example, based on the data from each training year... Calculate the mean and standard deviation of the training set, and then combine the results for each training year. Transformed into standardized nutritional germplasm memory The year to be evaluated will be calculated first. Then, using the mean and standard deviation obtained from the same training set, convert... The year being evaluated must not be included in the determination of the mean, standard deviation, or pooled parameters. The remaining eight process characteristics are handled in the same manner.
[0058] This step ultimately outputs nine standardized process features in the same order for each training year and the year to be evaluated, which serve as inputs to the prediction model in step S5. Step S5 no longer re-searches the variable window or merges the original environmental variables; instead, it directly uses the standardized process features generated in this step to establish a predictive relationship between them and the annual independent coverage area.
[0059] S5. Joint Evaluation Spatial Scene and Prediction Nodes: S5.1 Candidate Configurations and Prediction Models: Each spatial scene s is combined with each candidate prediction node k to obtain a candidate configuration C(j). In the example with 4 spatial scenes and 9 candidate prediction nodes, a total of 36 candidate configurations are formed; 4×9 is only an example and does not limit the number of scenes or the node setting method.
[0060] Under each candidate configuration, a continuous variable regression model is used to establish the mapping relationship between ecological memory features, process features, and annual independent cover area. The model input consists of the standardized ecological memory features, process features, and their interaction features obtained in step S4; the model output is the annual independent cover area A or its logarithmic transformation value. Preferably, the prediction model uses a sparse linear regression model capable of outputting the feature action state, including Lasso regression or Elastic Net regression, to record non-zero features, coefficient signs, and feature selection states, and to provide data for step S6 to calculate the process feature reproducibility.
[0061] When using logarithmic scaling for annual area modeling, let the logarithmic area of the historically labeled year \(y\) be \(L(y)=\ln A(y)\). In a single-round forecast for the year \(u\) to be evaluated, the logarithmic scaling forecast relationship is expressed as: (5) In the formula, For the year to be evaluated The predicted annual area on a logarithmic scale; For the year to be evaluated The A standardized process characteristic; This represents the total number of process features entering the current prediction model. For the intercept term; For the first The regression coefficients correspond to the characteristics of each process. Both the intercept term and the regression coefficients utilize only annual data. This was obtained by fitting historical annual samples with previously obtained true area labels. Equation (5) is only used to obtain log-scaled predicted values. This does not directly represent the final predicted value at the area scale after inverse transformation and bias correction. When the prediction model directly uses the annual independent coverage area... When used as output, no inverse logarithmic transformation is performed; when the prediction model uses the logarithmic area... As output, in step S7.3, an inverse exponential transform is performed on the log-scale predicted value for the target year, and a bias correction factor is calculated using the out-of-sample log-residuals formed according to the forward training rules for previous years, thus obtaining the final predicted value for the area scale. In each forward training round, all candidate features with missing values or zero variance in the current training year set are first removed, and then the mean and standard deviation of each feature are calculated using only the current training year set, and feature standardization is completed. The regularization strength of Lasso regression or Elastic Net regression is determined by the forward time partitioning within the current training year set. Years to be evaluated are temporarily considered unknown. They must not participate in the determination of feature mean, feature standard deviation, memory window, regularization strength, regression coefficient, or other model parameters.
[0062] S5.2 Forward calculation year by year See Figure 5 The evaluation is performed year by year in chronological order. For the year u to be evaluated, the training set can only contain labeled years where y < u; Path and Bloom can only use years up to u. The annual occupancy grid for one year generates the current round's weight W(s,u,p), and this weight is used to uniformly extract the environmental features of the current training year and year u. Memory window selection, feature standardization, feature filtering, and model parameters are all completed within this training set. Subsequently, only the environmental data of year u up to candidate node k is used to generate predicted values. This rule is repeated for each u. In other words, each round re-executes "current round weight generation—training set window selection—training set standardization—model fitting—prediction of the year to be evaluated," and the final weights for all periods cannot be used to calculate earlier evaluation rounds, nor can the window or standardized parameters obtained in a later round be backfilled into the previous round.
[0063] The task of S5 is to perform year-by-year prediction calculations and generate evaluation data for each candidate configuration. The optimal spatial scenario or prediction node is not directly determined in this step. The annual output includes the actual value, predicted value, error, selection features, and coefficient direction. After completing all evaluation years, a multi-year summary index for the scenario-node configuration is generated for gating and sorting in step S6.
[0064] S5.3 Multi-year evaluation indicators and results table: Each year to be evaluated has only one actual area and one predicted value. Therefore, absolute or relative errors can be calculated for that year, but a meaningful cross-year coefficient of determination cannot be formed independently. Let the effective year for participating in the allocation evaluation be... Let be the actual area and the predicted area for the i-th evaluation year, respectively, and let Ā be the average of the n actual areas. Then, the allocation index is calculated using the following formula:
[0065]
[0066] (6) When the variance of the actual area is 0, R² is not calculated; instead, RMSE and MRE are used for evaluation. When the actual area of individual instances is 0, the MRE denominator for that year is not included, and the absolute error is recorded separately. The annual prediction error is defined as the predicted area minus the actual area. The cross-year error standard deviation is calculated from the error sequence of the same candidate configuration in each valid evaluation year. Therefore, R², RMSE, MRE, and error standard deviation are all multi-year configuration indicators and cannot be used to evaluate the target year T, which has not yet obtained a true label. Step S5 outputs a candidate configuration evaluation result table. The result table includes at least the candidate spatial scene, candidate prediction node, predicted value for each evaluation year, actual value for each evaluation year, annual error, multi-year summary indicator, error standard deviation, selection feature, and coefficient direction record, and serves as the input for S6 reliability gating and target configuration determination.
[0067] S5.4 Forward evaluation example Taking the candidate configuration "Source_A—January 1st Node" as an example: When evaluating 2021, only labeled data from 2008 to 2020 is used to complete weight determination, window selection, feature construction, and model fitting. Then, environmental data available before January 1, 2021 is used to predict 2021. After the true labels are formed, only the error of this round is recorded. When evaluating 2022, the training data is expanded to 2008 to 2021, and the same process is repeated within this training set. Then, 2023 is predicted using 2008 to 2022, 2024 is predicted using 2008 to 2023, and 2025 is predicted using 2008 to 2024. After five rounds, the predicted values, true values, and error sequences for 2021 to 2025 are formed, and the multi-year index of formula (6) is calculated accordingly. The remaining 35 candidate configurations are executed according to the same annual boundary, and finally, a configuration evaluation table of "spatial scene × prediction node" is formed.
[0068] S6. Determine the target spatial scene and target prediction nodes: S6.1 Input and Filtering Objects: S6 reads the candidate configuration evaluation result table and model feature records from each round generated in S5, performs reliability gating and ranking on the candidate configurations, and does not retrain the prediction model in this step. S5 answers "How well does each spatial scene-prediction node configuration perform?", and S6 answers "Which configurations meet the release conditions, and which scene and node are ultimately adopted?"
[0069] S6.2 Performance, error, and feature reproducibility gating: One embodiment of the threshold determination method is as follows: During the initial development period, forward feedback is performed, and cross-year R², RMSE, MRE, and error standard deviation are calculated for 36 candidate configurations consisting of 4×9=36 regions and months. The threshold is then determined based on the out-of-sample distribution of the corresponding indicators for all candidate configurations. The R² threshold is not lower than 0, and the RMSE, MRE, and error standard deviation thresholds are taken from the upper quartiles of the corresponding distributions. The threshold is registered and frozen after the initial configuration development is completed. Step S8 first uses the threshold to review the current configuration. Only when the current configuration is judged to have degraded according to the original threshold and a new baseline version is established is a new threshold version formed according to the same preset rules.
[0070] The selection frequency of a process feature is only counted in the effective evaluation rounds in which the feature can complete window calculation and actually enter the candidate feature set. A preset lower limit is required for the number of effective evaluation rounds to avoid artificially high frequencies caused by a few rounds. Directional consistency is calculated based on the number of rounds in which the feature coefficient is non-zero. When the proportion of rounds in either the positive or negative direction reaches a threshold, the direction is considered stable; if no non-zero coefficient ever appears, the reproducibility condition is not met. In this embodiment, the selection frequency is no less than 60% of the effective evaluation rounds, the directional consistency is no less than 80%, and the number of effective evaluation rounds is no less than 60% of the total evaluation rounds. When at least one process feature in the candidate configuration simultaneously meets the above conditions, it is determined to have passed the process feature reproducibility check.
[0071] In the single-channel implementation, the same process feature is identified by the combination of "original environmental variables and process or interaction type," and changes in the specific window value during the rolling round do not change its identity. The multi-scale implementation also includes short-term, medium-term, or long-term channel identifiers in the feature identity. Candidate configurations must simultaneously meet prediction performance thresholds such as R², RMSE, and MRE, the annual error series standard deviation threshold, and the process feature reproducibility threshold. If any necessary condition is not met, the configuration will not be included in the set of releasable configurations.
[0072] S6.3 Lead Gating: The minimum forecast lead time is calculated based on the number of calendar days between the candidate node date and May 1st, the start date of the current outbreak season. The example requires at least 30 days. Therefore, candidate nodes less than 30 days from May 1st can be retained for performance comparison but cannot be included in the official releasable configuration set. The above values are for illustrative purposes; in practice, adjustments can be made based on historical sample size and business tolerance.
[0073] S6.4 Configure sorting and freezing content: Configurations that meet all conditions form a set of publishable configurations. For valid nodes in the same or different spatial scenarios, R² is compared first; if the difference in R² does not reach the preset tolerance, RMSE is compared; if the difference in RMSE still does not reach the preset tolerance, MRE is compared; if all three differences are within the tolerance, the node with the earlier time is selected. This determines the target spatial scenario and the target prediction node. Step S6 records and freezes the scenario type and generation rules, node date, feature identity and construction rules, memory window search rules, and model type; the specific window values, standardized parameters, and regression coefficients obtained again in each forward round or business target year are recorded separately as the fitting results of that round.
[0074] This step, in this embodiment, sequentially performs performance gating, error standard deviation gating, feature reproducibility gating, and 30-day lead time gating on the 2021-2025 evaluation table of 36 configurations; only configurations that pass all tests are ranked according to the three-level indicators. If no configuration passes, the output is "No publishable configuration under the current data conditions," instead of forcibly selecting one from invalid configurations.
[0075] S7, Target Year Unlabeled Forecast: S7.1 Generate current round input according to target configuration: For the actual business target year T, let u = T. Step S7 directly uses the target spatial scene and target prediction node determined in step S6, without recalculating all candidate scenes. If the target scene is Source_A or Source_B, then the corresponding binary weights are generated according to the determined boundaries; if the target scene is Path or Bloom, then according to their established rules, only the values up to T are used. The historical annual occupancy raster for one year generates W(s,T,p). In the single-round model for the target year, all environmental features with the same name for both historical training years and the target year are constructed using W(s,T,p) and the same extraction rules. Unselected spatial scenes can only be recalculated and evaluated when the current configuration is determined to be degraded in step S8 and a structure reselection is triggered.
[0076] S7.2 Final Window, Model Refitting, and Prediction: Subsequently, without altering the target spatial scene, target prediction node, and feature identity determined in step S6, the window search in step S4 is re-executed using all labeled years with y < T to determine the final memory window for each environmental variable. Then, historical training features and target year features are constructed according to this window, and the standardized parameters, model parameters, and regression coefficients are recalculated using the model type defined in step S5. This final window and fitting result are used only for target year prediction and do not participate in the candidate configuration re-ranking in step S6. The true area of the target year must be empty, and environmental data after the target node must not be included in window search, missing data processing, or feature construction. If a final window or model that meets the data requirements cannot be formed, a formal prediction is not forced, and the most recent valid configuration record is retained for verification.
[0077] S7.3 Predicted Values and Configuration Certificates: After model refitting in step S7.2, the target annual area prediction value is generated based on the model output scale. When the model directly outputs the annual independent coverage area, the model output is used as... When the model outputs log-scaled predicted values according to step S5.1 At that time, the inverse exponential transformation and deviation correction are performed according to the following formula:
[0078] In the formula, For historical evaluation of the year out-of-sample logarithmic residuals; These are out-of-sample logarithmic scale predictions obtained according to the year-by-year forward rule; This represents the number of effective out-of-sample residuals. This is the deviation correction factor. The labels, forecasts, and residuals used to calculate this deviation correction factor should all be in the target year. It was formed previously.
[0079] See Figure 6 In outputting the target year's predicted area value Simultaneously, the corresponding configuration records are saved. The configuration records include at least the target year, historical data cutoff time, target spatial scene, target prediction node, actual lead time, memory window, feature name and order, standardized parameters, model parameters, model output scale, bias correction factor, and configuration validity status, which are used to reproduce the data boundaries and calculation process of this prediction.
[0080] Example of this step: Taking the prediction of 2027 as January 1, 2027 as an example, this round only uses historical labels up to 2026 and environmental data that can be obtained before January 1, 2027; when the model outputs the logarithmic area, the bias correction factor is calculated using the out-of-sample residuals formed before 2027, and obtained according to equation (7). .
[0081] S8. Post-prediction evaluation and controlled updates: S8.1 Truth Arrival and Error Archiving: See Figure 7 After the outbreak season from May 1st to August 31st of the target year ends and a true annual scale label is generated, this true label is compared with the predicted value output in step S7. The annual absolute error, relative error, and other preset evaluation indicators are calculated, and the predicted value, actual value, error, data cutoff boundary, and configuration record are archived together. This true label must not be used to change the already published target year prediction results; it can only be used as a data source for subsequent year modeling and maintenance.
[0082] S8.2 Current Configuration Health Reassessment: During the update, the target space scene type, scene generation rules, target prediction nodes, feature identities, and window search rules are kept unchanged. The new year is added to the labeled historical set, and the health re-evaluation is performed only for the current configuration according to the year-by-year forward rule in step S5. The example prioritizes the five most recent years with valid real labels and corresponding out-of-sample prediction records; if less than five years are available, all available years are used, and the actual number of years is recorded. This re-evaluation recalculates R², RMSE, and MRE as multi-year prediction performance, calculates the standard deviation of the annual prediction error sequence as error stability, and re-forms process feature reproducibility records based on the selection status and coefficient direction of each round. A single full-history fitting result cannot replace the year-by-year forward evaluation.
[0083] S8.3 Stable Updates and Degradation Reselection: If the current configuration continues to meet the prediction performance threshold, error stability threshold, process feature reproducibility threshold, and minimum lead requirement frozen in step S6 on the same evaluation year set, then the configuration structure is maintained, and the final window and model parameters are re-determined using only the updated labeled years as the next target years. Configuration degradation is determined only if any necessary condition is not met, and S5 and S6 are re-executed for all candidate spatial scenarios and candidate prediction nodes.
[0084] When a structural review is triggered, the current configuration and the candidate new configuration must be compared under the same evaluation year set, the same data cutoff boundary, and the same forward evaluation rules. The current configuration can only be replaced if the candidate new configuration achieves the preset performance improvement and the predicted node is no later than the original node. The original predicted values, configuration records, and version numbers are all retained, and the new configuration does not overwrite the historical records.
[0085] S8.4 Update Output: S8 outputs the current configuration health status, the actual set of re-evaluation years, the updated parameter version or new configuration version, the reason for replacement, and the availability status for the next target year. This step distinguishes between a single annual error record, the current configuration's multi-year health re-evaluation, and the reselection of all candidate configuration structures, avoiding immediate replacement of the prediction framework due to a single abnormal year.
[0086] This step, in this embodiment, uses the five most recent valid sequences to predict the annual review; when the configuration is stable, only the final window and model parameters for the next year are updated; a full structure reselection is only performed when the configuration degrades. Candidate new configurations must also meet a preset improvement threshold and the target node must not be later than the original node before they can be replaced.
[0087] III. Enhanced Prediction Implementation Methods: The following enhanced implementation methods all refer to the basic steps S1-S8 in Part Two, only adding verification, channel, range, or version control mechanisms after the corresponding steps. Please refer to [link to relevant documentation]. Figure 3 Multi-scale stable memory is connected to S4; forward testing of the frozen configuration structure is connected to S6 to determine the target configuration; empirical prediction intervals are connected to the prediction output end of S7; and version-controlled updates are connected to S8. Basic processing not explicitly added will not be described again.
[0088] 3.1 Implementation method for forward testing of frozen configuration structure: Please see Figure 3 When stricter configuration access is required, labeled historical data can be divided into a basic training period, a configuration filtering and freezing period, and a year-by-year forward testing period, before entering the business prediction period. The basic training period is used to establish labels, scenario algorithms, candidate features, and model frameworks; the configuration filtering and freezing period is used to compare candidate scenario-node configurations; before entering the year-by-year forward testing period, the spatial scenario types and generation rules, target prediction nodes, feature identities and construction rules, memory window search range and selection rules, and model type and parameter selection rules are frozen. During the testing period, the above configuration structure is not reselected based on test results, but each test year can recalculate spatial weights, final windows, standardized parameters, and model coefficients using only the real labels formed before that year, according to the pre-frozen rules.
[0089] For example, 2008–2014 can be used as the initial training period, 2015–2020 as the configuration selection and freezing period, 2021–2025 as the forward testing period for the frozen configuration, and 2026 onwards as the business forecasting period. During the configuration selection and freezing period, 2015–2020 are used sequentially as the evaluation years, and forward evaluation is performed on the configuration of each candidate spatial scenario-prediction node. The target configuration is determined based on multi-year prediction performance, error stability, process feature reproducibility, and prediction lead. Before entering the forward testing period, the target spatial scenario type and generation rules, target prediction nodes, feature construction rules, window search rules, and model type are frozen.
[0090] 3.2 Implementation of Multi-Scale Stable Memory: Based on the single-window memory features in step S4, separate short-term, medium-term, and long-term local optimal windows can be retained for the same original environmental variable in the lag-cumulative evaluation surface. Each channel always retains the corresponding identifier of "original variable - process category - time scale". For example, both the short-term and long-term MWP channels belong to wave persistence, but they are calculated as two independent candidate features with window values and training set standardized parameters respectively.
[0091] Channel merging follows the process rules outlined in step S4.3: within the same time scale, NO3, PO4, and S... i Channels are used to form nutrient provenance memory at the corresponding scale, while U10 and V10 channels are used to form wind transport at the corresponding scale. The remaining variables retain their individual process identities. Channels at different time scales are not directly combined into a single mean; instead, channels that pass the stability check are input into the complete sparse regression model in step S5 along with the process characteristics of other variables. Process categories are used to record ecological implications and summarize stability, without altering the characteristics of individual variables or channels.
[0092] The comparison between the multi-channel model and the single-channel model is conducted at the level of the complete prediction model, rather than using the R² or RMSE of a single variable as the comparison object. For each annual evaluation year, the annual relative errors of the two models are compared; multi-scale channels are only activated when the multi-channel error improves in a predetermined proportion of evaluation years, and the earliest reliable prediction node of the multi-channel model is no later than that of the single-channel model. The example requires an improvement in MRE in at least three of the five evaluation years.
[0093] 3.3 Implementation method for empirical prediction intervals: Step S7.3 outputs the area-scale point prediction values after inverse transformation and bias correction. Based on the point prediction values, this embodiment further outputs an empirical prediction interval to characterize the prediction uncertainty. Specifically, it resamples with replacement from the out-of-sample logarithmic residuals formed according to the year-by-year forward evaluation rules before the target year; after adding each extracted residual to the logarithmic-scale prediction value of the target year, it directly performs an exponential inverse transformation to obtain multiple simulated area prediction values. Since the extracted out-of-sample residuals already characterize the logarithmic scale error and its inverse transformation effect, the simulated area prediction values are no longer multiplied by the bias correction factor in step S7.3. The lower and upper limits of the simulated area prediction intervals are determined according to preset quantiles. For example, the 2.5% quantile and the 97.5% quantile are used to form the 95% empirical prediction interval. This embodiment ultimately outputs the area point prediction values, the lower limit of the empirical prediction interval, and the upper limit of the empirical prediction interval simultaneously, enabling business personnel not only to obtain the predicted area for the target year but also to determine the possible fluctuation range of the prediction results. The residuals, predicted values, and true labels used for resampling must all be generated before the target year; information from the actual area of the target year or after the target prediction node must not be used. The intervals described are empirical prediction intervals generated under limited historical sample conditions and are not subject to statistical coverage guarantees exceeding the sample size and residual representativeness conditions.
[0094] 3.4 Versioned Controlled Update Implementation Method: This implementation further defines the configuration version management rules based on step S8. After the outbreak season from May 1st to August 31st of the target year ends and real labels are formed, the prediction error is archived. First, keeping the spatial scene type, scene generation rules, target prediction nodes, feature identities, and window search rules unchanged, the new year is added to the labeled set, and a forward health review is performed only for the current configuration. If the current configuration still meets the original threshold version in the unified evaluation year set, the original configuration structure continues to be used, and the final window and model parameters are redefined for the next target year.
[0095] Steps S5 and S6 are only re-executed for all candidate configurations if the current configuration, based on the original threshold version, no longer meets the necessary conditions for prediction performance, error stability, process characteristic reproducibility, or minimum lead time. The current configuration and candidate new configurations must be compared under the same evaluation year set, the same data cutoff boundary, and the same forward evaluation rules; a candidate new configuration is only replaced if it achieves a preset performance improvement and the prediction node is no later than the original node. A new threshold version can only be established according to pre-defined rules after the new configuration is determined; the threshold cannot be changed first and then the old configuration's degradation cannot be determined. The original prediction values, configuration records, and version numbers are all retained; the new configuration does not overwrite historical records.
[0096] IV. System and Business Process Implementation Methods: The system implementation description illustrates how the above method can be executed modularly by a computing system with a processor, memory, and data interface, without altering the algorithm order of S1-S8. Please refer to [link to documentation]. Figure 4 .
[0097] An annual green tide prediction system integrating spatial scenes and prediction nodes includes a label construction module, a scene weight construction module, an environmental sequence extraction module, a memory and process feature module, a configuration joint evaluation module, a target configuration determination module, an annual prediction module, and an update evaluation module.
[0098] The complete business process is divided into four stages: the data preparation stage (S1-S3) completes the construction of historical labels, spatial weights, and scene environment sequences; the model building stage (S4-S6) completes memory and process features, forward evaluation of candidate configurations year by year, and determination of target scenarios and nodes; the business forecasting stage (S7) outputs the annual scale forecast value and configuration certificate under the condition that there are no real labels in the target year; and the model validation and update stage (S8) completes error archiving, current configuration health review, and controlled updates after the arrival of real labels. Inputs include daily green tide detection results, candidate source boundaries, historical annual occupied grids, and marine and meteorological environmental data; outputs include annual forecast values, target configurations, data cutoff records, error archiving, and update status.
[0099] The system can also be configured with forward testing of the frozen configuration structure, multi-scale memory, experience prediction intervals, and version management modules, which are connected to these modules respectively. Figure 3 The basic steps are shown. Figure 5 This is a partial unfolded diagram of S4-S7 in the basic prediction process, illustrating the relationship between single-round unified weights, window selection within the training set, year-by-year forward evaluation, and the final window for the target year. It does not belong to the process that is only executed in the enhancement scheme. Figure 6 Explain the target annual information boundaries for S7; Figure 7 Explain the relationship between the evaluation and controlled update after the truth value of S8 is reached.
[0100] The above description is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for predicting annual green tides by integrating spatial scenes and prediction nodes, characterized in that, include: S1. Obtain the daily green tide detection results of the outbreak season in historical years, spatially combine the green tide detection grids of the valid observation days in the same year, calculate the annual independent coverage area, and form a historical annual scale label. S2. Based on different spatial information or weight generation rules that distinguish them from each other, construct at least two candidate spatial scene weights; S3. Acquire environmental data, perform spatial weighting according to the weight of each candidate spatial scenario, and form the environmental time series corresponding to each candidate spatial scenario. S4. Set up candidate prediction nodes, and construct ecological memory features and process features based on the environmental time series within the information boundary of each candidate prediction node. S5. Combine candidate spatial scenarios and candidate prediction nodes into a candidate configuration; for any year to be evaluated, use only the scale labels before that year and the environmental data up to the candidate node to fit a prediction model that can output the state of process characteristics and obtain the prediction value. By summarizing the prediction results and process feature status of multiple years to be evaluated, the prediction performance, error stability, feature selection and action direction of each candidate configuration are recorded. S6. Based on the prediction performance, error stability, process feature reproducibility determined by process feature selection and action direction recording, and minimum prediction lead, candidate configurations are screened to determine the target spatial scene and target prediction node. S7. For target years for which the true scale label has not yet been obtained, construct spatial weights according to the target spatial scenario rules, use historical scale labels and environmental data up to the previous year, redetermine the target year memory window and fit the prediction model, and output the target year green tide scale prediction value using only the target year environmental data up to the target prediction node.
2. The annual green tide prediction method integrating spatial scenes and prediction nodes according to claim 1, characterized in that, The candidate space scene weights include at least two of the following categories: The seed source category scene weight is obtained by rasterizing the candidate seed source boundary. The weight of each effective grid within the candidate range is 1, and the weight outside the candidate range is 0. The weights for transportation scenarios are obtained by accumulating the annual occurrence counts of each grid cell based on the green tide annual grid cells in previous historical years before the year to be evaluated or the target year, and then normalizing and spatially smoothing them. The weights for display-type scenes are obtained by performing high-frequency filtering and connected patch identification based on the annual occurrence frequency, and by assigning weights to the retained patches according to the annual occurrence frequency. Among them, the spatial ranges of different spatial scenarios are allowed to overlap; the seed source scenario is generated based on the candidate boundary, and the transport and manifestation scenarios are allowed to use the historical annual occupancy information, but they adopt different weight conversion rules and are not formed by mechanically cutting, expanding or copying the seed source boundary.
3. The annual green tide prediction method integrating spatial scenes and prediction nodes according to claim 1, characterized in that, The annual independent coverage area is the union area of the grid spaces detected as green tides on each valid observation day of the outbreak season in the same year. The same grid is only counted once, no matter how many times it is detected in the same year. The annual occupancy status is determined only when the grid meets the preset annual valid observation conditions. Grids that do not meet the preset annual valid observation conditions are marked as missing measurements and are not considered as non-occurring grids.
4. The annual green tide prediction method integrating spatial scenes and prediction nodes according to claim 1, characterized in that, The ecological memory features are obtained by setting lag length and cumulative length for environmental time series prior to candidate prediction nodes; the process features include at least one of nutrient supply, light and heat conditions, wind transport, wave persistence and disturbance forcing, and their interaction features.
5. The annual green tide prediction method integrating spatial scenes and prediction nodes according to claim 1, characterized in that, The prediction performance includes at least two of the following: coefficient of determination, root mean square error, and mean absolute relative error; the process feature reproducibility is determined based on the selection frequency and the consistency of the direction of action of the process features in several effective year-by-year forward evaluation rounds; the candidate configuration is marked as a valid configuration only when the prediction performance, error stability, process feature reproducibility, and minimum prediction lead all meet the preset conditions.
6. The annual green tide prediction method based on the fusion of spatial scenes and prediction nodes according to claim 1, characterized in that, The prediction model is a sparse regression model that can output the state of feature action; when performing a logarithmic transformation on the annual independent coverage area, the inverse transformation correction factor is determined based on the out-of-sample residuals before the target year to obtain the area-scale prediction value.
7. The annual green tide prediction method integrating spatial scenes and prediction nodes according to claim 1, characterized in that, In the forward testing implementation of the frozen configuration structure, the labeled historical data is divided into the basic training period, the configuration filtering and freezing period, and the forward testing period. Candidate configurations are compared only during the configuration screening and freezing period. After entering the forward testing period, the spatial scene type and generation rules, target prediction nodes, feature identities and construction rules, memory window search rules and model types are frozen. The configuration structure is not reselected based on the test year results. The standardized parameters, memory window and model parameters for each test year are only re-determined using the real labels formed before that test year.
8. The annual green tide prediction method integrating spatial scenes and prediction nodes according to claim 1, characterized in that, Also includes: S8. After the actual scale label for the target year is formed, calculate and archive the forecast error for that year; The current configuration is re-evaluated according to the same forward rules as the historical evaluation. If the current configuration still meets the preset conditions, the configuration structure is maintained and new labels are added to refit the model parameters. If the current configuration degenerates, the candidate configurations are re-evaluated, and the target configuration is re-determined only if the new candidate configuration meets the preset improvement conditions.
9. The annual green tide prediction method based on the fusion of spatial scenes and prediction nodes according to claim 8, characterized in that, After obtaining the target year's true scale label, keep the target spatial scene type, target prediction node, and feature construction rules unchanged, and re-evaluate the health of the current configuration according to the historical forward evaluation rules. If the current configuration has not degraded, add new labels to refit the model parameters. Only when the current configuration degrades will the candidate configuration evaluation be re-executed, and the current configuration will be replaced only if the new candidate configuration meets the preset improvement conditions relative to the current configuration.
10. An annual green tide prediction system integrating spatial scenes and prediction nodes, characterized in that, include: The tag building module is used to generate historical year-scale tags based on the daily green tide detection results of historical years; The scene weight construction module is used to construct candidate spatial scene weights based on different spatial information or weight generation rules that distinguish them from each other. The environmental sequence extraction module is used to spatially weight environmental data according to the candidate spatial scene weights to form an environmental time series. The memory and process feature module is used to construct ecological memory features and process features within the information boundaries of candidate prediction nodes; Configure a joint evaluation module to combine candidate spatial scenes and candidate prediction nodes into candidate configurations and perform year-by-year forward evaluation; The target configuration determination module is used to determine the target spatial scene and target prediction nodes based on prediction performance, error stability, process feature reproducibility and minimum prediction lead. The annual forecast module is used to output the target annual green tide scale forecast value using only the target annual environmental data before the target forecast node.
11. The annual green tide prediction system integrating spatial scenes and prediction nodes according to claim 10, characterized in that, It also includes an update evaluation module, which archives prediction errors, reviews the current configuration, and updates the target configuration when the current configuration degrades and the candidate new configuration meets the preset improvement conditions after the target year's true scale labels are formed.