Meteorological metadata storage method and system based on machine learning

Through the machine learning-based method, a dynamic hierarchical index structure and adaptive storage strategy are constructed, which solves the problem that meteorological data storage systems are difficult to dynamically adapt to the spatial and temporal evolution characteristics in the existing technology, and realizes efficient data retrieval and optimized storage resource utilization.

CN120104579AActive Publication Date: 2025-06-06HUAYUNSHENGDA(BEIJING)METEROLOGICAL TECH CO LTD

Patent Information

Application Number
CN202510583997.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

Existing meteorological data storage systems are difficult to dynamically adapt to the spatio-temporal evolution characteristics of meteorological data, resulting in the static index structure being unable to effectively adjust storage priorities when data access hotspots change, resulting in increased search delays and waste of storage resources.

Method used

Using a machine learning-based method, the model and spatial topology mapping model are extracted through time-series feature, the timing distribution features and spatially associated topology features are generated, the dynamic hierarchical index structure is constructed, and the meteorological storage optimization model is used to generate adaptive storage strategies.

Benefits of technology

It realizes dynamic adaptation of space-time correlation, improves data retrieval efficiency and storage resource utilization, reduces operation and maintenance costs, and improves disaster recovery and recovery capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104579A_ABST
    Figure CN120104579A_ABST
Patent Text Reader

Abstract

The invention provides a meteorological metadata storage method and system based on machine learning, and the method comprises the steps: obtaining a meteorological metadata set of a target region, inputting the meteorological metadata set into a pre-trained time sequence feature extraction model, generating the time sequence distribution features of all meteorological file files, and inputting a spatial topology mapping model, generating spatial correlation topological features of the meteorological elements; based on a cross fusion result of the time sequence distribution features and the space correlation topological features, a hierarchical index structure of meteorological metadata is constructed, a meteorological storage optimization model is called to generate a storage node distribution strategy of a meteorological file, and the meteorological file is stored according to the storage node distribution strategy. And writing each meteorological file file in the meteorological metadata set into the corresponding distributed storage node, and dynamically adjusting the mapping relationship between the physical position and the index of the stored file based on a real-time updating mechanism of the hierarchical index. The meteorological metadata storage efficiency and retrieval accuracy can be improved, and storage resource waste and operation and maintenance cost are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a meteorological metadata storage method and storage system based on machine learning. Background Art

[0002] In the current meteorological data storage system, meteorological monitoring agencies manage massive unstructured meteorological file files through distributed storage nodes, and generally adopt indexing technology based on spatiotemporal dimensions to improve data retrieval efficiency; in related technologies, fixed rules are usually used to divide the storage hierarchy structure, such as presetting vertical classification indexes by meteorological element types or dividing static spatial indexes by geographic grids. Although such methods can achieve basic data organization, they are difficult to dynamically adapt to the spatiotemporal evolution characteristics of meteorological data: on the one hand, the periodic fluctuations of meteorological elements and sudden events cause data access hotspots to change dramatically over time, and the static index structure cannot dynamically adjust the storage priority according to the time series characteristics, causing high-frequency access data to be retained in low-performance storage nodes, and the retrieval delay is significantly increased; on the other hand, the spatial correlation between meteorological sites and the cross-regional data dependency have not been effectively modeled, resulting in a mismatch between the spatial index density distribution and actual business needs, serious waste of storage resources and insufficient disaster recovery capabilities. Summary of the invention

[0003] In view of this, an embodiment of the present invention provides a meteorological metadata storage method and storage system based on machine learning. The technical solution of the embodiment of the present invention is implemented as follows:

[0004] On the one hand, an embodiment of the present invention provides a meteorological metadata storage method based on machine learning, the method comprising: obtaining a meteorological metadata data set of a target area, the meteorological metadata data set comprising a plurality of unstructured meteorological file files, each meteorological file file comprising monitoring data of at least one meteorological element and corresponding spatiotemporal identification information; inputting the meteorological metadata data set into a pre-trained temporal feature extraction model to generate temporal distribution features of each meteorological file file, and synchronously inputting the meteorological metadata data set into a spatial topological mapping model to generate spatial correlation topological features of each meteorological element; constructing a hierarchical index structure of meteorological metadata based on the cross-fusion result of the temporal distribution features and the spatial correlation topological features, the hierarchical index structure comprising a vertical hierarchy divided by meteorological element type and a horizontal hierarchy divided by spatiotemporal dimensions; calling a meteorological storage optimization model, and generating a storage node allocation strategy for meteorological file files according to the hierarchical weight distribution of the hierarchical index structure, the storage node allocation strategy comprising a compression rate threshold and a storage location priority of each meteorological element data; writing each meteorological file file in the meteorological metadata set into a corresponding distributed storage node according to the storage node allocation strategy, and dynamically adjusting the physical location and index mapping relationship of the stored files based on the real-time update mechanism of the hierarchical index.

[0005] On the other hand, an embodiment of the present invention provides a storage system, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps in the above method when executing the program.

[0006] The meteorological metadata storage method based on machine learning provided by the present invention constructs a dynamic hierarchical index structure by cross-analysis of integrating temporal distribution features and spatial correlation topological features, and generates an adaptive storage strategy using a meteorological storage optimization model, thereby solving the core problems of spatiotemporal correlation separation and rigid storage resource allocation in meteorological data storage: the temporal distribution features provide dynamic weight factors for the index hierarchy by capturing the periodic fluctuations and mutation events of meteorological elements, so that high-frequency access data is preferentially allocated to high-performance storage nodes, while the spatial correlation topological features generate a spatial density grid by characterizing the spatial conduction relationship between meteorological sites to optimize the geographical distribution density of data. The cross-fusion mechanism of the two converts the nonlinear correlation of spatiotemporal features into the hierarchical weight distribution of the hierarchical index, thereby realizing the dynamic adaptation of the storage structure to the meteorological business scenario; at the same time, the hierarchical index structure enables meteorological file files to be quickly located based on composite index tags through bidirectional mapping between the vertical hierarchy (meteorological element type) and the horizontal hierarchy (space-time grid), and the meteorological storage optimization model optimizes the location of meteorological files based on the hierarchical weight distribution and the node The compression rate threshold, redundant copy strategy and storage location priority queue are generated for point load fluctuation data to ensure that high-frequency data is stored preferentially in high-bandwidth nodes and low-frequency data is stored at a high compression rate. The dynamic adjustment mechanism avoids storage node overload and ensures the retrievability of historical data by real-time migration of low-temperature data and updating index mapping relationships. In addition, in view of the strong timeliness of meteorological data, the storage priority of expired data is automatically reduced through the dynamic weight factor decay mechanism to release storage space. Combined with the core storage area division of the spatial density grid, the storage redundancy of data in disaster-prone areas is prioritized to improve disaster recovery capabilities. At the same time, the model-driven adaptive capability dynamically adjusts the compression algorithm and storage location allocation strategy to adapt to the differentiated needs of meteorological data centers of different sizes. Finally, through the technical collaboration of spatiotemporal feature fusion, dynamic index construction, model strategy generation and closed-loop feedback mechanism, the storage efficiency and retrieval accuracy of meteorological metadata are significantly improved, the waste of storage resources and operation and maintenance costs are reduced, and high-reliability and high-adaptability data management support is provided for meteorological services. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 A schematic diagram of the implementation flow of a meteorological metadata storage method based on machine learning provided in an embodiment of the present invention.

[0008] Figure 2 A hardware entity diagram of a storage system provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0009] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present invention. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present invention.

[0010] The embodiment of the present invention provides a meteorological metadata storage method based on machine learning, which can be executed by a processor of a storage system, wherein the storage system can refer to a device with data processing capability, such as a server or a desktop computer.

[0011] Figure 1 A schematic diagram of the implementation flow of a meteorological metadata storage method based on machine learning provided in an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method comprises the following steps: Step S100: Acquire a meteorological metadata set of a target area, where the meteorological metadata set includes a plurality of unstructured meteorological file files, each of which contains monitoring data of at least one meteorological element and corresponding spatiotemporal identification information.

[0012] The process of obtaining the meteorological metadata set of the target area is to collect a collection of raw meteorological data from a specific geographical range, where the meteorological metadata set consists of multiple unstructured meteorological file files. Unstructured meteorological file files are, for example, raw data files that are not organized according to a unified format or standard, and their content may contain text, numerical values, charts, and other forms. Each meteorological file contains at least one meteorological element monitoring data. Meteorological elements are, for example, basic physical quantities that describe the state of the atmosphere, such as temperature, humidity, wind speed, precipitation, etc. Monitoring data refers to real-time or historical measurements collected by sensors, satellites, or ground observation equipment. Spatiotemporal identification information is, for example, time and space tags associated with meteorological element monitoring data. The time tag includes the specific time point or time period of data collection (for example, October 1, 2023, 12:00 to 18:00), and the space tag includes the geographic location coordinates of data collection (for example, longitude 116.40°E, latitude 39.90°N) or regional code (for example, city B meteorological zone code A01). Specifically, the acquisition of meteorological metadata sets can be achieved through the interface of data sources such as meteorological observation stations, remote sensing satellites or meteorological radars, where unstructured meteorological file files are stored in various file formats such as CSV, JSON, NetCDF, etc. For example, a meteorological file may contain temperature monitoring data of City B in the summer of 2023, and its spatiotemporal identification information is marked as the time interval "June 1 to August 31, 2023" and the geographical scope "the entire area of ​​City B". In this process, the meteorological metadata set needs to ensure the integrity and originality of the data so that the spatiotemporal features can be accurately extracted in the subsequent processing stage.

[0013] As an implementation manner, in step S100, after obtaining the meteorological metadata set of the target area, a process of performing data preprocessing on the meteorological metadata set may also be included, wherein the data preprocessing process may include: Step S110: identifying the time stamp interval of the missing monitoring data in the meteorological file, and generating interpolation data based on the fluctuation trend of the monitoring data of the adjacent meteorological stations; Identifying the timestamp interval of missing monitoring data in the meteorological file is, for example, by analyzing the time series of monitoring data recorded in the meteorological file, locating the continuous time period in which no valid data was collected. The timestamp interval is defined by a starting time point and an ending time point. For example, a certain meteorological file does not record any humidity monitoring data in the time period from 14:00 to 16:00 on July 10, 2023. Adjacent meteorological stations are, for example, other stations in the target area that are geographically adjacent to the current meteorological station and whose meteorological element monitoring data have spatial continuity. For example, the Chaoyang District Meteorological Station and the Haidian District Meteorological Station in City B are adjacent to each other. The fluctuation trend of monitoring data is, for example, the change law of meteorological element data of adjacent stations in the same time period. For example, the temperature data of the Chaoyang District Meteorological Station from 14:00 to 16:00 on July 10, 2023 showed a trend of rising by 1°C per hour, while the temperature data of the Haidian District Meteorological Station during the same period showed a trend of rising by 0.8°C per hour. The process of generating interpolation data based on this fluctuation trend is, for example, to infer the meteorological element values ​​of the current station in the missing timestamp interval through linear regression, time series prediction or spatial kriging interpolation algorithm. For example, if the Chaoyang District meteorological station is missing temperature data from 14:00 to 16:00, and the data in Haidian District during the same period is linearly increasing, the interpolation data can generate missing values ​​based on the historical correlation between the two stations (such as the temperature change slope ratio of 1:0.8). The generation of interpolation data must ensure spatial continuity constraints with adjacent stations to avoid the interpolation results deviating from the actual observation range due to local terrain or microclimate differences.

[0014] Step S120: Detect abnormal jump points in the meteorological element monitoring data, and use a sliding window algorithm to smooth and correct the data sequence in the window; Detecting abnormal jump points in meteorological element monitoring data is, for example, to identify sudden changes in data sequences that do not conform to meteorological laws or the normal measurement range of equipment through statistical analysis methods. Abnormal jump points are manifested as drastic fluctuations in the numerical values ​​of meteorological elements in a short period of time, such as wind speed suddenly increasing from 3m / s to 20m / s in 10 minutes, or precipitation suddenly increasing from 0mm to 50mm in 5 consecutive time points (collected once every 10 minutes). The sliding window algorithm is, for example, to divide the data sequence into continuous subsequences (windows) of fixed length and perform statistical analysis on the data in each window. For example, for the temperature data sequence, the window length is set to 24 hours, the window sliding step is 1 hour each time, and the standard deviation and mean of the temperature are calculated in each window. If the deviation of a data point from the mean exceeds 3 times the standard deviation, it is determined to be an abnormal jump point. Smoothing correction is, for example, to replace or adjust the value of the abnormal jump point through a filtering algorithm (such as moving average, median filtering or exponential smoothing) to make it conform to the natural variation law of meteorological elements. For example, for the above-mentioned wind speed anomaly, after detecting the jump point, the sliding window algorithm can replace the anomaly with the average value of the data before and after 1 hour in the window, thereby eliminating the noise caused by equipment failure or instantaneous interference.

[0015] Step S130: Convert meteorological file files of different formats into a unified binary encoding format, and add a metadata tag containing spatiotemporal identification information to the file header; wherein the preprocessed meteorological metadata set meets the input format specifications of the temporal feature extraction model and the spatial topology mapping model.

[0016] Meteorological archive files of different formats are, for example, raw files from heterogeneous data sources, and their formats may include text formats (such as CSV, TXT) and scientific data formats (such as NetCDF, HDF5). A unified binary encoding format is, for example, converting all files into binary streams arranged in fixed byte lengths, such as using the encoding rules defined by Protocol Buffers or Apache Avro, in which the value type (such as floating point number, integer), byte order and data length of each field are strictly standardized. The file header is, for example, a specific byte area at the beginning of a binary file, which is used to store metadata tags describing the file content. The metadata tags of spatiotemporal identification information include the geographical range of data collection (such as longitude 116.20°E to 116.60°E, latitude 39.80°N to 40.20°N), time interval (such as 00:00 on August 1, 2023 to 23:59 on August 31, 2023) and meteorological element type code (such as temperature code TEMP_01, humidity code HUMI_02). For example, after conversion, the first 128 bytes of a binary-coded header of a meteorological file in the original CSV format record the metadata tag "City B Entire Domain - August 2023 - Temperature Data", and the subsequent bytes store hourly temperature values ​​as floating point numbers. The preprocessed meteorological metadata set must meet the input format specifications of the time series feature extraction model and the spatial topology mapping model, including data dimension alignment (such as timestamps sorted by UTC time), numerical range normalization (such as uniform conversion of temperatures to degrees Celsius), and missing value placeholders (such as NaN filling). Format unification ensures that subsequent models can parse data efficiently and avoid parsing errors or feature extraction biases caused by format differences.

[0017] Step S200: input the meteorological metadata data set into the pre-trained temporal feature extraction model to generate the temporal distribution features of each meteorological file, and synchronously input the meteorological metadata data set into the spatial topology mapping model to generate the spatial correlation topology features of each meteorological element.

[0018] The pre-trained time series feature extraction model is a machine learning model trained with historical meteorological data, such as a recurrent neural network, which is used to extract representative patterns or laws from time series data. When the meteorological metadata dataset is input into the model, the model first parses the monitoring data and corresponding timestamps in each meteorological file to identify the changing trend of meteorological elements in the time dimension. The time series distribution feature is, for example, a vector or matrix output by the model that represents the temporal evolution characteristics of meteorological elements, such as periodic fluctuations, mutation events, or long-term trends. For example, for temperature monitoring data, the time series distribution feature may include the amplitude of daily temperature changes, the seasonal temperature rise rate, or the duration of extreme high temperature events. The spatial topological mapping model that synchronously inputs the meteorological metadata dataset is, for example, another independently trained model, which is used to analyze the spatial correlation of meteorological elements. The spatial correlation topological feature is, for example, a structure generated by the model that represents the mutual influence relationship between meteorological elements in different geographical locations, such as the data similarity of adjacent meteorological stations, the wind field propagation path, or the movement trajectory of precipitation clouds. For example, for wind speed data, the spatial correlation topological feature may be expressed as a network structure of wind direction conduction, in which the nodes represent meteorological stations and the edge weights represent the correlation coefficient of wind speed changes between stations. In this process, the temporal feature extraction model and the spatial topology mapping model need to run in parallel to ensure the independence and complementarity of temporal and spatial features, thereby providing a multi-dimensional analysis basis for subsequent fusion.

[0019] As an implementation method, the training process of the above-mentioned time series feature extraction model may include: Step S201: Collect continuous time series monitoring data of different meteorological elements in the historical meteorological data set as training samples, where the time series monitoring data includes the numerical fluctuation sequence of the meteorological elements at a preset time granularity and the corresponding timestamp mark.

[0020] The process of collecting historical meteorological data sets is, for example, to screen the continuous time series data of multiple meteorological elements in the target area from long-term stored meteorological archives, where different meteorological elements include independent physical quantities such as temperature, humidity, wind speed, and precipitation. Continuous time series monitoring data requires that the data points are uninterrupted in the time dimension and arranged at a fixed time granularity, such as once every hour, day, or month. The preset time granularity is, for example, the minimum time unit predefined in the model training phase. For example, when the granularity is hours, the time series monitoring data consists of 24 data points per day, and each data point corresponds to the measurement value at the hour. The numerical fluctuation sequence is, for example, the numerical change curve of meteorological elements in time granularity, such as the daily average temperature sequence of a certain area from January 1, 2020 to December 31, 2023, which contains a total of 1461 data points. The timestamp is marked with the precise time information associated with each data point, such as "2023-07-15 14:00:00 UTC". The construction of training samples needs to cover data from different seasons, climate events (such as El Niño) and geographical regions to ensure that the model can learn the multi-scale variation patterns of meteorological elements. For example, the temperature training sample can include the temperature monitoring data of city B in the past ten years with hourly granularity, with timestamps covering all valid collection moments, and the numerical fluctuation sequence reflects the fluctuation characteristics of temperature at hourly, daily, monthly and interannual scales.

[0021] Step S202: Input the time series monitoring data into the forward propagation path of the bidirectional recurrent neural network to extract the long-term trend characteristics of the meteorological elements. The long-term trend characteristics represent the periodic change law of the meteorological elements on the seasonal scale or the interannual scale.

[0022] The forward propagation path of a bidirectional recurrent neural network is, for example, the computational path of the model processing input data in chronological order (from the past to the future). After the time series monitoring data is arranged in the order of timestamp marks, it is input into the recurrent unit (such as LSTM or GRU) of the forward propagation path to update the hidden state step by time. Long-term trend features refer to the macroscopic change patterns extracted by the model from the input sequence, such as the slow rising trend of temperature on the interannual scale (such as the increase in average annual temperature caused by global warming) or the periodic fluctuations on the seasonal scale (such as high temperatures in summer and low temperatures in winter). For example, for the ten-year temperature data of city B, the forward propagation path can capture the long-term trend of the average temperature in July each year rising by 0.1℃ compared with the previous year, and the pattern of periodic peak temperatures from June to August each year. The extraction of long-term trend features relies on the memory capacity of the recurrent neural network, which ignores short-term noise and strengthens the correlation across years or seasons by retaining the cumulative effect of historical information.

[0023] Step S203: Input the time series monitoring data into the back propagation path of the bidirectional recurrent neural network to extract the short-term fluctuation characteristics of the meteorological elements. The short-term fluctuation characteristics represent the correlation of sudden changes in meteorological elements at the hourly or daily scale.

[0024] The back propagation path of a bidirectional recurrent neural network is, for example, the computational path of the model processing input data in reverse chronological order (from the future to the past). After the time series monitoring data is reversely input into the recurrent unit, the model iteratively calculates the hidden state from the end of the sequence to the beginning. Short-term fluctuation characteristics refer to the rapid change pattern of meteorological elements captured by the model within a fine-grained time window, such as a sudden drop in temperature within a few hours (such as the passage of a cold front) or a sudden increase in wind speed on a daily scale (such as an approaching typhoon). For example, for the temperature data of the same city B, the back propagation path can identify a short-term event in which the temperature suddenly dropped by 5°C from 14:00 to 16:00 on a certain day in July 2023, and analyze the correlation between the event and the change in air pressure on that day. The extraction of short-term fluctuation characteristics focuses on data mutations in local time windows, and the sensitivity of the reverse path is used to enhance the model's ability to respond to recent events, thereby supplementing the long-term trend analysis of the forward path.

[0025] Step S204: Perform multi-scale time window fusion on the long-term trend characteristics and the short-term fluctuation characteristics to generate a fused time series feature vector, wherein the first dimension of the fused time series feature vector corresponds to the intensity coefficient of the periodic variation law, and the second dimension corresponds to the confidence score of the correlation of the mutation event.

[0026] Multi-scale time window fusion, for example, integrates long-term trend features (such as interannual trends) with short-term fluctuation features (such as hourly mutations) through feature concatenation, weighted summation, or convolution operations. The fused time series feature vector is a numerical vector of fixed dimensions, whose first dimension quantifies the intensity of the periodicity of meteorological elements, such as the annual cycle amplitude coefficient extracted by Fourier transform; the second dimension evaluates the statistical significance of mutation events, such as the p-value calculated based on hypothesis testing or the probability score predicted by the model. For example, after fusing the long-term trend and short-term fluctuations of the temperature data of city B, the feature vector may be expressed as [0.85, 0.93], where 0.85 represents the annual periodicity intensity (1.0 is the strongest), and 0.93 represents the confidence level of the association between the sudden drop in temperature on that day and similar historical events. The fusion process needs to retain the independence of multi-scale features to avoid information confusion, and at the same time ensure the consistency of the numerical range of different dimensions through normalization.

[0027] Step S205: Introduce an adaptive attention allocation mechanism in the output layer of the neural network, and dynamically adjust the feature weight distribution of meteorological elements in different time windows according to the product relationship between the intensity coefficient of the fused time series feature vector and the confidence score.

[0028] The adaptive attention allocation mechanism, for example, automatically allocates the contribution of features in different time windows to the final output through learnable weight parameters. The multiplication relationship between the intensity coefficient and the confidence score generates the attention weight. For example, the product of the intensity coefficient of 0.85 and the confidence score of 0.93 is 0.7905, indicating that the feature weight of this time window is high. The dynamic adjustment process allocates weights according to the product results. For example, in windows with strong interannual periodicity and high confidence in mutation events, the weight is increased to 0.9; while in windows with weak periodicity or low confidence in mutation, the weight is reduced to 0.3. For example, for the summer data of City B in 2023, the model may give a high weight to the temperature drop event on July 15 (due to high-intensity periodicity and high-confidence mutations), while giving a low weight to the autumn stable period data. The attention mechanism adjusts the numerical distribution of the feature vector through scaling and translation operations, so that the model pays more attention to key time windows in prediction or classification tasks.

[0029] Step S206: Optimize network parameters through supervised learning algorithm so that the time series distribution characteristics output by the time series feature extraction model can simultaneously reflect the periodic stability and mutation sensitivity of meteorological elements, and use the optimized time series distribution characteristics as the basis for generating dynamic weight factors of the main index nodes of the vertical level in the hierarchical index structure.

[0030] The supervised learning algorithm uses labeled training data, measures the difference between the model output and the true value through a loss function (such as mean square error or cross entropy), and adjusts the network parameters through back propagation. Cyclic stability requires that the long-term trend characteristics of the model output are consistent with the cyclical law of historical data, for example, the prediction error of annual temperature fluctuations is less than 1%; mutation sensitivity requires that the model's detection accuracy for short-term events exceeds 95%. The optimized time series distribution feature is a vector that integrates periodic and mutation information. For example, [0.85, 0.93] becomes [0.88, 0.95] after parameter adjustment, which enhances the ability to represent actual meteorological changes. In the hierarchical index structure, the dynamic weight factor of the vertical level main index node (such as the temperature element node) is generated by the weighted dimension of the time series distribution feature, for example, 0.88×0.6 (periodic weight coefficient) + 0.95×0.4 (mutation weight coefficient) = 0.904, so that elements with high cyclic stability and high mutation sensitivity receive higher storage and retrieval priority. The real-time update mechanism of dynamic weight factors ensures that the index structure can adapt to changes in meteorological data. For example, the weight factor of wind speed elements in typhoon season increases significantly, triggering the priority adjustment of storage nodes.

[0031] Step S300: Based on the cross-fusion results of temporal distribution features and spatial correlation topological features, a hierarchical index structure of meteorological metadata is constructed, wherein the hierarchical index structure includes vertical levels divided by meteorological element types and horizontal levels divided by time and space dimensions.

[0032] The cross-fusion of temporal distribution features and spatial correlation topological features is to combine the features of time and space dimensions through mathematical operations or model integration, such as feature splicing, weighted superposition or attention mechanism. The goal of building a hierarchical index structure is to establish a multi-dimensional retrieval and storage framework for meteorological metadata. The vertical level is divided according to the type of meteorological elements, such as temperature, humidity, wind speed and other categories. Each category is an independent main index node, and the node stores the temporal distribution features of the element and the associated spatial topological features. The horizontal level is divided according to the time and space dimensions. The time dimension can be subdivided into granularities such as hours, days, and months, and the spatial dimension can be based on geographic grids (such as 1km×1km units) or administrative divisions (such as provinces, cities, and counties). For example, in the vertical level, the main index node of the temperature element contains its periodic change characteristics and the conductive relationship with the surrounding areas; in the horizontal level, the temperature data of City B in June 2023 may be classified into the "North China-Summer-Core Grid" unit. The cross-fusion process needs to ensure the dynamic association between the vertical and horizontal levels. For example, when a meteorological element has a spatial anomaly in a specific time period, the weight of the corresponding vertical main index node will be automatically increased, and the storage priority of the relevant spatiotemporal grid in the horizontal level will be adjusted synchronously. The construction of the hierarchical index structure needs to rely on feature fusion algorithms (such as tensor decomposition or graph neural networks) to achieve efficient mapping of cross-dimensional features and dynamic updating of index nodes.

[0033] As an implementation mode, the above step S300, based on the cross-fusion result of the temporal distribution feature and the spatial correlation topology feature, constructs a hierarchical index structure of meteorological metadata, which may include: Step S310: Decompose the time series distribution characteristics into periodic characteristics and mutation characteristics of meteorological elements, wherein the periodic characteristics represent the fluctuation pattern of meteorological elements that recurs in historical monitoring data, and the mutation characteristics represent the abnormal fluctuation range that exceeds a preset threshold.

[0034] The decomposition process of time series distribution features is, for example, to separate independent characteristic components of periodic changes and sudden changes from the fused time series feature vector output by the model. Periodic features are extracted by analyzing the regular fluctuation patterns of meteorological elements that recur in historical monitoring data, such as the annual seasonal changes in temperature data (high temperature in summer and low temperature in winter) or the monthly average distribution of precipitation data (such as concentrated precipitation during the plum rain period). Mutation features are determined by detecting abnormal fluctuation intervals in the numerical sequence of meteorological elements that exceed the preset threshold. The preset threshold can be dynamically set according to the statistical distribution of historical data (such as three times the standard deviation) or domain knowledge (such as the critical value of typhoon wind speed). For example, for wind speed data, if the wind speed increases from 5m / s to 25m / s in a certain time period and exceeds the preset threshold of 20m / s, the interval is marked as a mutation feature interval. During the decomposition process, the periodic features can be quantified by Fourier transform or autocorrelation analysis to quantify their fluctuation amplitude and frequency, while the mutation features can be located by combining the extreme value detection algorithm (such as peak recognition) in the sliding window with the threshold comparison. The separation of periodic features and mutation features provides a basis for differentiated weights for subsequent hierarchical division.

[0035] Step S320: decomposing the spatial correlation topological features into geographic correlation features and cross-regional conduction features, wherein the geographic correlation features represent the spatial continuity of the monitoring data of adjacent meteorological stations, and the cross-regional conduction features represent the correlation strength of meteorological elements between non-adjacent regions.

[0036] The decomposition of spatial correlation topological features aims to distinguish the local correlation and long-distance transmission of meteorological elements in the spatial dimension. Geographic correlation features are generated by calculating the spatial continuity index (such as Pearson correlation coefficient or covariance matrix) of the monitoring data of adjacent meteorological stations. For example, the temperature data of the meteorological stations in Chaoyang District and Haidian District of City B show a high degree of synchronization in the daily variation trend (correlation coefficient 0.92), indicating that the two stations have a strong geographical correlation. Cross-regional transmission features are extracted by analyzing the statistical dependence or physical transmission path (such as wind field and ocean current) of meteorological elements between non-adjacent regions. For example, when a typhoon moves from Sea Area C to the southeast coast, its wind speed and precipitation data show a lagged correlation (transmission intensity 0.75) between the non-adjacent meteorological stations in Xiamen and Fuzhou. The decomposition process uses a graph neural network or a spatial lag model to divide the edge weights in the spatial topological structure into local adjacent edges (geographic correlation features) and long-distance connection edges (cross-regional transmission features). For example, in the wind speed transmission network, the geographical correlation feature corresponds to the edge weight between adjacent sites (0.92), and the cross-regional transmission feature corresponds to the edge weight between inter-provincial sites (0.68). The independent extraction of geographical correlation features and cross-regional transmission features provides a multi-scale spatial relationship basis for the division of spatial density grids.

[0037] Step S330: In the vertical hierarchical division, a main index node is created according to the meteorological element type, and a dynamic weight factor is assigned to each main index node. The dynamic weight factor is determined by the ratio of the periodic feature to the mutation feature, so that the main index node of the high-frequency mutation element obtains a higher vertical hierarchical priority.

[0038] The vertical hierarchical division constructs independent main index nodes based on the meteorological element type (such as temperature, humidity, and wind speed), and each node represents a data set of a type of meteorological element. The dynamic weight factor is calculated by the ratio of periodic characteristics to mutation characteristics. For example, if the periodic characteristic intensity of a wind speed element is 0.6 (weak annual seasonal fluctuations) and the mutation characteristic intensity is 0.9 (frequent sudden strong wind events), the ratio is 0.9 / 0.6=1.5. This ratio reflects the degree of advantage of the mutation frequency of the element over the periodic stability. The higher the ratio, the more priority the element needs to be processed. The vertical hierarchical priority is sorted according to the ratio. For example, when the wind speed element ratio is 1.5 and the temperature element ratio is 0.8, the wind speed main index node is assigned a higher hierarchical position to ensure that high-frequency mutation data is given priority during retrieval and storage. The calculation of the dynamic weight factor needs to be combined with real-time data updates. For example, the intensity of the wind speed mutation characteristic in the typhoon season is increased to 1.2, and its ratio becomes 1.2 / 0.6=2.0, further consolidating its hierarchical priority. The creation and weight allocation of the main index node are implemented through the database index management module, which supports the linkage between dynamic adjustment of node position and storage resource allocation strategy.

[0039] Step S340: In the horizontal hierarchical division, a spatial density grid is generated based on the superposition results of geographic correlation features and cross-regional conduction features. Each grid unit records the statistical distribution characteristics of meteorological elements within the corresponding geographical range, and the grid unit is divided into a core storage area and an edge cache area according to the spatial density value.

[0040] The generation of spatial density grid is achieved by superimposing the spatial weights of geographic correlation features and cross-regional transmission features. The spatial weight matrix of geographic correlation features (such as 0.92 for adjacent site areas) and the spatial weight matrix of cross-regional transmission features (such as 0.68 for remote transmission areas) are linearly weighted and fused to generate a comprehensive spatial density value. For example, the comprehensive spatial density value of area B in the city is , while the density of the suburbs of city T is . Each grid unit (such as 1km×1km) records its density value and the statistical distribution characteristics of the corresponding meteorological elements (such as mean, variance, and extreme values). The division of the core storage area and the edge cache area is based on a preset density threshold. For example, grid cells with a density value ≥0.8 are designated as core storage areas (such as urban area B) to store high-access frequency or high-priority data; cells with a density value <0.8 are designated as edge cache areas (such as urban T suburbs) to store low-frequency or archived data. The division of spatial density grids needs to be updated periodically. For example, during the passage of a typhoon, the conduction feature weight of the affected area is increased, and the core area expansion is triggered when the density value exceeds the threshold.

[0041] Step S350: Establish a bidirectional mapping rule between the main index node of the vertical level and the spatial density grid of the horizontal level. When the meteorological file matches the dynamic weight factor threshold of the main index node and the core storage area condition of the spatial density grid at the same time, the cross-level joint index mark is triggered to generate a composite index tag containing spatiotemporal dimensions and element attributes.

[0042] The bidirectional mapping rule is implemented by associating the main index node attributes (such as dynamic weight factors) of the vertical level with the spatial density grid attributes (such as core storage area identifiers) of the horizontal level. The cross-level joint index tag is triggered when the meteorological file meets the following conditions at the same time: 1) the dynamic weight factor of the main index node of the meteorological element type to which it belongs exceeds the preset threshold (such as ≥1.0); 2) the spatial density grid unit corresponding to its geographical scope is the core storage area. For example, a typhoon wind speed file is marked as a cross-level joint index object because the dynamic weight factor of the wind speed main index node is 2.0 (exceeding the threshold 1.0) and is located in the core storage area (density value 0.85). The composite index label is generated by splicing the spatiotemporal dimension code (such as "2023-07-15_City B Core Area") and the feature attribute code (such as "wind speed_high frequency mutation"), such as "ARPU_WIND_20230715_BJ_CORE". After the label is generated, the index management system optimizes the storage location (such as allocating it to high-bandwidth nodes) and the retrieval path (such as preferentially matching core area data) based on the label content. Bidirectional mapping rules are implemented through database triggers or rule engines to ensure that label marking and resource allocation logic are automatically executed when data is written or updated.

[0043] Step S400: calling the meteorological storage optimization model, generating a storage node allocation strategy for the meteorological file according to the hierarchical weight distribution of the hierarchical index structure, the storage node allocation strategy including the compression rate threshold and storage location priority of each meteorological element data.

[0044] The meteorological storage optimization model is a decision model based on machine learning training. Its input is the weight distribution data of each level in the hierarchical index structure, and its output is the allocation rules of storage resources. Optionally, it can be implemented as a hybrid model combining deep learning and reinforcement learning. The long-term trend characteristics and short-term fluctuation characteristics of meteorological elements are extracted through bidirectional LSTM (BiLSTM), and the spatial topology model is constructed using graph convolutional network (GCN) to capture the geographical correlation and cross-regional transmission characteristics of adjacent meteorological stations. Then, the multi-head self-attention mechanism is used to fuse the temporal and spatial features to generate multidimensional encoding; the strategy generation module is based on the Actor-Critic reinforcement learning framework. The Actor network generates storage allocation strategies according to the real-time hierarchical weight distribution and node load status, and the Critic network optimizes the decision by evaluating the strategy effect; the feedback learning module adjusts the model parameters according to the changes in node capacity utilization after migration, and dynamically updates the feature fusion weights through meta-learning, and ensures the data integrity and query continuity of the migration process through index consistency verification (such as hash value comparison) and atomic lock mechanism, finally realizing efficient storage, real-time migration and intelligent retrieval of meteorological data.

[0045] The hierarchical weight distribution includes the heat score of meteorological element types in the vertical hierarchy (e.g., elements with high frequency of access or high mutation have higher weights) and the access frequency prediction value of the spatiotemporal grid in the horizontal hierarchy. The storage node allocation strategy specifically includes two parts: the compression rate threshold determines the selection of compression algorithms for different meteorological element data (e.g., lossless compression is used for high-precision temperature data, and lossy compression is used for historical wind speed data), and the storage location priority determines the distribution of data in distributed storage nodes (e.g., high-bandwidth nodes store real-time typhoon data, and low-cost nodes store archived precipitation data). For example, for the core grid data of temperature elements with higher weights, the model may specify a low compression rate (retaining high precision) and allocate it to high-performance storage nodes; while for the historical humidity data of the edge grid, a high compression rate is used and stored in the cold backup node. The training of the meteorological storage optimization model needs to be combined with the historical storage load data, the node hardware performance parameters, and the hierarchical weight change trend of the index structure, and the accuracy and resource utilization of the strategy generation process are optimized through supervised learning.

[0046] As an implementation mode, in step S400, the meteorological storage optimization model is called to generate a storage node allocation strategy for the meteorological file according to the hierarchical weight distribution of the hierarchical index structure, which may specifically include: Step S410: extract the dynamic weight factor of each main index node in the vertical hierarchy, and generate a feature type heat weight sequence, wherein the heat weight is exponentially positively correlated with the dynamic weight factor.

[0047] The dynamic weight factor of each main index node in the vertical hierarchy is, for example, a meteorological element type priority score calculated based on the fusion result of the temporal distribution characteristics and the spatial correlation topological characteristics. For example, the dynamic weight factor of the temperature element is 1.5, and the dynamic weight factor of the wind speed element is 2.3. The generation process of the element type heat weight sequence maps the dynamic weight factor to the storage priority weight through an exponential function. The specific formula is heat weight = e (k×动态权重因子) , where k is the scaling factor determined during the model training phase (e.g., k=0.8). The exponential positive correlation means that the higher the dynamic weight factor, the exponential growth of the heat weight, which significantly increases the storage resource allocation weight of high-priority elements. For example, when the dynamic weight factor of the wind speed element is 2.3, its heat weight is e (0.8×2.3) ≈8.17; and the dynamic weight factor of the temperature element is 1.5, which corresponds to the heat weight e (0.8×1.5) ≈3.32. The heat weight sequence is sorted by element type to form a list, such as [wind speed: 8.17, precipitation: 5.43, temperature: 3.32], which provides a quantitative basis for subsequent storage node type matching.

[0048] Step S420: extracting the boundary threshold between the core storage area and the edge cache area of ​​the spatial density grid in the horizontal level, generating a regional storage density gradient map, the gradient map reflecting the data access frequency prediction values ​​of different geographic grid units.

[0049] The threshold between the core storage area and the edge cache area of ​​the spatial density grid is determined by the statistics of historical data access frequency. For example, the threshold of the core storage area is set to a spatial density value ≥ 0.8, and the threshold of the edge cache area is set to a spatial density value < 0.8. The generation process of the regional storage density gradient map maps the spatial density value of each geographic grid cell (such as 1km×1km) into a color gradient or a numerical matrix. For example, the grid cell with a density value of 0.9 is represented by dark red (high-frequency access to the core area), and the grid cell with a density value of 0.6 is represented by light blue (low-frequency access to the edge area). The data access frequency prediction value is based on the time series analysis model (such as ARIMA or LSTM) to predict the historical access log. For example, it is predicted that the number of visits to the core grid cell of city B in the next week will be 1000 times per day, while the number of visits to the edge grid cell of city T will be 200 times per day. The gradient map records the density value and predicted access frequency of each grid cell in the form of visualization or matrix coding. For example, the cell in the i-th row and j-th column of the matrix stores the value 0.9 (density value) and the predicted value 1000 (number of visits), providing spatial dimension input for the geographic storage strategy.

[0050] Step S430: Input the element type heat weight sequence into the first feature processing layer of the meteorological storage optimization model to generate an element priority queue, which arranges the best storage node type for each meteorological element type in descending order of heat weight.

[0051] The first feature processing layer of the meteorological storage optimization model is composed of a fully connected neural network, whose input is a numerical vector of the heat weight sequence of the feature type (such as [8.17, 5.43, 3.32]), and the output is the storage node type label matching each feature type (such as high-bandwidth node, medium-performance node, cold storage node). The generation process of the feature priority queue arranges the feature types from high to low according to the heat weight through a sorting algorithm (such as quick sort), and assigns a preset node type rule to each type. For example, features with a heat weight ≥ 5.0 (such as wind speed and precipitation) are assigned to high-bandwidth storage nodes (SSD arrays), features with a weight of 2.0~5.0 (such as temperature) are assigned to medium-performance nodes (SATA hard drives), and features with a weight <2.0 (such as historical humidity) are assigned to cold storage nodes (tape libraries). The specific form of the queue can be a list structure, such as [wind speed: high-bandwidth node, precipitation: high-bandwidth node, temperature: medium-performance node], to ensure that high-heat features take up high-performance storage resources first.

[0052] Step S440: inputting the regional storage density gradient map into the second feature processing layer of the meteorological storage optimization model to generate a geographic storage strategy matrix, in which each cell in the matrix records the storage position adjustment frequency of the corresponding grid within a preset time window; The second feature processing layer uses a convolutional neural network to extract features from the regional storage density gradient map and output a geographic storage strategy matrix. Each cell of the matrix corresponds to a geographic grid, recording the number of times the grid needs to adjust its storage location within a preset time window (such as the next 24 hours). For example, the core grid cell of city B has a high predicted access frequency (1000 times / day), so its adjustment frequency is set to once per hour (i.e., 24 times per day), while the edge grid cell of city T has an adjustment frequency set to once per day. The calculation of the adjustment frequency combines the spatial density value, the predicted access frequency value, and the node load balancing strategy. For example, when the grid density value is ≥0.8 and the predicted access frequency exceeds the node throughput threshold, the adjustment frequency is automatically increased to 2 times per hour. The geographic storage strategy matrix is ​​stored in a two-dimensional array or database table. For example, the cell in the i-th row and j-th column of the matrix stores the value 24 (number of adjustments), which provides a basis for the execution frequency of dynamic migration operations.

[0053] Step S450: Integrate the element priority queue and the geographic storage strategy matrix to generate a storage node allocation strategy, in which meteorological file files of element types with heat weights greater than a preset value are allocated to the target bandwidth storage node, and based on the grid cells in the geographic storage strategy matrix whose adjustment frequency exceeds the threshold, increase the number of redundant copies and the frequency of cross-node backup for the associated files.

[0054] The fusion process logically superimposes the rules of the feature priority queue and the geographic storage strategy matrix through the strategy generation algorithm. The meteorological file files of the feature types with heat weight greater than the preset value, that is, the high heat feature types (such as wind speed and precipitation), are allocated to the target bandwidth storage node according to the queue indication. The target bandwidth storage node is a high bandwidth storage node. For example, typhoon monitoring data is stored in the SSD array to ensure low-latency access. At the same time, the grid unit whose adjustment frequency exceeds the preset threshold (such as once per hour) in the geographic storage strategy matrix triggers the redundant copy strategy, such as creating 3 copies of the typhoon data of the core grid of City B, and performing a cross-node backup every 30 minutes (such as synchronizing from node A to node B and node C). The specific output of the storage node allocation strategy includes the storage location path (such as " / ssd_node1 / typhoon_data"), the number of copies (such as 3) and the backup period (such as 30 minutes), and is sent to the distributed storage system for execution through the configuration file or API instruction. For example, a meteorological file belongs to the wind speed element (high heat) and is located in the core grid with a high adjustment frequency. Its allocation strategy is marked as "high-bandwidth node, 3 copies, 30-minute backup", so as to meet both performance and disaster recovery requirements.

[0055] As an implementation method, the training process of the meteorological storage optimization model may include: Step S401: Acquire historical meteorological metadata storage records, which include hierarchical weight distribution samples of a hierarchical index structure, storage node allocation strategy labels, and node load fluctuation time series data.

[0056] The acquisition process of historical meteorological metadata storage records refers to extracting historical operation records related to meteorological file storage from the log database of the distributed storage system. The hierarchical weight distribution samples of the hierarchical index structure include the dynamic weight factor change sequence of each main index node in the vertical layer (such as the dynamic weight factor of the wind speed element node in the typhoon season increases from 1.8 to 2.5) and the core storage area coverage ratio of the spatial density grid in the horizontal layer (such as the coverage rate of the core grid of City B expands from 60% to 85%). The storage node allocation strategy label is a storage rule manually labeled or automatically generated by the system, such as allocating wind speed data with high dynamic weight to high-bandwidth nodes and creating 3 redundant copies for the core grid with high-frequency adjustment. The node load fluctuation time series data records the curves of the CPU occupancy rate, disk throughput and network latency of each storage node over time. For example, the CPU occupancy rate of a certain SSD node peaked at 90% on July 15, 2023, while the disk throughput remained at 500MB / s during the same time period. The collection of historical data needs to cover different meteorological events (such as typhoons and rainstorms) and hardware load scenarios (such as node expansion and fault switching) to ensure the diversity of training samples and the generalization ability of the model.

[0057] Step S402: Convert the hierarchical weight distribution samples into a multidimensional training feature vector, wherein the first component of the multidimensional training feature vector corresponds to the dynamic weight factor change trend of the main index node in the vertical hierarchy, the second component corresponds to the core storage area coverage ratio of the spatial density grid in the horizontal hierarchy, and the third component corresponds to the compatibility parameter between the meteorological element type and the storage node hardware configuration.

[0058] The construction process of the multi-dimensional training feature vector realizes the multi-dimensional quantization of the hierarchical weight distribution samples through mathematical coding. The first component dynamic weight factor change trend is generated by calculating the dynamic weight factor mean, variance and slope statistics of the main index node within a time window. For example, the dynamic weight factor mean of the wind speed node in the typhoon season is 2.3, variance 0.5 and daily growth rate 0.2. The second component core storage area coverage ratio is calculated by counting the number of grid cells in the horizontal layer whose spatial density value exceeds the preset threshold (such as ≥0.8) to the total grid. For example, when the core grid coverage ratio of city B increases from 60% to 85%, this component is 0.85. The third component compatibility parameter is defined based on the matching rule between the meteorological element type and the hardware performance of the storage node. For example, the wind speed data has a matching score of 0.9 due to the high real-time requirements and the low latency characteristics of the SSD node, while the compatibility score of the historical temperature data with the high-capacity mechanical hard disk is 0.7. The multi-dimensional training feature vectors are normalized to a unified numerical range. For example, the mean of the dynamic weight factor is mapped to the interval [0,1] and the compatibility parameter is converted by percentage, so as to avoid the interference of feature scale differences on model training.

[0059] Step S403: In the model initialization stage, the multi-dimensional training feature vector is input into the convolutional feature extraction layer of the meteorological storage optimization model, and the local correlation pattern of the hierarchical weight distribution is captured through the sliding window to generate the initial feature code.

[0060] The convolution feature extraction layer uses a one-dimensional convolution kernel to slide in the time dimension to capture the local correlation between the dynamic weight factor change trend and the core storage area coverage ratio. The length of the sliding window is set according to the data period, such as a 24-hour window to capture daily cycle load fluctuations, or a 30-day window to analyze monthly cycle weight changes. The output of each convolution kernel corresponds to a local correlation pattern. For example, a convolution kernel may identify the correlation rule that "when the daily growth rate of the dynamic weight factor exceeds 0.1, the core storage area coverage ratio increases by 5% simultaneously." The initial feature encoding is generated by splicing the results of multi-channel convolution. For example, the input feature vector dimension is 3 (dynamic weight, coverage ratio, compatibility), and a 32-dimensional encoding vector is generated after being processed by 32 convolution kernels. The encoding process retains the key patterns of hierarchical weight distribution, such as the strong correlation between the surge in wind speed node weights and the expansion of the core grid during typhoon events, providing a local feature basis for subsequent attention allocation.

[0061] Step S404: Input the initial feature code into the multi-head attention allocation layer of the meteorological storage optimization model, calculate the influence weights of different hierarchical dimensions on the storage node allocation strategy, and generate a global feature code with attention weights.

[0062] The multi-head attention allocation layer focuses on the contribution of features at different levels through multiple parallel self-attention mechanisms. Each group of attention heads independently calculates the query vector, key vector, and value vector. For example, the first group of attention heads focuses on analyzing the impact of the dynamic weight factor change trend on the storage strategy (for example, for every 0.1 increase in the weight factor slope, the probability of high-bandwidth node allocation increases by 15%), and the second group of attention heads focuses on the relationship between the core storage area coverage ratio and the number of redundant copies (for example, for every 10% increase in the coverage ratio, the number of copies needs to increase by 1). The attention weights are normalized by the Softmax function to generate the contribution distribution of each feature, such as the dynamic weight factor feature contribution of 0.6, the core storage area feature contribution of 0.3, and the compatibility feature contribution of 0.1. The global feature encoding is generated by weighted summing the outputs of all attention heads, for example, multiplying the 32-dimensional initial encoding with the attention weight matrix and reducing the dimension to 16 dimensions, thereby integrating the global correlation of multi-level features.

[0063] Step S405: input the global feature code into the strategy generation layer of the meteorological storage optimization model, and generate a candidate storage node allocation strategy set in combination with the periodic law of the node load fluctuation time series data.

[0064] The policy generation layer consists of a fully connected neural network and a rule engine, and its input is the global feature encoding and node load fluctuation time series data. The periodicity of the node load fluctuation time series data is extracted through Fourier transform or period detection algorithm. For example, the CPU utilization of a node shows a periodic peak from 9:00 am to 12:00 pm every day, and the disk throughput decreases periodically at the end of the month due to data archiving operations. The policy generation layer predicts the load status of the future time window based on the periodicity (such as the CPU utilization is expected to reach 85% at 9:00 tomorrow) and generates an adaptive candidate policy. For example, for the predicted high CPU load period, a policy of "limiting the amount of new data written to the node to 50GB" is generated; for the low disk throughput period, a policy of "starting cross-node backup to alleviate I / O bottlenecks" is generated. The candidate strategy set is generated by enumerating possible node allocation rules (such as the number of replicas, storage location, and backup frequency) and screening solutions that meet the load constraints. For example, three strategies are generated: Strategy A (high-bandwidth nodes, 2 replicas, and hourly backups), Strategy B (medium-performance nodes, 3 replicas, and daily backups), and Strategy C (cold storage nodes, 1 replica, and no backups).

[0065] As an implementation mode, in step S405, the process of generating a candidate storage node allocation strategy set in combination with the periodicity of the node load fluctuation time series data may include: Step S4051: input the global feature code and the node load fluctuation time series data into the load balancing prediction submodule of the meteorological storage optimization model to predict the CPU occupancy rate curve and disk throughput change trend of each storage node in the future time window.

[0066] The load balancing prediction submodule uses a time series prediction model (such as LSTM or Transformer), whose input is the global feature encoding (characterizing the correlation between the hierarchical weight and the storage strategy) and the node load fluctuation time series data (historical CPU usage, disk throughput). The prediction process analyzes the periodicity (such as daily peaks and weekly troughs) and trend (such as long-term hardware performance decay) of the historical load data, and outputs the CPU usage curve (such as 85% predicted value at 9:00 and 60% predicted value at 15:00) and the disk throughput change trend (such as 500MB / s in the morning and 200MB / s in the night) of each node in the future time window (such as the next 24 hours). For example, for a certain SSD node, the model predicts that the CPU usage will be 88% and the disk throughput will be 550MB / s at 10:00 the next day, triggering the high-voltage node identification.

[0067] Step S4052: Generate a storage node load pressure level identifier based on the predicted CPU occupancy rate curve, and the identifier is divided into three categories: high-voltage node, medium-voltage node and low-voltage node.

[0068] For example, the load pressure level identification can be generated by dividing the predicted CPU usage curve by threshold. For example, CPU usage ≥ 80% is a high-pressure node (red identification), 50%~80% is a medium-pressure node (yellow identification), and <50% is a low-pressure node (green identification). During the forecast period, a node with a CPU usage of 88% at 10:00 is marked as a high-pressure node, a node with a usage of 75% at 14:00 is marked as a medium-pressure node, and a node with a usage of 45% at 20:00 is marked as a low-pressure node. After the identification is generated, the storage policy needs to be dynamically adjusted according to the pressure level, for example, high-pressure nodes limit data writing to avoid overload.

[0069] Step S4053: Generate a storage node data transmission efficiency score based on the predicted disk throughput change trend, where the score is positively correlated with the slope value of the throughput increase trend.

[0070] The data transfer efficiency score is calculated by quantifying the change trend of disk throughput. For example, the throughput of a node in the morning period increases from 300MB / s to 600MB / s, the slope value is (600-300) / 3 hours = 100MB / s², and the score is set to 0.9 (full score 1.0); the throughput of another node decreases from 200MB / s to 150MB / s, the slope is negative, and the score is set to 0.3. The scoring formula can be defined as score = tanh(k×slope), where k is the scaling factor (such as k=0.01), ensuring that the score range is within [-1,1] and mapped to [0,1] through offset. Efficient nodes (score ≥ 0.7) take priority in real-time data transfer tasks, and inefficient nodes (score < 0.4) trigger backup or migration operations.

[0071] Step S4054: In the strategy generation layer, the load pressure level identifier and the data transmission efficiency score are jointly constrained and analyzed: for high-voltage nodes, the maximum capacity threshold for allocating new data is limited; for low-voltage nodes, the weight coefficient of receiving target priority data is increased; for nodes with a data transmission efficiency score lower than the preset standard, the cross-node replica backup mark is triggered.

[0072] Joint constraint analysis combines load pressure and transmission efficiency conditions through the rule engine to generate policy constraints. For example: 1) The threshold for new data capacity of high-voltage nodes (CPU ≥ 80%) is set to 30% of the current remaining capacity (e.g., 300GB is allowed to be written if there is 1TB remaining); 2) The weight coefficient of target priority data (that is, high priority data, such as wind speed elements) of low-voltage nodes (CPU < 50%) is increased from 0.6 to 0.9, so that it is selected first during allocation; 3) Nodes with a transmission efficiency score < 0.4 need to create at least 2 cross-node replicas for storage data. The analysis process needs to balance load balancing and storage efficiency. For example, while restricting writes on high-voltage nodes, the weight of low-voltage nodes can be increased to ensure that high-priority data can still be stored in a timely manner.

[0073] Step S4055: Based on the results of the joint constraint analysis, a set of candidate storage node allocation strategies is generated. Each strategy in the set contains the load pressure compatibility parameters, data transmission efficiency guarantee parameters and redundant backup execution conditions of the target node, and the parameter scale consistency between different strategies is ensured through normalization processing of the strategy generation layer.

[0074] The candidate strategy set is generated by enumerating all feasible node allocation schemes and applying constraints to filter them. The parameters of each strategy include: 1) load pressure compatibility parameters (such as the maximum write volume allowed for high-pressure nodes is 300GB); 2) data transmission efficiency guarantee parameters (such as the minimum throughput of efficient nodes is 500MB / s); 3) redundant backup execution conditions (such as the need to create 2 copies for inefficient nodes). Normalization processing unifies the parameter scale through Min-Max scaling or Z-score standardization, for example, mapping the write volume from 0GB~1000GB to the range of 0~1, and the throughput from 0MB / s~1000MB / s to the range of 0~1. The final strategy set is sorted by comprehensive score (such as compatibility × efficiency × backup guarantee), and the Top-K (such as Top 5) optimal strategies are output for system execution. For example, the comprehensive score of strategy A is 0.92, strategy B is 0.85, and strategy C is 0.78. The system prefers strategy A to implement storage allocation.

[0075] Step S406: In the model optimization stage, the candidate storage node allocation strategy set is compared with the storage node allocation strategy label in the storage record, the strategy offset loss value is calculated, and the network parameters of the convolutional feature extraction layer, the multi-head attention allocation layer and the strategy generation layer are adjusted through error back propagation until the matching degree between the allocation strategy output by the model and the storage node allocation strategy label reaches a preset threshold.

[0076] The policy shift loss value is measured by comparing the difference between the candidate policy and the policy label, such as using the cross entropy loss function to compare the probability distribution of the policy, or the mean square error function to compare numerical parameters such as the number of replicas and storage location. For example, if the policy label specifies "high bandwidth nodes, 3 replicas", and candidate policy A is "high bandwidth nodes, 2 replicas", the loss value caused by the difference in the number of replicas is 1. The error backpropagation process adjusts the parameters of each network layer through chain derivation, such as the convolution kernel weights, the query matrix of the attention head, and the bias term of the fully connected layer. The optimization iteration continues until the policy output by the model matches the preset threshold on the validation set (such as accuracy ≥ 95% or loss value ≤ 0.05). After training, the model can generate a storage allocation plan that is highly consistent with the historical optimal policy based on the real-time layer weights and node load data.

[0077] Step S500: According to the storage node allocation strategy, each meteorological file in the meteorological metadata set is written into the corresponding distributed storage node, and based on the real-time update mechanism of the hierarchical index, the physical location and index mapping relationship of the stored files are dynamically adjusted.

[0078] Distributed storage nodes are networked storage systems composed of multiple physical or virtual storage devices, whose nodes may be distributed in different geographical locations or data centers. The writing process follows the compression rate threshold and location priority in the storage node allocation strategy. For example, high-priority meteorological file files are written to the SSD storage array through a dedicated transmission channel, and the Zstandard or GZIP algorithm is called to encode the file at a specified compression rate. The real-time update mechanism of the hierarchical index is, for example, to continuously monitor the weight changes of the index structure (such as new meteorological events causing the weight of a certain spatiotemporal grid to increase) and the load status of the storage node (such as insufficient node capacity or increased access latency), and trigger data migration or index reconstruction. For example, when the rainstorm monitoring data in a certain area is frequently accessed due to real-time warning needs, the system automatically migrates it from the edge cache node to the core storage node, and updates the index mark of the corresponding grid in the horizontal hierarchy as "high priority". The dynamic adjustment process needs to ensure the atomicity of data migration (avoid query failure during migration) and the consistency of index mapping (ensure that both new and old paths can be resolved after migration). In addition, the real-time update mechanism needs to be linked with the meteorological storage optimization model to optimize the generation efficiency of subsequent allocation strategies and the overall performance of the storage system through feedback loops.

[0079] As an implementation mode, in step S500, the process of dynamically adjusting the mapping relationship between the physical location and the index of the stored file includes: Step S510: Collect the capacity utilization rate and data access frequency of the distributed storage nodes in real time to generate a node load state vector, which includes the remaining capacity percentage of the storage node, the number of input and output operations per unit time, and the average response delay time.

[0080] The process of real-time collection of capacity utilization and data access frequency of distributed storage nodes is executed periodically by the monitoring agent deployed on the storage node. The monitoring agent collects the hardware resource indicators of the node at preset time intervals (such as every minute). Capacity utilization refers to the ratio of the used capacity of the node storage medium to the total capacity. For example, the total capacity of a node is 10TB, the used capacity is 7TB, and the remaining capacity percentage is 30%. Data access frequency is calculated by counting the number of data read and write operations (IOPS) on the storage node within a unit time (such as per second). For example, a node processes 1000 read operations and 500 write operations per second, and the total number of input and output operations is 1500 times / second. Average response delay time refers to the average time from receiving a request to returning a result. For example, the average delay of a node's read operation is 20 milliseconds, the average delay of a write operation is 50 milliseconds, and the comprehensive average response delay time is 35 milliseconds. The node load status vector encapsulates the above indicators through a structured data format (such as JSON or Protocol Buffers). For example, the vector is represented as [remaining capacity percentage: 30%, number of input and output operations: 1500, average response delay time: 35ms], providing real-time input for subsequent migration decisions.

[0081] Step S520: Input the node load state vector into the real-time decision submodule of the meteorological storage optimization model, and calculate the heat priority score of the file to be migrated in combination with the heat decay curve of the hierarchical weight distribution in the hierarchical index structure. The heat priority score is determined by the decay rate of the dynamic weight factor of the main index node of the vertical level to which the file belongs and the decreasing trend of the access frequency of the horizontal spatial density grid.

[0082] Exemplarily, the real-time decision submodule processes the node load state vector through a machine learning model (such as a random forest or a gradient boosting tree) and integrates the heat decay curve of the hierarchical index structure. The heat decay curve is modeled based on a time decay function (such as an exponential decay), reflecting the rate of decrease of the dynamic weight factor of the meteorological element over time. For example, the dynamic weight factor of the typhoon wind speed element decays by 5% every day after the typhoon passes. The access frequency decline trend of the horizontal spatial density grid is calculated by analyzing the sliding window mean of the historical access log. For example, the access frequency of a core grid unit drops from 1,000 times per day to 200 times per day, with a decline rate of 80 times per day. The calculation formula for the heat priority score is: score = dynamic weight factor decay rate × α + access frequency decline trend × β, where α and β are weighting coefficients determined in the model training phase (such as α=0.6, β=0.4). For example, if the dynamic weight factor decay rate of a typhoon file is 5% / day and the access frequency decreases by 80 times / day, the score = 5×0.6 +80×0.4=3+32=35. The lower the score, the faster the popularity of the file decays and the higher the migration priority.

[0083] Step S530: Generate a migration task queue according to the heat priority score, arrange the files to be migrated in ascending order according to the score, and mark the hardware performance matching parameters of the target migration node; Exemplarily, the generation of the migration task queue can sort the files to be migrated from low to high according to the heat priority score through a sorting algorithm to ensure that low-scoring files are migrated first. The hardware performance matching parameter is calculated based on the remaining capacity, input and output operation capability and delay characteristics of the target node. For example, the remaining capacity of target node A is 40%, the upper limit of the number of input and output operations is 2000 times / second, and the average delay is 25ms, and its matching score is 0.8 (full score 1.0); the remaining capacity of node B is 20%, the upper limit of the number of input and output operations is 1200 times / second, the average delay is 50ms, and the matching score is 0.5. Each entry in the queue records the source node path, the target node path, the migration file size and the matching parameters, such as [source node: / node1 / typhoon_20230715.data, target node: / node3 / archive / , file size: 50GB, matching: 0.8]. After the queue is generated, the migration scheduler selects the optimal target node based on the matching parameters, for example, the file is preferentially migrated to the node with a score ≥0.7 to ensure performance.

[0084] Step S540: When performing the incremental migration operation, the meteorological file with the lowest score in the queue and which has not been accessed for a period exceeding a preset threshold is preferentially migrated, and after creating a migration copy in the target node, the spatial density grid association mark of the hierarchical index structure is synchronously updated.

[0085] Exemplarily, the incremental migration operation only migrates the newly added or modified parts of the file to reduce network bandwidth consumption. The preset threshold is set according to business needs. For example, files that have not been accessed for 30 consecutive days trigger migration. For example, a historical humidity file has a heat priority score of 10 (the lowest score) and was last accessed on June 1, 2023, which exceeds the 30-day threshold and is included in the first migration task. When the migration is executed, the source node transfers the file to the target node in blocks, and after the target node completes the copy verification, it updates the spatial density grid mark associated with the file in the hierarchical index structure. For example, after a file originally marked as "Core Storage Area-City B Grid A01" is migrated to the edge cache area, the mark is changed to "Edge Cache Area-City T City Grid B02". The update operation ensures atomicity through database transactions to avoid inconsistent index status.

[0086] Step S550: During the migration process, an atomic lock mechanism is implemented on the index mapping relationship between the source node and the target node to ensure that the query request always points to the valid storage location during the migration, and triggers the index consistency check after the migration is completed. When a cross-node index path conflict is detected, it automatically rolls back to the pre-migration state and regenerates the migration task queue.

[0087] Exemplarily, the atomic lock mechanism can be implemented through a distributed lock service (such as ZooKeeper or Etcd). When the migration starts, the source node file is locked to prohibit concurrent write or delete operations. The query request is redirected to the source node or the target node replica during the lock period. For example, if the migration is not completed, the request is still responded by the source node; if the migration is completed, the request is switched to the target node. The index consistency check detects inconsistent items by comparing the file hash values, index labels and path mapping relationships between the source node and the target node. For example, when it is detected that the hash value of the target node replica is inconsistent with the source node, it is determined to be a conflict and triggers an automatic rollback: delete the target node replica, restore the source node index mark, and reinsert the file into the head of the migration queue. The rollback log records the exception information for subsequent analysis, and the migration task queue is regenerated according to the latest node load status. For example, after a rollback due to a target node failure, the standby node C is selected as the new target.

[0088] As an implementation method, the retrieval process of meteorological metadata includes the following steps: Step S600: receiving a search request submitted by a user, parsing the time and space range constraints and meteorological element type combination conditions in the search request, and generating a multi-dimensional search condition vector, which includes a geographic boundary coordinate set, a time interval stamp, and an element type coding sequence.

[0089] Exemplarily, the process of receiving the search request submitted by the user is implemented through the front-end interface or API gateway, and the parsing process extracts the spatiotemporal range constraints and meteorological element type combination conditions in the request. The spatiotemporal range constraints include a geographic boundary coordinate set and a time interval stamp. The geographic boundary coordinate set defines the target search area with a sequence of longitude and latitude coordinate pairs (for example, the boundary coordinate set of city B is [116.20°E, 39.80°N], [116.60°E, 39.80°N], [116.60°E, 40.20°N], [116.20°E, 40.20°N]), and the time interval stamp defines the target period with a start timestamp and an end timestamp (for example, July 1, 2023 00:00:00 to July 31, 2023 23:59:59). The meteorological element type combination condition is represented by a predefined coding sequence (such as temperature is coded as TEMP_01 and wind speed is coded as WIND_02). For example, when the user requests to retrieve temperature and wind speed data at the same time, the element type coding sequence is [TEMP_01, WIND_02]. The multi-dimensional retrieval condition vector encapsulates the above parameters through structured data. For example, the vector format is {geographic boundary coordinate set: [[116.20,39.80], [116.60,39.80], [116.60,40.20], [116.20,40.20]], time interval stamp: [1688169600, 1690847999], element type coding sequence: [TEMP_01, WIND_02]}, providing standardized input for subsequent joint queries.

[0090] Step S700: Input the multi-dimensional search condition vector into the joint query engine of the hierarchical index structure, match all main index nodes whose dynamic weight factors are greater than the preset threshold in the vertical level, and simultaneously screen the geographic grid cells covered by the core storage area of ​​the spatial density grid in the horizontal level.

[0091] Exemplarily, the joint query engine of the hierarchical index structure processes the index conditions of the vertical and horizontal levels in parallel, and the vertical level matches the dynamic weight factor of the main index node based on the meteorological element type. The dynamic weight factor threshold is set according to real-time business needs (for example, the threshold is 1.0), and the main index nodes with a weight factor greater than this value are screened out (for example, the dynamic weight factor of the wind speed node in the typhoon season is 2.3, which meets the threshold). The horizontal level screening is based on the core storage area identification of the spatial density grid, and the core storage area is composed of geographic grid cells whose spatial density value exceeds the preset threshold (for example, ≥0.8) (for example, the density value of the core grid cell A01 of city B is 0.9, and the density value of the edge grid cell B02 of city T is 0.6). The joint query engine matches the element type coding sequence in the retrieval condition vector with the type of the main index node of the vertical level (for example, TEMP_01 matches the temperature node, WIND_02 matches the wind speed node), and at the same time, the geographic boundary coordinate set is intersected with the geographic range of the horizontal level grid cell (for example, the boundary of city B covers grid cells A01, A02, and A03). The matching result is the intersection set that meets the longitudinal weight condition and the transverse core area condition. For example, the intersection of the wind speed main index node (dynamic weight 2.3) and the core grid unit A01 of city B is identified as the effective search range.

[0092] Step S800: Generate a candidate meteorological file set based on the intersection result of the matched main index node and the geographic grid unit, each file in the set is marked with the compression algorithm type and storage location path recorded in the storage node allocation strategy.

[0093] Exemplarily, the candidate meteorological file set is extracted from the bidirectional mapping rules of the hierarchical index structure by the joint query engine, and the meteorological file corresponding to the intersection result must meet both the element type weight condition and the core storage area location condition. For example, the typhoon wind speed file (element type WIND_02) stored in the core grid unit A01 of city B in July 2023 is included in the candidate set because it matches the dynamic weight threshold and the core area condition. The metadata tag of each file contains the compression algorithm type defined in the storage node allocation strategy (such as the Zstandard compression algorithm identifier ZSTD_01) and the storage location path (such as the distributed node path / node_ssd_01 / typhoon_wind_202307.data). The candidate set is generated by a database query statement, such as the SQL query statement filtering "WHERE element type IN (TEMP_01, WIND_02) AND grid unit IN (A01, A02, A03)", and the result set is returned in a list form and annotated with compression and storage information.

[0094] Step S900: calling the real-time decompression interface of the distributed storage node, performing parallel decompression operations on the meteorological file files in the candidate set based on the compression algorithm type, and generating a standardized meteorological element monitoring data stream.

[0095] The real-time decompression interface of the distributed storage node calls the corresponding decoding library according to the compression algorithm type of the file, for example, the Zstandard compression algorithm calls the libzstd library, and the GZIP compression algorithm calls the zlib library. The parallel decompression operation distributes the files in the candidate set to multiple computing nodes for simultaneous processing through the task scheduler, for example, 10 wind speed file files are assigned to 10 CPU cores for parallel decompression. Standardized meteorological element monitoring data streams require that the decompressed data format is unified into a preset structure (such as NetCDF format or Parquet format), including fields such as timestamps, geographic coordinates, and element values. For example, the data stream generated after decompression of a typhoon wind speed file contains hourly timestamps (2023-07-15T12:00:00), longitude and latitude coordinates (116.40°E, 39.90°N) and wind speed values ​​(15m / s), and is arranged in chronological order as a time series data matrix. The standardization of data streams ensures that the subsequent correlation scoring model can consistently process monitoring data from different sources.

[0096] Step S1000: Input the standardized data stream into the factor relevance scoring model established in the meteorological storage optimization model training phase, calculate the spatiotemporal coverage completeness, factor matching degree and historical access heat weight of each meteorological file and the retrieval condition vector, generate the final retrieval result sorting list and return it to the user end.

[0097] Exemplarily, the factor relevance scoring model calculates multi-dimensional scoring indicators through machine learning algorithms (such as random forests or neural networks). The spatiotemporal coverage completeness evaluates the coverage ratio of the geographical and temporal scope of the file data to the search conditions (such as a file covering 80% of the area of ​​city B and 90% of the period in July); the factor matching evaluates whether the number of data points and the fluctuation of values ​​meet the search factor type (such as the number of wind speed data points in the search period accounts for 95% and the value fluctuation is within the historical mean ±2σ); the historical access heat weight is based on the access frequency of the file in the storage log and the recent attenuation trend (such as a file has been accessed 100 times in the past 30 days, but only 5 times in the last 7 days). The model weights and sums the above scores to generate a total relevance score. For example, file A scores 0.92 (coverage 0.8×0.4 + matching 0.95×0.3 + heat 0.7×0.3), and file B scores 0.85. The sorted list is sorted in descending order by total score (such as file A>file B>file C) and returned in JSON or table format through the user interface, supporting paging and highlighting of key data.

[0098] As an implementation method, the calculation process of the factor relevance scoring model includes: Step S1001: extracting the spatiotemporal identification information of the meteorological file, calculating the overlap ratio of its geographical coverage and the geographical boundary coordinate set in the search request, and generating a first dimension coverage score; The spatiotemporal identification information includes the coordinate set of the polygon vertices of the geographic area recorded in the file and the time interval stamp. The overlapping area ratio of the geographic coverage is calculated through the spatial overlay analysis of the geographic information system (GIS). For example, the geographic scope of file A is the Dongcheng District of City B (an area of ​​41.84 square kilometers), and the entire area of ​​City B requested for retrieval (an area of ​​16410.54 square kilometers). The overlapping area is 41.84 square kilometers, and the ratio is 41.84 / 16410.54≈0.255%. The coverage score is normalized to map the ratio to the range of 0~1 (such as 0.255% corresponds to a score of 0.00255), or logarithmic scaling is used to alleviate the deviation of the ratio of small areas (such as log(1+0.255)=0.23). If the file covers the entire retrieval area (accounting for 100%), the coverage score is 1.0; if there is no overlap, the score is 0.

[0099] Step S1002: Analyze the number of data points in the meteorological element monitoring data stream that match the retrieval request element type coding sequence, combine the numerical fluctuation range of the data points with the degree of deviation from the historical average, and generate a second dimension element matching score.

[0100] The matching degree of the number of data points is calculated by counting the total number of data points in the monitoring data stream that meet the search element type code (such as WIND_02) to the expected number of points in the search period. For example, the search period is 720 hours (30 days), and 1 data point is expected per hour. A file actually contains 700 wind speed data points, and the matching degree ratio is 700 / 720≈0.972. The value fluctuation range is calculated by calculating the standard deviation of the data point value and the historical mean (such as the average wind speed of the same period in the past five years is 12m / s). The deviation degree scoring formula is 1-(actual standard deviation / historical standard deviation). For example, if the actual standard deviation is 3m / s and the historical standard deviation is 4m / s, the score is 1-3 / 4=0.25. The total score of the element matching degree is the weighted average of the quantity ratio and the deviation score (such as 0.972×0.7+0.25×0.3=0.785).

[0101] Step S1003: Obtain the dynamic weight factor attenuation curve of the main index node in the hierarchical index structure, and generate a third-dimensional heat attenuation compensation coefficient in combination with the access frequency decrease rate of the file in the historical storage log.

[0102] The dynamic weight factor attenuation curve is fitted by an exponential function. For example, the dynamic weight factor of the wind speed node decays by 5% every day after the typhoon passes. The attenuation function is W(t)=W0×e (-0.05t) . The file access frequency decline rate is analyzed by linear regression to analyze the daily access times in the historical logs. For example, the access times of a certain file in the past 30 days have dropped from 10 times per day to 2 times per day, and the decline rate is Δ=-0.27 times / day. The heat decay compensation coefficient is calculated as the absolute value of the product of the slope of the decay curve (such as -0.05) and the access decline rate (-0.27) (|(-0.05)×(-0.27)|=0.0135), and then mapped to the range of 0~1 (such as 0.0135→0.503) by the Sigmoid function. The higher the coefficient, the slower the file heat decays, and it needs to be compensated in the score to delay the decline in ranking.

[0103] Step S1004: Input the first dimension coverage score, the second dimension element matching score and the third dimension heat attenuation compensation coefficient into the weight allocation function defined in the meteorological storage optimization model training phase. The function dynamically adjusts the weight ratio of each dimension score according to the business priority of each category in the element type coding sequence.

[0104] Exemplarily, the weight allocation function dynamically sets the weight of each dimension through the business priority table encoded by the element type. For example, the priority of the wind speed element (WIND_02) in the typhoon warning stage is 0.6, and the temperature element (TEMP_01) is 0.4. The function defines coverage weight = 0.4 × priority, matching weight = 0.3 × priority, and heat weight = 0.3 × priority. For the wind speed element, the weights of each dimension are coverage 0.24 (0.4 × 0.6), matching 0.18 (0.3 × 0.6), and heat 0.18 (0.3 × 0.6); for the temperature element, the weights are 0.16, 0.12, and 0.12 respectively. The weighted total score is calculated as coverage score × 0.24 + matching score × 0.18 + heat coefficient × 0.18, ensuring that high-priority elements have an advantage in the sorting.

[0105] Step S1005: normalize the total score after weighted summation to generate a standardized relevance score value, and sort the candidate meteorological file files in descending order according to the score value to form a final search result sorting list.

[0106] Exemplarily, the normalization process maps the total score to the interval of 0~1 through Min-Max scaling, for example, the highest total score of 0.92 is mapped to 1.0, and the lowest 0.75 is mapped to 0.75 / 0.92≈0.815. The standardized score value retains two decimal places (such as 0.92→0.92, 0.85→0.85), and the sorted list is sorted in descending order (0.92, 0.85, 0.78). The list entry contains the file name, relevance score, storage path and key data summary (such as "Typhoon Wind Speed_20230715.data, score 0.92, path / node_ssd_01, maximum wind speed 25m / s"). The user end can directly access or download highly relevant data by clicking on the entry. The sorting results are cached in a distributed memory database (such as Redis) to support rapid response and result reuse of subsequent retrieval requests.

[0107] In an optional derivative implementation, after dynamically adjusting the mapping relationship between the physical location and the index of the stored file in step S500, the method provided by the embodiment of the present invention may further include: Step S1100: Perform cross-node data integrity verification on the migrated meteorological file. The verification process includes: Exemplarily, cross-node data integrity verification is, for example, to ensure that the data is not damaged or tampered with during the transmission process by comparing the hash check codes of the source node and the target node files after the meteorological file is migrated from the source storage node to the target storage node. The hash check code is generated using a cryptographic hash function (such as SHA-256 or MD5). The source node calculates the hash value of the original file before migration (such as the SHA-256 result is "a1b2c3d4..."), and the target node performs the same hash calculation on the migrated copy after receiving it. The consistency verification result is generated by comparing whether the hash value strings of the two nodes are completely consistent. For example, if the hash value of the source node is "a1b2c3d4..." and the hash value of the target node is "a1b2c3d4...", the verification passes; if the hash value of the target node is "e5f6g7h8...", the verification fails. The verification process needs to be performed in an independent secure environment to avoid temporary files or network caches during the migration process interfering with the results. For example, after the 2023 typhoon wind speed file of city B is migrated from node A to node B, the system calls the SHA-256 algorithm to calculate the hash values ​​of the two node files respectively, and returns the consistency verification result as "success" or "failure".

[0108] Step S1200: extracting the hash check code of the original file in the source storage node and comparing it with the hash check code of the migration copy to generate a consistency verification result.

[0109] Exemplarily, the hash check code extraction of the source storage node and the target storage node is implemented through a file system interface or a dedicated verification tool. When the migration task is triggered, the source node generates a hash value of the original file and temporarily stores it in the log database. After the migration is completed, the target node calls the same hash function to generate a copy hash value. The comparison process is completed by character-by-character string matching. For example, when the hash value of the original file "a1b2c3d4" is exactly the same as the copy hash value "a1b2c3d4", the generated consistency verification result is "consistent"; if there is any character difference (such as "a1b2c3d4" and "a1b2c3d5"), the result is "inconsistent". The verification results are recorded in the migration task log and trigger subsequent index updates or exception retry processes. For example, when a temperature history data file is migrated from node C to node D, the hash check code comparison finds inconsistencies, the system marks the task as "failed", and triggers the exception handling mechanism.

[0110] Step S1300: When the consistency verification result passes, the storage location path of the migration copy and the spatial density grid association mark of the hierarchical index structure are synchronously updated, and the migration effective timestamp is marked in the index mapping relationship table.

[0111] Exemplarily, after the consistency verification is passed, the system needs to update the spatial density grid association tag in the hierarchical index structure to reflect the new storage location of the migrated copy. The spatial density grid association tag includes the geographic grid unit code (such as city B core grid A01) and the storage node path (such as / node_ssd_02 / typhoon_wind_202307.data). The index mapping relationship table updates the migration effective timestamp (such as 2023-07-15T14:30:00 UTC) through database transactions to ensure that query requests can be accurately routed to the target node. For example, after the typhoon wind speed file is migrated to node B, its association tag is updated from " / node_ssd_01 / typhoon_wind_202307.data" to " / node_ssd_02 / typhoon_wind_202307.data", and the effective timestamp is recorded in the mapping table. Subsequent retrieval requests directly access the data copy of node B.

[0112] Step S1400: When the consistency verification result fails, the abnormal retry mechanism of the migration task queue is triggered. The mechanism automatically increases the heat priority score of the failed migration file to a preset level and reinserts it to the head of the queue to wait for the next migration scheduling.

[0113] Exemplarily, the abnormal retry mechanism processes files that fail verification by modifying the priority and retry strategy of the migration task queue. The preset level of the heat priority score is increased by, for example, multiplying the original score by a coefficient (such as 1.5) or increasing a fixed value (such as +20). For example, the original score of a file from 35 is increased to 52.5 (35×1.5), so that its ranking in the queue rises. Reinserting the head of the queue ensures that the file is processed first in the next migration scheduling cycle. For example, a humidity history file fails the hash check due to a network interruption. The system increases its score from 10 to 15 and inserts it to the top of the queue. It will be tried to be transferred first during the next migration. The upper limit of the number of retries can be configured (such as up to 3 times). If the upper limit is exceeded, it will be marked as "permanent failure" and a manual intervention alarm will be triggered.

[0114] Step S1500: Input the storage node load state vector of the migration copy that finally passes the verification into the feedback learning module of the meteorological storage optimization model to generate a storage node load distribution feature vector for updating the third component compatibility parameter of the multidimensional training feature vector in the model training phase.

[0115] Exemplarily, the storage node load state vector includes the capacity utilization rate, input and output operation times and average response delay time of the target node after migration. For example, the remaining capacity of node B after migration is reduced from 40% to 35%, the input and output operation times are increased from 1500 times / second to 1800 times / second, and the average delay is increased from 25ms to 30ms. The feedback learning module generates a load distribution feature vector (such as [capacity change: -5%, operation times change: +300, delay change: +5ms]) by analyzing the impact of load changes on storage strategies, and adjusts the compatibility parameter of the third component of the multi-dimensional training feature vector. The compatibility parameter reflects the degree of adaptation between the meteorological element type and the storage node hardware. For example, the compatibility parameter of the high-bandwidth node and the real-time wind speed data is adjusted from 0.9 to 0.85 (performance degradation due to increased load), thereby reducing the priority of the node for high-heat elements in subsequent strategy generation.

[0116] As an implementation manner, the process of updating the storage node load distribution feature vector may include: Step S1501: According to the change range of the capacity utilization rate of the target node where the migration copy is located, the weight coefficient of the load balancing prediction submodule of the strategy generation layer in the meteorological storage optimization model is adjusted.

[0117] Exemplarily, the change in capacity utilization is obtained by calculating the difference in the remaining capacity of the target node before and after migration. For example, the remaining capacity of node B drops from 40% to 35%, with a change of -5%. The weight coefficient of the load balancing prediction submodule is used to balance the relationship between capacity, performance and heat priority. For example, the original weight coefficient is capacity: 0.4, performance: 0.4, heat: 0.2; if the capacity utilization rate deteriorates (such as three consecutive migrations resulting in a decrease in capacity), the capacity weight coefficient is adjusted to 0.5, the performance weight is reduced to 0.3, and the heat is maintained at 0.2. The adjustment process is optimized through the gradient descent algorithm, so that the model pays more attention to capacity constraints when predicting and avoids node overload. For example, a node has tight capacity due to repeated reception of high-heat wind speed data. The model increases the capacity weight coefficient and gives priority to nodes with higher remaining capacity when generating subsequent strategies.

[0118] Step S1502: based on the correlation between the migration effective timestamp and the access frequency of similar files in the historical storage log, the slope parameter of the dynamic weight factor attenuation curve of the main index node in the hierarchical index structure is recalculated.

[0119] Exemplarily, the correlation analysis between the migration effective timestamp and the historical access log is achieved through time series alignment. For example, after a typhoon wind speed file is migrated to node B on July 15, 2023, its access frequency drops from 100 times per day to 50 times per day between July 16 and July 20. The slope parameter of the dynamic weight factor decay curve reflects the decay rate of the element heat. The original slope parameter is -0.05 (daily decay 5%), which is adjusted to -0.06 (daily decay 6%) after recalculation to reduce the storage priority of inactive files more quickly. The slope parameter is calculated by linear regression fitting the downward trend of access frequency, for example, slope = Δ number of accesses / Δ time = (50-100) / (5 days) = -10 times / day, mapped to a dynamic weight factor decay rate of -0.06 / day.

[0120] Step S1503: input the updated dynamic weight factor attenuation curve into the core storage area coverage ratio calculation process of the spatial density grid, and re-divide the boundary threshold between the core storage area and the edge cache area of ​​the geographic grid unit.

[0121] Exemplarily, the demarcation threshold between the core storage area and the edge cache area is adjusted according to the slope of the dynamic weight factor attenuation curve. For example, the original threshold density value is 0.8. When the slope parameter changes from -0.05 to -0.06, the density value threshold is increased to 0.82 to more strictly screen high-heat grid units. The re-division process is analyzed through the density distribution histogram. For example, the density value of the core grid unit A01 of city B drops from 0.9 to 0.85 (due to the acceleration of dynamic weight decay). If the new threshold 0.82 is still higher than this value, it will be retained as the core storage area; if the density value of a grid unit drops from 0.81 to 0.79, it will be downgraded to the edge cache area. Threshold adjustment ensures that the core area only retains the current highly active meteorological data and optimizes the allocation of storage resources.

[0122] Step S1504: According to the re-divided demarcation threshold, trigger the horizontal level reconstruction operation of the hierarchical index structure, update the composite index tag of the meteorological file, and simultaneously optimize the number of redundant copies and the cross-node backup frequency in the storage node allocation strategy.

[0123] Exemplarily, the horizontal hierarchical reconstruction operation is completed by batch updating the storage area identifiers of the spatial density grid, for example, downgrading A08 and A09 in the original core area grids A01 to A10 to edge areas, and adding A11 as the core area. The composite index label is updated to "core storage area-wind speed-A11" or "edge cache area-temperature-A08". The number of redundant copies is adjusted according to the new demarcation threshold, for example, the number of core area copies is reduced from 3 to 2 (due to the reduction of the number of core areas after the density threshold is increased), and the number of edge area copies is increased from 1 to 2 (to improve disaster recovery capabilities). The cross-node backup frequency is optimized synchronously, for example, the core area backup frequency is adjusted from once per hour to once every 30 minutes, and the edge area is adjusted from once a day to once every 12 hours. The reconstruction operation ensures atomicity through database transactions to avoid query errors caused by inconsistent index status. For example, a typhoon file is downgraded to an edge area because grid A09 is downgraded to an edge area, and its number of copies is adjusted from 3 to 2, and an additional backup task from node B to node C is triggered.

[0124] Figure 2 A hardware entity diagram of a storage system provided by an embodiment of the present invention, such as Figure 2 As shown, the hardware entity of the storage system 1000 includes: a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can be run on the processor 1001, and the processor 1001 implements the steps in the method of any of the above embodiments when executing the program.

[0125] The memory 1002 stores computer programs that can be run on the processor. The memory 1002 is configured to store instructions and applications executable by the processor 1001. It can also cache data to be processed or processed by the processor 1001 and various modules in the storage system 1000 (for example, image data, audio data, voice communication data, and video communication data). This can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0126] When the processor 1001 executes the program, the steps of any of the above-mentioned meteorological metadata storage methods based on machine learning are implemented. The processor 1001 generally controls the overall operation of the storage system 1000.

[0127] An embodiment of the present invention provides a computer storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the meteorological metadata storage method based on machine learning in any of the above embodiments.

[0128] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0129] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0130] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0131] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.

[0132] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0133] The above description is only an implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A meteorological metadata storage method based on machine learning, characterized in that: The method comprises: Acquire a meteorological metadata set of a target area, wherein the meteorological metadata set includes a plurality of unstructured meteorological file files, each meteorological file file includes monitoring data of at least one meteorological element and corresponding spatiotemporal identification information; Input the meteorological metadata data set into a pre-trained temporal feature extraction model to generate temporal distribution features of each meteorological file, and synchronously input the meteorological metadata data set into a spatial topological mapping model to generate spatial correlation topological features of each meteorological element; Based on the cross-fusion results of the temporal distribution characteristics and the spatial correlation topological characteristics, a hierarchical index structure of meteorological metadata is constructed, wherein the hierarchical index structure includes a vertical level divided by meteorological element type and a horizontal level divided by time and space dimensions; Calling the meteorological storage optimization model, generating a storage node allocation strategy for the meteorological file file according to the hierarchical weight distribution of the hierarchical index structure, wherein the storage node allocation strategy includes a compression rate threshold and a storage location priority for each meteorological element data; According to the storage node allocation strategy, each meteorological file in the meteorological metadata set is written into the corresponding distributed storage node, and based on the real-time update mechanism of the hierarchical index, the physical location and index mapping relationship of the stored files are dynamically adjusted.

2. The method according to claim 1, characterized in that After obtaining the meteorological metadata set of the target area, the method further includes: Performing data preprocessing on the meteorological metadata set, the data preprocessing comprising: Identify the timestamp intervals of missing monitoring data in the meteorological file, and generate interpolated data based on the fluctuation trend of monitoring data from adjacent meteorological stations; Detect abnormal jump points in meteorological element monitoring data, and use sliding window algorithm to smooth and correct the data sequence in the window; Convert meteorological archive files of different formats into a unified binary encoding format and add metadata tags containing time and space identification information to the file header; The preprocessed meteorological metadata set meets the input format specifications of the temporal feature extraction model and the spatial topology mapping model.

3. The method according to claim 1, characterized in that The step of constructing a hierarchical index structure of meteorological metadata based on the cross-fusion result of the temporal distribution feature and the spatial correlation topological feature includes: Decomposing the time series distribution characteristics into periodic characteristics and mutation characteristics of meteorological elements, wherein the periodic characteristics represent the fluctuation pattern of meteorological elements that recurs in historical monitoring data, and the mutation characteristics represent the abnormal fluctuation range that exceeds a preset threshold; Decomposing the spatial correlation topological features into geographic correlation features and cross-regional conduction features, wherein the geographic correlation features represent the spatial continuity of the monitoring data of adjacent meteorological stations, and the cross-regional conduction features represent the correlation strength of meteorological elements between non-adjacent regions; In the vertical hierarchical division, a main index node is created according to the meteorological element type, and a dynamic weight factor is assigned to each main index node. The dynamic weight factor is determined by the ratio of the periodic feature to the mutation feature, so that the main index node of the high-frequency mutation element obtains a higher vertical hierarchical priority; In the horizontal hierarchical division, a spatial density grid is generated based on the superposition of geographical correlation characteristics and cross-regional conduction characteristics. Each grid unit records the statistical distribution characteristics of meteorological elements within the corresponding geographical range, and the grid unit is divided into a core storage area and an edge cache area according to the spatial density value. A bidirectional mapping rule between the main index node of the vertical level and the spatial density grid of the horizontal level is established. When the meteorological file matches both the dynamic weight factor threshold of the main index node and the core storage area condition of the spatial density grid, the cross-level joint index tag is triggered to generate a composite index tag containing spatiotemporal dimensions and feature attributes.

4. The method according to claim 3, characterized in that The calling of the meteorological storage optimization model to generate a storage node allocation strategy for the meteorological file according to the hierarchical weight distribution of the hierarchical index structure includes: Extract the dynamic weight factor of each main index node in the vertical hierarchy to generate a feature type heat weight sequence, where the heat weight is exponentially positively correlated with the dynamic weight factor; Extracting the boundary threshold between the core storage area and the edge cache area of ​​the spatial density grid in the horizontal level, generating a regional storage density gradient map, wherein the gradient map reflects the predicted values ​​of data access frequency of different geographic grid cells; Inputting the element type heat weight sequence into the first feature processing layer of the meteorological storage optimization model to generate an element priority queue, wherein the queue arranges the best storage node type of each meteorological element type in descending order of heat weight; Inputting the regional storage density gradient map into the second feature processing layer of the meteorological storage optimization model to generate a geographic storage strategy matrix, wherein each cell in the matrix records the storage position adjustment frequency of the corresponding grid within a preset time window; The element priority queue and the geographic storage strategy matrix are integrated to generate a storage node allocation strategy, in which meteorological file files of element types with heat weights greater than a preset value are allocated to target bandwidth storage nodes, and based on the grid cells whose adjustment frequency exceeds a threshold in the geographic storage strategy matrix, the number of redundant copies and the frequency of cross-node backup for associated files are increased.

5. The method according to claim 1, characterized in that The training process of the meteorological storage optimization model includes: Acquire historical meteorological metadata storage records, wherein the historical meteorological metadata storage records include hierarchical weight distribution samples of a hierarchical index structure, storage node allocation strategy labels, and node load fluctuation time series data; The hierarchical weight distribution sample is converted into a multidimensional training feature vector, wherein the first component of the multidimensional training feature vector corresponds to the dynamic weight factor change trend of the main index node in the vertical hierarchy, the second component corresponds to the core storage area coverage ratio of the spatial density grid in the horizontal hierarchy, and the third component corresponds to the compatibility parameter of the meteorological element type and the storage node hardware configuration; In the model initialization stage, the multi-dimensional training feature vector is input into the convolutional feature extraction layer of the meteorological storage optimization model, and the local correlation pattern of the hierarchical weight distribution is captured through a sliding window to generate an initial feature code; Input the initial feature code into the multi-head attention allocation layer of the meteorological storage optimization model, calculate the influence weights of different hierarchical dimensions on the storage node allocation strategy, and generate a global feature code with attention weights; Input the global feature code into the strategy generation layer of the meteorological storage optimization model, and generate a candidate storage node allocation strategy set in combination with the periodic law of the node load fluctuation time series data; In the model optimization stage, the candidate storage node allocation strategy set is compared with the storage node allocation strategy label in the storage record, the strategy offset loss value is calculated, and the network parameters of the convolutional feature extraction layer, the multi-head attention allocation layer and the strategy generation layer are adjusted through error back propagation until the allocation strategy output by the model matches the storage node allocation strategy label to a preset threshold.

6. The method according to claim 5, characterized in that The step of generating a candidate storage node allocation strategy set in combination with the periodic law of the node load fluctuation time series data includes: Input the global feature code and node load fluctuation time series data into the load balancing prediction submodule of the meteorological storage optimization model to predict the CPU occupancy rate curve and disk throughput change trend of each storage node in the future time window; Generate a storage node load pressure level identifier according to the predicted CPU occupancy rate curve, wherein the storage node load pressure level identifier includes a high-voltage node, a medium-voltage node, and a low-voltage node; Generate a storage node data transmission efficiency score based on the predicted disk throughput change trend, where the score is positively correlated with the slope value of the throughput increase trend; In the strategy generation layer, the load pressure level identifier and the data transmission efficiency score are jointly constrained and analyzed: for high-pressure nodes, the maximum capacity threshold for allocating new data is limited; for low-pressure nodes, the weight coefficient of receiving target priority data is increased; for nodes with data transmission efficiency scores lower than the preset standard, the cross-node replica backup mark is triggered; According to the results of the joint constraint analysis, a set of candidate storage node allocation strategies is generated. Each strategy in the set contains the load pressure compatibility parameters of the target node, data transmission efficiency guarantee parameters and redundant backup execution conditions, and the parameter scale consistency between different strategies is ensured through normalization processing of the strategy generation layer.

7. The method according to claim 1, characterized in that The process of dynamically adjusting the mapping relationship between the physical location and index of the stored files includes: Collect the capacity utilization rate and data access frequency of the distributed storage nodes in real time to generate a node load state vector, which includes the remaining capacity percentage of the distributed storage node, the number of input and output operations per unit time, and the average response delay time; The node load state vector is input into the real-time decision submodule of the meteorological storage optimization model, and the heat priority score of the file to be migrated is calculated in combination with the heat decay curve of the hierarchical weight distribution in the hierarchical index structure. The heat priority score is determined by the decay rate of the dynamic weight factor of the main index node of the vertical level to which the file belongs and the access frequency decline trend of the horizontal spatial density grid; Generate a migration task queue according to the heat priority score, arrange the files to be migrated in ascending order according to the score, and mark the hardware performance matching parameters of the target migration node; When performing incremental migration operations, the meteorological file with the lowest score in the queue and the continuous non-access time exceeding the preset threshold is preferentially migrated, and the spatial density grid association mark of the hierarchical index structure is synchronously updated after the migration copy is created in the target node; During the migration process, an atomic lock mechanism is implemented on the index mapping relationship between the source node and the target node, so that the query request always points to the valid storage location during the migration, and triggers the index consistency check after the migration is completed. When a cross-node index path conflict is detected, it rolls back to the pre-migration state and regenerates the migration task queue.

8. The method according to claim 1, characterized in that: The meteorological metadata retrieval process includes: Receiving a search request submitted by a user, parsing the time and space range constraints and the meteorological element type combination conditions in the search request, and generating a multi-dimensional search condition vector, wherein the vector includes a geographic boundary coordinate set, a time interval stamp, and an element type coding sequence; Input the multi-dimensional search condition vector into a joint query engine of a hierarchical index structure, match all primary index nodes whose dynamic weight factor is greater than a preset threshold in the vertical level, and simultaneously screen the geographic grid cells covered by the core storage area of ​​the spatial density grid in the horizontal level; Generate a candidate meteorological file set according to the intersection result of the matched main index node and the geographic grid unit, each file in the set is marked with the compression algorithm type and storage location path recorded in the storage node allocation strategy; Calling the real-time decompression interface of the distributed storage node, performing parallel decompression operations on the meteorological file files in the candidate set based on the compression algorithm type, and generating a standardized meteorological element monitoring data stream; The standardized data stream is input into the factor relevance scoring model established in the meteorological storage optimization model training phase, and the spatiotemporal coverage completeness, factor matching degree and historical access heat weight of each meteorological file and the retrieval condition vector are calculated to generate a final retrieval result sorting list and return it to the user end.

9. The method according to claim 8, characterized in that The calculation process of the factor relevance scoring model includes: Extract the spatiotemporal identification information of the meteorological file, calculate the overlap ratio of its geographical coverage and the geographical boundary coordinate set in the search request, and generate the first dimension coverage score; Analyze the number of data points in the meteorological element monitoring data stream that match the retrieval request element type coding sequence, and generate the second dimension element matching score by combining the numerical fluctuation range of the data points with the degree of deviation from the historical average value; Obtain the dynamic weight factor attenuation curve of the main index node in the hierarchical index structure, and generate the third dimension heat attenuation compensation coefficient by combining the access frequency decrease rate of the file in the historical storage log; The first dimension coverage score, the second dimension element matching score and the third dimension heat attenuation compensation coefficient are input into the weight allocation function defined in the meteorological storage optimization model training phase, and the function dynamically adjusts the weight ratio of each dimension score according to the business priority of each category in the element type coding sequence; The total score after weighted summation is normalized to generate a standardized relevance score value, and the candidate meteorological file files are sorted in descending order according to the score value to form a final search result sorting list.

10. A storage system comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps in the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Processing method and search method for multiple-source and multi-temporal satellite image tile data

    CN104899282A

  • Multi-dimensional space meteorological grid data distributed storage query method and system

    CN116126942A

  • Financial big data optimization storage method

    CN118363961A

  • Method and system for realizing self-adaptive object storage data life cycle management based on deep learning large model

    CN118820200A

  • Parallel synchronization method and system for unstructured files

    CN119513057A

Cited By

  • Enterprise financial document unified management system and method based on distributed storage

    CN120430878A

  • Density map data visual query method and device based on flow model

    CN120448430A

  • Framing DOM (Document Object Model) data edge matching method and device, equipment and storage medium

    CN120455679A

  • Emergency command rescue data distributed storage method and system based on artificial intelligence

    CN120492546A

  • Distributed storage method and system for emergency command and rescue data based on artificial intelligence

    CN120492546B