A Fiber Optic Data Storage Management System and Method Based on Big Data
Through a fiber optic data storage management system based on big data, using LSTM network and deep reinforcement learning algorithms, a spatio-time joint indexing and resource allocation model is built, which solves the problems of inefficiency and poor stability of the fiber optic storage system in complex access modes, and realizes efficient data scheduling and index optimization, improving the overall performance and stability of the system.
Patent Information
- Application Number
- CN202510559566.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-30
AI Technical Summary
When existing fiber optic storage systems face complex data access modes and dynamic load changes, there are problems of inefficiency and poor stability in resource management, data scheduling and storage optimization. Especially in large-scale concurrent access, uneven data hot and cold data, and sudden access hotspots, the existing indexing mechanism is difficult to adapt to complex query needs.
The fiber data storage management system based on big data is adopted to extract the spatiotemporal characteristics of historical access records through the LSTM network, calculate the data heat value with the time attenuation function, generate a thermal distribution map, and use deep reinforcement learning algorithm to build a storage resource allocation model, define bandwidth utilization, node load rate and predicted access as state space parameters, output optical path priority matrix and replica distribution topology instructions, build a spatiotemporal joint index of structured and unstructured data, integrate data access delay, index hit rate and replica migration frequency, obtain fiber storage efficiency index, and dynamically adjust data scheduling and indexing strategies.
It realizes efficient data scheduling and index optimization, improves the throughput and retrieval efficiency of the fiber storage system, ensures the long-term efficient and stable operation of the system, and adapts to complex data access modes and dynamic load changes.
Smart Images

Figure CN120085812B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of storage management, and particularly to a fiber optic data storage management system and method based on big data. Background Art
[0002] With the advent of the big data era, the demand for data storage has shown an explosive growth. The traditional storage architecture is difficult to meet the increasing high-throughput and low-latency access requirements. Fiber optic storage technology has been widely used in fields such as data centers, cloud computing, and edge computing due to its high bandwidth, low latency, and high reliability. However, in the face of complex data access patterns and dynamic load changes, there are still many challenges in resource management, data scheduling, and storage optimization in existing fiber optic storage systems. There is an urgent need for a more intelligent and dynamic management mechanism to improve storage efficiency and optimize data access performance.
[0003] Currently, fiber optic storage management mainly relies on traditional rule-based scheduling strategies, such as static replica allocation, fixed bandwidth allocation, and linear index management. These methods perform well when the data access pattern changes little, but in the face of large-scale concurrent access, uneven data heat and cold, sudden access hotspots, etc., they are prone to cause a decline in resource utilization and low storage efficiency. In addition, the existing index mechanism is usually based on static data structures such as B+ trees and hash indexes, which are difficult to adapt to complex query requirements, resulting in a decline in query efficiency. At the same time, the lack of an intelligent storage efficiency evaluation mechanism makes storage optimization rely on manual adjustment, with a slow response speed and unable to meet the requirements of efficient management. Therefore, there is an urgent need for a fiber optic storage management system based on big data analysis and intelligent optimization to improve data scheduling capabilities, optimize the index structure, and dynamically evaluate storage efficiency to ensure the high efficiency and stability of the system. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a fiber optic data storage management system and method based on big data, which solves the problems in the above background art.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A fiber optic data storage management system based on big data, comprising the following modules: a data access analysis module, a data scheduling module, an index optimization module, and a storage efficiency evaluation module; the data access analysis module is used to extract the spatio-temporal features of the historical access records of the fiber optic storage system through an LSTM network, calculate the data heat value in combination with a time decay function, and generate a heat distribution map identifying high-frequency and low-frequency data blocks; the data scheduling module is used to construct a storage resource allocation model through a deep reinforcement learning algorithm according to the heat distribution map, define the bandwidth utilization rate, node load rate, and predicted access volume as state space parameters, output an optical path priority matrix and a replica distribution topology instruction, and monitor the data flow fluctuation through a sliding window mechanism to trigger the migration of high-frequency data to low-load nodes; the index optimization module is used to construct a spatio-temporal joint index for structured data and a vector similarity index for unstructured data based on the heat distribution map and the replica distribution topology instruction, and optimize the prefetch strategy according to the data heat value to generate a high-frequency data cache index distribution map; the storage efficiency evaluation module is used to comprehensively consider the data access latency, index hit rate, and replica migration frequency to obtain a fiber optic storage efficiency index, and when the fiber optic storage efficiency index is lower than a preset threshold, generate a structure optimization instruction set and feedback it to the data scheduling module and the index optimization module.
[0006] Further, the specific process of extracting the spatio-temporal features of the historical access records of the fiber optic storage system through an LSTM network and calculating the data heat value in combination with a time decay function is as follows: perform spatio-temporal encoding on the historical access records, extract timestamps, geographical coordinates, and the number of concurrent requests, and construct a spatio-temporal feature vector; model the time dimension features through a bidirectional LSTM network to capture the periodic pattern of data access, and at the same time extract the spatial dimension features through a convolutional network; input the spatio-temporal features into a spatio-temporal attention mechanism to dynamically allocate time decay weights and generate a spatio-temporal correlation weight matrix; input the spatio-temporal correlation weight matrix into the time decay function to dynamically decay the influence of historical data according to the time distance, and perform normalization processing on the decayed data to generate a globally comparable data heat value.
[0007] Further, the specific process of generating a heat distribution map identifying high-frequency and low-frequency data blocks is as follows: perform geographical grid division on the storage space based on the data heat value, and map the access location to a hierarchical spatial grid through geographical hash encoding; identify high-frequency access grid clusters through a density clustering algorithm, and merge adjacent high-density grids to form hot blocks; separate low-frequency cold data regions through a spatio-temporal joint clustering algorithm, and dynamically adjust the block boundaries in combination with the stability of the access interval; mark the high-frequency blocks as red hot spots and the low-frequency blocks as blue cold areas to generate a visual heat distribution map.
[0008] Furthermore, according to the heat distribution map, the specific process of constructing a storage resource allocation model through a deep reinforcement learning algorithm is as follows: construct a hierarchical state space, and encode the block heat value, fiber optic network topology structure, and load data in the heat distribution map into a multi-dimensional state vector; design a multi-objective reward function, with minimizing transmission delay, balancing node load, and maximizing throughput as the joint optimization objectives, and increase the delay penalty weight of high-frequency blocks; train the model through the deep deterministic policy gradient algorithm, the action decision-making network generates a resource allocation strategy, and constructs a storage resource allocation model.
[0009] Furthermore, define the bandwidth utilization rate, node load rate, and predicted access volume as state space parameters, and the specific process of outputting the optical path priority matrix and replica distribution topology instructions is as follows: jointly encode the bandwidth utilization rate, node load rate, and predicted access volume to construct a dynamically perceived state space matrix; calculate the fiber optic path priority through the optical path game model, dynamically adjust the path weight in combination with the real-time load, and generate the optical path priority matrix; design a replica distribution algorithm with spatio-temporal constraints, generate replica topology instructions in combination with the heat value distribution and node geographical distance, and dynamically correct the instruction parameters through a sliding window verification mechanism.
[0010] Furthermore, based on the heat distribution map and replica distribution topology instructions, the specific process of constructing a spatio-temporal joint index for structured data and a vector similarity index for unstructured data is as follows: parse the high-frequency block coordinates and heat values in the heat map, and generate spatio-temporal key values in combination with the replica topology instructions; construct a multi-level B+ tree index based on the spatio-temporal key values, extract the semantic features of unstructured data, generate a weighted vector space in combination with the heat value, construct a vector index through the distributed HNSW algorithm, and preferentially store high-heat data to low-latency nodes; dynamically associate the spatio-temporal joint index and the vector similarity index to construct a cross-modal retrieval path.
[0011] Furthermore, optimize the prefetch policy according to the data heat value, and the specific process of generating a high-frequency data cache index distribution map is as follows: divide the cache levels according to the data heat value, and define the rules for preloading high-frequency data into memory; design a dynamic prefetch window to predict future hot data and generate a prefetch task queue; dynamically adjust the cache weight through a spatio-temporal attention mechanism, and map the cache location and validity period into a visual distribution map.
[0012] Furthermore, comprehensively consider the data access delay, index hit rate, and replica migration frequency to obtain the fiber optic storage efficiency index. The specific process is as follows: comprehensively adjust the weight ratio of access delay, index hit rate, and replica migration frequency through the analytic hierarchy process; introduce a time decay factor to perform exponential smoothing processing on historical performance data, and generate the fiber optic storage efficiency index through linear weighted fusion.
[0013] Further, when the fiber optic storage efficiency index is lower than the preset threshold, the specific process of generating a structure optimization instruction set and feeding it back to the data scheduling module and the index optimization module is as follows: Establish a multi-objective decision tree, and select optimization strategies according to the branches of the fiber optic storage efficiency index, including: too high latency: trigger optical path priority reset instructions and edge cache expansion instructions; too low hit rate: generate index level migration instructions and prefetch policy update instructions; too frequent migration: start replica distribution topology reconstruction instructions and load balancing policies.
[0014] A fiber optic data storage management method based on big data includes the following steps: S1. Extract the spatio-temporal features of the historical access records of the fiber optic storage system through an LSTM network, calculate the data heat value in combination with a time decay function, and generate a heat distribution map identifying high-frequency and low-frequency data blocks; S2. According to the heat distribution map, construct a storage resource allocation model through a deep reinforcement learning algorithm, define bandwidth utilization rate, node load rate, and predicted access volume as state space parameters, output an optical path priority matrix and replica distribution topology instructions, and monitor the data flow fluctuation through a sliding window mechanism to trigger the migration of high-frequency data to low-load nodes; S3. Based on the heat distribution map and replica distribution topology instructions, construct a spatio-temporal joint index for structured data and a vector similarity index for unstructured data, and optimize the prefetch policy according to the data heat value to generate a high-frequency data cache index distribution map; S4. Synthesize the data access latency, index hit rate, and replica migration frequency to obtain the fiber optic storage efficiency index. When the fiber optic storage efficiency index is lower than the preset threshold, generate a structure optimization instruction set and feed it back to the data scheduling module and the index optimization module.
[0015] The present invention has the following beneficial effects:
[0016] (1) A fiber optic data storage management system based on big data models the historical access records through the LSTM network of the data access analysis module, calculates the data heat value in combination with a time decay function, and generates a heat distribution map, thereby accurately identifying high-frequency access data blocks and low-frequency cold data regions, providing an optimization basis for data scheduling. The data scheduling module uses a deep reinforcement learning algorithm to construct a storage resource allocation model, combines state space parameters such as bandwidth utilization rate, node load rate, and predicted access volume, and dynamically adjusts the optical path priority and replica allocation strategy to achieve adaptive optimization of the storage path and improve the system throughput capacity.
[0017] (2) A fiber optic data storage management method based on big data. By constructing a spatio-temporal joint index for structured data and a vector similarity index for unstructured data, and optimizing the prefetching strategy in combination with data popularity, it realizes fast indexing and cache optimization of high-frequency data, improving the retrieval efficiency. Comprehensively analyzing data access latency, index hit rate, and replica migration frequency, calculating the fiber optic storage efficiency index, and when the index is lower than the preset threshold, automatically generating a structure optimization instruction set to dynamically adjust data scheduling and indexing strategies to ensure the long-term efficient and stable operation of the storage system.
[0018] Of course, it is not necessary for any product implementing the present invention to simultaneously achieve all the above-mentioned advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flowchart of a fiber optic data storage management system based on big data according to the present invention.
[0020] Figure 2 It is a flowchart of a fiber optic data storage management method based on big data according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The embodiments of the present application solve the problems of low data access efficiency, uneven storage resource allocation, insufficient index retrieval performance, and poor system stability in existing fiber optic storage systems through a fiber optic data storage management system and method based on big data.
[0022] The general idea of the solution in the embodiments of the present application is as follows:
[0023] Extract the spatio-temporal features of the historical access records of the fiber optic storage system through the LSTM network, calculate the data popularity value in combination with the time decay function, and generate a heat distribution map indicating high-frequency and low-frequency data blocks.
[0024] According to the heat distribution map, construct a storage resource allocation model through the deep reinforcement learning algorithm, define the bandwidth utilization rate, node load rate, and predicted access volume as state space parameters, output the optical path priority matrix and replica distribution topology instructions, and monitor the data flow fluctuations through the sliding window mechanism to trigger the migration of high-frequency data to low-load nodes.
[0025] Based on the heat distribution map and replica distribution topology instructions, construct a spatio-temporal joint index for structured data and a vector similarity index for unstructured data, and optimize the prefetching strategy according to the data popularity value to generate a cache index distribution map of high-frequency data.
[0026] Comprehensively considering data access latency, index hit rate, and replica migration frequency, obtain the fiber optic storage efficiency index. When the fiber optic storage efficiency index is lower than the preset threshold, generate a structure optimization instruction set and feedback it to the data scheduling module and the index optimization module.
[0027] Please refer to Figure 1 , an embodiment of the present invention provides a technical solution: a big data-based optical fiber data storage management system, including the following modules: a data access analysis module, a data scheduling module, an index optimization module, and a storage efficiency evaluation module; the data access analysis module is used to extract the spatio-temporal characteristics of the historical access records of the optical fiber storage system through an LSTM network, calculate the data heat value in combination with a time decay function, and generate a heat distribution map identifying high-frequency and low-frequency data blocks; the data scheduling module is used to construct a storage resource allocation model through a deep reinforcement learning algorithm according to the heat distribution map, define the bandwidth utilization rate, node load rate, and predicted access volume as state space parameters, output an optical path priority matrix and a replica distribution topology instruction, and monitor the data flow fluctuation through a sliding window mechanism to trigger the migration of high-frequency data to low-load nodes; the index optimization module is used to construct a spatio-temporal joint index for structured data and a vector similarity index for unstructured data based on the heat distribution map and the replica distribution topology instruction, and optimize the prefetch policy according to the data heat value to generate a high-frequency data cache index distribution map; the storage efficiency evaluation module is used to comprehensively consider the data access delay, index hit rate, and replica migration frequency to obtain the optical fiber storage efficiency index. When the optical fiber storage efficiency index is lower than a preset threshold, a structure optimization instruction set is generated and fed back to the data scheduling module and the index optimization module.
[0028] In this implementation plan, the data access analysis module: By analyzing historical access data, it identifies the hot regions with high-frequency access and the cold data regions with low-frequency access in the system, providing data support for resource scheduling and storage optimization. LSTM network: A type of recurrent neural network that is good at capturing the long-term dependencies of time series data and is used to extract the time periodicity features of data access (such as daily traffic peaks). Spatiotemporal features: Comprehensive features that combine the time dimension (such as access timestamps) and the space dimension (such as access geographical coordinates). Time decay function: A mathematical function that dynamically adjusts the weights of historical data, with higher weights for recent data and lower weights for long-term data, reflecting the timeliness of data heat. Heat distribution map: A visualization of the data heat distribution map, which uses color gradients (such as red for high frequency and blue for low frequency) to identify the access heat of different blocks. Spatiotemporal feature fusion modeling: For the first time, it combines the time decay function with the LSTM network to dynamically fuse spatiotemporal features, solving the problem of the separation of time and space analysis in traditional methods. Nonlinear decay mechanism: Design a function with an adaptive decay rate to flexibly handle sudden traffic scenarios (such as a sudden increase in data heat caused by an emergency). Data scheduling module: Dynamically allocates fiber bandwidth and storage node resources to ensure that high-heat data is preferentially transmitted and stored for disaster recovery. Deep reinforcement learning algorithm: A machine learning method that automatically optimizes decision-making strategies through simulating environmental interactions and is used to build a resource allocation model. State space parameters: Input variables on which the model's decisions depend, including bandwidth utilization rate (the proportion of fiber bandwidth used), node load rate (the occupancy rate of storage node resources), and predicted access volume (prediction of future data request volume). Optical path priority matrix: Defines the transmission priority rules of fiber paths (such as paths with low load have high weights) and dynamically optimizes the data transmission path. Replica distribution topology instruction: Specifies the rules for the storage location distribution of data replicas (such as storing across geographical regions) to ensure data disaster tolerance and access efficiency. Sliding window mechanism: A real-time monitoring technology that triggers resource adjustment strategies through the analysis of data flow fluctuations within a fixed time window (such as 5 minutes). Multi-objective joint optimization: Simultaneously optimizes delay, throughput, and load balancing through a reinforcement learning model, breaking through the limitations of traditional single-objective scheduling strategies. Dynamic disaster recovery rule: Generates replica distribution instructions by combining geographical distance and network topology to improve the data survival rate under regional disasters (such as earthquakes and power outages). Index optimization module: Improves data retrieval efficiency and supports efficient queries for structured data (such as logs) and unstructured data (such as images, gene sequences). Structured data: Data with a fixed format (such as database tables, time series logs). Unstructured data: Data without a fixed format (such as text, images, videos). Spatiotemporal joint index: An index structure that combines the time and space dimensions and supports joint queries of "time range + geographical region". Vector similarity index: Generates vectors by extracting data features and constructs an index to support similarity retrieval.Prefetch Strategy: Predict future access demands based on data popularity, and load high-frequency data into the cache (such as memory, SSD) in advance. Cross-modal Index Fusion: Unify the retrieval frameworks of structured and unstructured data for the first time, supporting "time-space-semantics" compound queries. Heat-driven Dynamic Prefetch: Dynamically adjust prefetch rules based on heat values, reduce redundant cache occupancy, and improve cache hit rate. Storage Efficiency Evaluation Module: Monitor system performance in real time, trigger adaptive optimization strategies, and ensure the long-term efficient operation of the system. Fiber Storage Efficiency Index (FSEI): A quantitative performance index generated by integrating access latency (data retrieval time), index hit rate (cache hit ratio), and replica migration frequency (data replica scheduling times). Structure Optimization Instruction Set: A set of executable instructions including index migration, replica expansion, path reset, etc., used to dynamically optimize system configuration. Multi-objective Decision Tree: A logical model that matches optimization strategies based on the reasons for the decline of performance indices (such as excessive latency, low hit rate). Closed-loop Feedback Mechanism: Feed back the execution effects of optimization instructions to the evaluation module in real time to form a closed loop of self-iterative optimization. Dynamic Weighted Evaluation Model: Dynamically adjust the index weights according to the real-time load (such as focusing on latency optimization under high load), breaking through the rigidity of the fixed weight model.
[0029] Specifically, the specific process of calculating the data heat value by combining the spatio-temporal features of the historical access records of the fiber storage system extracted by the LSTM network and the time decay function is as follows: Perform spatio-temporal encoding on the historical access records, extract timestamps, geographical coordinates, and concurrent request volumes, and construct a spatio-temporal feature vector; Model the temporal dimension features through a bidirectional LSTM network to capture the periodic patterns of data access, and at the same time extract the spatial dimension features through a convolutional network; Input the spatio-temporal features into the spatio-temporal attention mechanism to dynamically allocate time decay weights and generate a spatio-temporal correlation weight matrix; Input the spatio-temporal correlation weight matrix into the time decay function, dynamically decay the influence of historical data according to the time distance, and normalize the decayed data to generate a globally comparable data heat value.
[0030] In this implementation, the historical access records of the optical fiber storage system are extracted through the LSTM network to obtain the spatio-temporal characteristics of the data, and the time decay function is combined to calculate the popularity value of the data, thereby generating the heat distribution of data access. The specific process is as follows. Spatio-temporal feature encoding: The historical access records contain information of time and space. The main encoding contents include: Time information: The timestamp of data access. Space information: The storage location where the data is accessed (such as the geographical coordinates of the storage node or the location of the data center). Access features: The concurrent request volume and the access frequency. These information are combined to form a spatio-temporal feature vector. Time dimension feature extraction (bidirectional LSTM), Data access usually has periodic patterns (such as more frequent data access during peak hours), so the bidirectional LSTM is used to model the time dimension features to capture the long-term dependencies and access patterns. Through the memory gating mechanism, LSTM can effectively extract the evolution trend of time series data and learn the influence of different time points on data access behavior. Space dimension feature extraction, Since data access may show an aggregation effect in space (such as some storage nodes being accessed more frequently than others), the convolutional network (CNN) is used to extract the spatial features to identify the access hotspots. Through the convolution operation, CNN can extract the access patterns of different data blocks in the storage system, identify local hotspots, and enhance the expression ability of spatio-temporal feature learning. Spatio-temporal attention mechanism to calculate the spatio-temporal correlation weight matrix, In order to dynamically measure the influence of time and space factors on data access patterns, this method uses the spatio-temporal attention mechanism to calculate the weights of data access. This mechanism can capture the influence degree of data access at different time points and different spatial regions and assign corresponding weights. The spatio-temporal attention mechanism forms a spatio-temporal correlation weight matrix by calculating the correlation between data points: ; where: represents the historical data point the influence degree of the access behavior on the data point . This weight matrix is used to measure the spatio-temporal correlation between data at different time points and storage locations. In order to ensure that the data access records in the relatively recent time have a greater impact on the popularity value, the time decay function is used to weight and decay the historical data. The time decay weight is calculated as follows: ; where: represents the data point the time decay weight. is the decay coefficient, which determines the influence degree of the data access history. is the current time, is the timestamp of the data point a. When is closer to the current time , is larger; on the contrary, the farther the time is, the more obvious the decay is. The data popularity value is calculated as follows: ; where: represents the location The data heat value at Calculated by the spatio-temporal attention mechanism. Calculated by the time decay function. Is the access request volume. The calculated heat value is further normalized for global comparison and visual analysis. Based on the calculated data heat value, a heat distribution map of data access is constructed to visually display the hot data areas of the storage system and provide a basis for subsequent data scheduling and index optimization.
[0031] Specifically, the specific process of generating a heat distribution map that identifies high-frequency and low-frequency data blocks is as follows: Geographically grid the storage space based on the data heat value, and map the access location to a hierarchical spatial grid through Geohash encoding; Identify high-frequency access grid clusters through the density clustering algorithm, and merge adjacent high-density grids to form hot blocks; Separate low-frequency cold data areas through the spatio-temporal joint clustering algorithm, and dynamically adjust the block boundaries in combination with the stability of the access interval; Mark high-frequency blocks as red hot spots and low-frequency blocks as blue cold areas to generate a visual heat distribution map.
[0032] In this implementation plan, the access pattern of the storage system is analyzed through the data heat value, and technologies such as spatial gridification, density clustering, high- and low-frequency block identification, and spatio-temporal joint clustering are used to generate a heat distribution map to realize the visualization of data access hot spots. The specific steps are as follows: Geographical grid division of the storage space. In order to accurately identify the spatial distribution of data access, Geohash encoding is used for hierarchical grid division of the storage space. Geohash encoding (Geohash encoding): Divide the storage space into L-level grids, and the size of each level of grid is determined by the length of the Geohash code. Higher-level Geohash codes represent larger areas, and lower-level Geohash codes represent finer areas. Through the Geohash code, each access location (x, y) is mapped to a grid cell for subsequent clustering analysis. Identify high-frequency access blocks (density clustering). Based on the grid division, the density clustering algorithm (DBSCAN) is used to identify high-frequency access blocks. Calculate the access density of each grid: ; where: Represents the network Access density. Is the grid Inside the The heat value of the Is the number of data points in the grid. Set the density threshold , when the network density Is greater than When this occurs, the network is marked as a high-frequency grid. The DBSCAN algorithm is used to aggregate adjacent high-frequency grids to form a hotspot area. Identify low-frequency cold data blocks (spatiotemporal joint clustering). To distinguish low-frequency data blocks, a spatiotemporal joint clustering algorithm is used to identify cold data areas: Calculate the access interval stability of the grid: ; where: represents the access interval stability of the grid . is the number of data accesses within this grid. represents the th timestamp of access. is the average access time of this grid. Set the stability threshold . When , it indicates that the access frequency of this grid is low and the time distribution is unstable, and it is determined as a low-frequency cold data block. Combining with the density clustering algorithm, low-frequency cold data grids are aggregated to form a cold data area.
[0033] Specifically, according to the heat distribution map, the specific process of constructing a storage resource allocation model through the deep reinforcement learning algorithm is as follows: Construct a hierarchical state space, and encode the block heat value, fiber optic network topology structure, and load data in the heat distribution map into a multi-dimensional state vector; Design a multi-objective reward function, with minimizing transmission delay, balancing node load, and maximizing throughput as the joint optimization objectives, and increasing the delay penalty weight of high-frequency blocks; Train the model through the deep deterministic policy gradient algorithm, and the action decision-making network generates a resource allocation strategy to construct a storage resource allocation model.
[0034] In this implementation plan, a hierarchical state space is constructed. The state space of the storage resource allocation model consists of multiple factors, including data heat value, fiber optic network topology structure, node load data, etc. This information needs to be encoded as a multi-dimensional vector to represent the state of the current system. The specific steps are as follows: Data heat value: Obtain the heat value of each block from the heat distribution map, which is used to measure the high or low data access frequency. A high heat value corresponds to a high-frequency data block, and a low heat value corresponds to a low-frequency data block. Fiber optic network topology structure: Encode the connection situation between each node and optical path in the network to reflect the structure of the fiber optic network. The network topology structure helps the model understand the communication delay and bandwidth resources between nodes. Node load data: The load situation of each node (such as storage load, processing capacity, etc.) affects the priority of nodes during data allocation. Integrate these factors to form a multi-dimensional state vector , where: represents the heat value distribution vector at the current time point t. represents the fiber optic network topology structure. Represents the load data of each current node. A multi-objective reward function is designed. To balance the optimization performance of storage resource allocation among different objectives, a multi-objective reward function is designed. Its goal is to jointly minimize the transmission delay, balance the node load, and maximize the throughput. Specifically as follows: Minimize the transmission delay: Minimize the delay during data transmission as much as possible. The goal is to allocate the frequently accessed data to the nodes with minimized delay to improve the system's response speed. Balance the node load: Ensure the load balance of each node and avoid some nodes being overloaded while other nodes are idle. Load balance can improve the overall processing capacity of the system. Maximize the throughput: Improve the overall throughput of the network and increase the amount of data that can be processed. During this process, a higher penalty weight is given to the delay of high-frequency blocks, that is, increase the optimization intensity of high-frequency data to ensure its minimum transmission delay. Multi-objective reward function Is expressed as: Where: Is a metric of the transmission delay in the current state. Is a metric of the node load balance. Is the current network throughput. Is the weight coefficient, used to control the relative importance of different objectives, and . Through the Deep Deterministic Policy Gradient algorithm (DDPG), a deep reinforcement learning model is trained to optimize the storage resource allocation. This algorithm is used for optimization problems in continuous action spaces and is suitable for tasks such as storage resource allocation that require choosing specific resource allocation strategies. Actor network: Generates the storage resource allocation strategy. By continuously updating the strategy to maximize the long-term reward, the optimal resource allocation scheme is selected. The output of the Actor network is an allocation strategy π(S) in a continuous action space, that is, how much resources should be allocated to each node. Critic network: Evaluates the value function of the current state-action pair ( , ). Through this evaluation, the Critic network can give the predicted value of the long-term reward of the current action and assist the Actor network in optimization. In each training cycle, the Actor network generates an action based on the current state = π( ), and then the Critic network calculates the Q value of this action, representing the expectation of the long-term reward. Update formula: Update of the Actor network: By the gradient ascent method, maximize the Q value given by the Critic network. ; Update of the Critic network: Optimize the Q value by minimizing the Bellman error. Where: , Is the learning rate for the Actor network and the Critic network. is the reward value at the current time step. γ is the discount factor used to calculate the long-term return. After training through deep reinforcement learning, the storage resource allocation model can dynamically adjust the resource allocation strategy according to the real-time state. Through continuous exploration and optimization, the model can find the optimal solution in multi-objective optimization, and intelligently allocate storage and bandwidth resources, thereby improving the performance and efficiency of the entire optical fiber storage system.
[0035] Specifically, defining the bandwidth utilization rate, node load rate, and predicted access volume as state space parameters, the specific process of outputting the optical path priority matrix and the replica distribution topology instruction is as follows: jointly encode the bandwidth utilization rate, node load rate, and predicted access volume to construct a dynamically perceived state space matrix; calculate the optical fiber path priority through the optical path game model, dynamically adjust the path weights in combination with the real-time load, and generate the optical path priority matrix; design a replica distribution algorithm with spatio-temporal constraints, generate the replica topology instruction in combination with the heat value distribution and the geographical distance between nodes, and dynamically correct the instruction parameters through a sliding window verification mechanism.
[0036] In this implementation plan, first, the state space matrix is constructed by jointly encoding the bandwidth utilization rate, node load rate, and predicted access volume. These state parameters reflect the current resource usage of the storage system and can dynamically perceive and adapt to changes in the system state. Bandwidth utilization rate: Reflects the degree of bandwidth usage on the optical fiber path. The bandwidth usage directly affects the priority of optical path selection. The higher the bandwidth utilization rate, the higher the likelihood that the path is in a high-load state and its priority may need to be reduced. Node load rate: Indicates the current processing load of each storage node. Nodes with a high load rate may need to transfer a portion of the traffic to other nodes to prevent overload. Predicted access volume: Based on historical data or model predictions of future access demands. The predicted access volume can help the system schedule and allocate resources in advance to cope with future access peaks. Generation of the optical path priority matrix. In this model, the design of the optical path priority depends on the optical path game model. This model dynamically adjusts the real-time data of the bandwidth utilization rate and the node load in the middle to generate the optical path priority matrix, determines the weight of each optical fiber path, and ensures efficient and balanced path selection. The key point of the optical path game model is to analyze the competition relationship between different paths through game theory and calculate the priority of each optical path. This priority is not only based on the bandwidth usage but also needs to be dynamically adjusted according to the load and transmission delay on each optical path. The calculation formula for the optical path priority matrix P is: ; where: is a function of the bandwidth utilization rate, describing the impact of the bandwidth utilization rate on the priority. is a function of the node load rate, describing the impact of the load on path selection. It is a function for predicting the access volume, which describes the impact of the access volume on the path. is the weight coefficient, which is used to adjust the influence degree of bandwidth, load, and access volume on the priority. This formula dynamically adjusts the path weight by comprehensively considering bandwidth, load, and access requirements, and then optimizes the priority of the optical path. The goal of replica distribution is to reasonably select the replica storage location according to the data popularity value and the geographical distance between nodes, so as to improve the storage efficiency of the system and reduce the access latency. The spatio-temporal constrained replica distribution algorithm combines the popularity value distribution with the geographical distance information of nodes and generates replica distribution topology instructions through the following steps: Popularity value distribution: High-popularity data should be as close as possible to the frequently accessed nodes to reduce the access latency. Node geographical distance: According to the geographical distribution of nodes, replicas are preferentially stored in nodes closer to users or high-access-frequency regions, thereby improving the efficiency of data access. To ensure the effectiveness and adaptability of the replica distribution instructions, a sliding window verification mechanism is adopted. This mechanism can dynamically adjust the replica storage scheme according to the actual access pattern, popularity change, and network topology. The core of sliding window verification is to verify the replica distribution instructions in batches, and the replica strategy within each window will be dynamically corrected according to the actual data access situation. By regularly adjusting the replica storage scheme, it is ensured that data can be efficiently stored and quickly respond to access requests.
[0037] Specifically, based on the heat distribution map and the replica distribution topology instructions, the specific process of constructing the spatio-temporal joint index for structured data and the vector similarity index for unstructured data is as follows: Parse the high-frequency block coordinates and popularity values in the heat map, and generate spatio-temporal key values in combination with the replica topology instructions; Based on the spatio-temporal key values, construct a multi-level B+ tree index, extract the semantic features of unstructured data, generate a weighted vector space in combination with the popularity value, construct a vector index through the distributed HNSW algorithm, and preferentially store high-popularity data to low-latency nodes; Dynamically associate the spatio-temporal joint index with the vector similarity index to construct a cross-modal retrieval path.
[0038] In this implementation, the thermal map and replica topology instructions are parsed to generate spatio-temporal key values. The system identifies the blocks with higher access frequencies by parsing the high-frequency block coordinates and heat values in the thermal map. These high-frequency blocks represent the most active areas in the system with high data access demands. By combining the replica distribution topology instructions, the storage locations and access paths of the data can be determined. This information will be used to generate spatio-temporal key values, which serve as the basis for subsequent index construction. The spatio-temporal key value is a combined identifier consisting of the geographical coordinates of the block (representing the storage location) and the heat value (representing the access frequency of the data), used to distinguish and index different data blocks. In this process, the spatio-temporal key value not only contains the geographical location of the data but also reflects the heat of the data, thus improving the data retrieval efficiency. Construct a multi-level B+ tree index. After generating the spatio-temporal key values, use these spatio-temporal key values to construct a multi-level B+ tree index. The B+ tree is an index structure widely used in databases and file systems, which has high query efficiency and good range query characteristics. By constructing a B+ tree index, specific spatio-temporal blocks can be quickly located, improving the access efficiency of structured data. Each node of the B+ tree contains spatio-temporal key value information, and the child nodes of each node point to data blocks with similar spatio-temporal key values, thus supporting efficient query operations, especially when data needs to be retrieved in spatio-temporal order. Extract the semantic features of unstructured data and generate a weighted vector space. For unstructured data such as text, pictures, or videos, semantic feature extraction is first required. This process extracts the semantic hierarchical features of the data through natural language processing (NLP) or computer vision (CV) techniques to form a vector representation that can express the data content. After obtaining the semantic feature vectors, the system combines the heat values and assigns weights to each vector. The heat value, as an important weight parameter, reflects the access frequency of the data. Combining the heat value with the semantic feature vectors forms a weighted vector space, thus enhancing the weight of high-heat data in the index. This means that high-frequency accessed data will be given priority and stored in nodes with higher performance and lower latency to improve the overall performance of the system. Construct a vector similarity index. Using the weighted vector space, the system constructs a vector similarity index through the distributed HNSW algorithm (Hierarchical Navigable Small World). HNSW is a graph-based approximate nearest neighbor search algorithm that can effectively handle high-dimensional vectors in large-scale datasets. Through the HNSW algorithm, fast similarity search for unstructured data can be achieved to find high-heat data similar to the query vector. The HNSW algorithm constructs a hierarchical graph structure, enabling the query to find data blocks with higher similarity in a shorter time. This method is efficient and scalable, especially suitable for systems that need to process massive data. Dynamically associate the spatio-temporal joint index with the vector similarity index. After constructing the spatio-temporal joint index and the vector similarity index, the system will perform dynamic association.The spatio-temporal joint index mainly processes structured data, indexing and retrieving data through spatio-temporal key-value pairs. The vector similarity index, on the other hand, processes unstructured data and provides a retrieval path based on semantic similarity. By dynamically associating these two indexes, the system can provide users with a cross-modal retrieval path, that is, support retrieval requests based on both spatio-temporal features and semantic similarity. Users can, through a single query request, locate the storage location of the data in the spatio-temporal joint index and then find the most relevant unstructured data through the vector similarity index, thus achieving efficient cross-modal data retrieval.
[0039] Specifically, the specific process of optimizing the prefetch strategy according to the data heat value and generating the high-frequency data cache index distribution map is as follows: divide the cache levels according to the data heat value and define the rules for preloading high-frequency data into memory; design a dynamic prefetch window to predict future hot data and generate a prefetch task queue; dynamically adjust the cache weights through the spatio-temporal attention mechanism and map the cache location and validity period into a visual distribution map.
[0040] In this implementation, the cache levels are divided according to the data heat value. The system classifies and hierarchically divides the data through the data heat value. The data heat value is calculated based on the data access frequency or access pattern, reflecting the activity level of the data. The part with a higher data heat value represents the hot data with frequent access, and the part with a lower data heat value is the cold data or the data with less access. According to the high and low heat values, the system divides the data into different cache levels. For example, the following levels can be set: High-frequency data: These data are frequently accessed and should be pre-loaded into the memory cache to ensure fast access. Medium-frequency data: These data are not frequently accessed but have a certain access frequency and can be stored in the local disk cache. Low-frequency data: Data with a lower access frequency is stored in a remote storage or archival system. This division is achieved by setting heat value thresholds. The thresholds can be dynamically adjusted according to the actual application and requirements to ensure that the data can be reasonably stored according to the access frequency. Define the rules for pre-loading high-frequency data into memory. After the cache level division, the system needs to define the rules for pre-loading high-frequency data into memory. For high-frequency data, the system must ensure that it is in the memory for fast access. Therefore, the rules are defined as follows: Rule 1: Loading based on the heat threshold: If the heat value of a certain data exceeds a certain threshold, it will be loaded into the memory cache. Rule 2: Prediction based on the access frequency: According to the trend of historical access records, the system can predict the data that will become hot in the future and pre-load these data into the memory cache in advance. The goal of these rules is to ensure that high-frequency data can quickly respond to access requests and improve the overall performance of the system. Design a dynamic prefetch window to predict future hot data. To cope with future hot data, the system designs a dynamic prefetch window. The dynamic prefetch window is based on the current data access history and predicts the data that may become hot in the future for a certain period of time. The specific process is as follows: Data access pattern analysis: By analyzing the access patterns of historical data, identify which data has a high access trend. Time series prediction: Use time series analysis methods (such as LSTM, ARIMA, etc.) to predict which data will have an increasing access volume in the future for a certain period of time, thus forming future hot data. Task queue generation: According to the prediction results, the system generates a prefetch task queue, and these tasks will pre-load the predicted hot data into the cache in advance for subsequent access. Through the dynamic prefetch window, the delay caused by cache misses can be effectively reduced, and the data processing efficiency can be improved. Dynamically adjust the cache weights through the spatio-temporal attention mechanism. To further optimize the cache efficiency, the system introduces the spatio-temporal attention mechanism to dynamically adjust the cache weights. The spatio-temporal attention mechanism combines the time and space dimensions and dynamically adjusts the data distribution in the cache. The specific process is as follows: Time dimension: As time goes by, the access frequency of some data may change. For example, some data may be frequently accessed during a specific time period and less accessed during other time periods.The spatio-temporal attention mechanism can dynamically adjust the caching strategy according to the changes in time, increase the weight of hot data, and reduce the weight of low-frequency data. Spatial dimension: The access of data is usually related to its storage location. Spatially, some data may need to be stored and accessed at specific geographical locations (e.g., locations close to users). The spatio-temporal attention mechanism can dynamically adjust the caching location according to the spatial distance and the load conditions of storage nodes. By combining time and space information, the spatio-temporal attention mechanism can automatically adjust the weight of each data block in the cache to achieve an optimal storage strategy. Mapping the caching location and the expiration period into a visualization distribution map, the system displays the distribution map of the caching location and the expiration period through visualization means. This map intuitively shows the storage location and expiration period of each data block in the cache, enabling the system administrator to quickly understand the cache distribution and adjust the strategy in a timely manner. Caching location mapping: According to the location where the data is stored (such as memory, disk, remote storage, etc.), visualize the location of the data block to show which data is stored on efficient storage media and which is stored on slower media. Expiration period mapping: Show the expiration period of the data in the cache, indicating which data needs to be updated and which data can be cleared, thereby improving the utilization rate of the cache space. This visualization distribution map provides an intuitive basis for data management, facilitating the administrator to make optimization adjustments.
[0041] Specifically, the specific process of obtaining the optical fiber storage efficiency index by integrating data access latency, index hit rate, and replica migration frequency is as follows: Comprehensively adjust the weight ratios of access latency, index hit rate, and replica migration frequency through the analytic hierarchy process; Introduce a time decay factor to perform exponential smoothing on historical performance data, and generate the optical fiber storage efficiency index through linear weighted fusion.
[0042] In this implementation plan, the formula for the optical fiber storage efficiency index: ; Explanation of the formula: : The weight coefficients obtained through the analytic hierarchy process (AHP), corresponding to data access latency, index hit rate, and replica migration rate respectively. These weights reflect the relative importance of different indicators in storage efficiency. : The data access latency at the current moment. : The data access latency at the previous moment. : The index hit rate at the current moment. : The replica migration rate at the current moment. : The decay factor, representing the contribution ratio of the data access latency at the current moment and the previous moment to the efficiency respectively. Usually, will be larger, indicating that the latency at the current moment has a greater impact on the efficiency index. : The introduced time decay factor, where is the decay rate, is the time step. This factor attenuates historical data to ensure that the storage efficiency index mainly focuses on current performance data and reduces the impact of historical data.
[0043] Specifically, when the fiber optic storage efficiency index is lower than the preset threshold, the specific process of generating the structure optimization instruction set and feeding it back to the data scheduling module and the index optimization module is as follows: Establish a multi-objective decision tree, and select optimization strategies according to the branches of the fiber optic storage efficiency index, including: Too high latency: Trigger the optical path priority reset instruction and the edge cache expansion instruction; Too low hit rate: Generate the index level migration instruction and the prefetch policy update instruction; Too frequent migration: Start the replica distribution topology reconstruction instruction and the load balancing policy.
[0044] In this implementation, a multi-objective decision tree is established: The multi-objective decision tree is a decision-making model constructed based on multiple storage performance objectives (such as latency, hit rate, migration frequency, etc.). This tree makes branch selections according to the changes in the fiber optic storage efficiency index based on a predetermined threshold. When the efficiency index is lower than the threshold, each branch of the decision tree corresponds to different optimization strategies. Decision tree branches and strategy selection: Excessive latency: If the access latency of the fiber optic storage system exceeds the preset threshold, the system will trigger two main instructions: Optical path priority reset instruction: By readjusting the priority of the fiber optic path, optimize the access path of the fiber optic storage data and reduce the transmission latency. Edge cache expansion instruction: Increase the cache capacity of the edge nodes, reduce the frequency of data transmission to the storage center, and thus reduce the latency. Low hit rate: If the index hit rate is low, it means that the cache or storage hierarchy is often not hit during data access. At this time, the system will take the following optimization measures: Index level migration instruction: Adjust the data level structure of the index, migrate the index level of hot data upward, and improve the access hit rate. Prefetch policy update instruction: Update the prefetch policy, predict the data that may be needed in the future based on historical access data, and preload it into the cache in advance, thereby improving the data hit rate. Excessive migration: If the replica migration frequency is too high, it means that there are frequent replica relocations in the system, which may lead to waste of network bandwidth or uneven load. In response to this, the system will trigger the following optimization instructions: Replica distribution topology reconstruction instruction: Reconstruct the replica distribution, reasonably distribute the data replicas according to characteristics such as geographical location and access frequency, and reduce unnecessary data migrations. Load balancing strategy: Reallocate resources according to the load situation, avoid overloading of some nodes, ensure system load balancing, and reduce performance degradation caused by resource imbalance. Feedback to the data scheduling module and the index optimization module: The above optimization instructions are input into the data scheduling module and the index optimization module through a feedback mechanism. These modules will make adjustments according to the optimization instructions and perform operations such as resource reallocation, path optimization, and cache policy update. Data scheduling module: Dynamically optimize the data scheduling according to the optical path priority and cache expansion instructions to ensure the minimum system response latency. Index optimization module: Optimize the data index strategy according to the index level migration and prefetch policy update instructions, improve the cache hit rate, and improve the overall performance of the system.
[0045] Please refer to Figure 2, A fiber optic data storage management method based on big data, comprising the following steps: S1. Extract the spatio-temporal features of the historical access records of the fiber optic storage system through the LSTM network, calculate the data heat value in combination with the time decay function, and generate a heat distribution map identifying high-frequency and low-frequency data blocks; S2. According to the heat distribution map, construct a storage resource allocation model through the deep reinforcement learning algorithm, define the bandwidth utilization rate, node load rate, and predicted access volume as state space parameters, output the optical path priority matrix and replica distribution topology instructions, and monitor the data flow fluctuations through the sliding window mechanism to trigger the migration of high-frequency data to low-load nodes; S3. Based on the heat distribution map and replica distribution topology instructions, construct a spatio-temporal joint index for structured data and a vector similarity index for unstructured data, and optimize the prefetch strategy according to the data heat value to generate a high-frequency data cache index distribution map; S4. Comprehensively consider the data access latency, index hit rate, and replica migration frequency, obtain the fiber optic storage efficiency index, and when the fiber optic storage efficiency index is lower than the preset threshold, generate a structure optimization instruction set and feedback it to the data scheduling module and index optimization module.
[0046] In this implementation, a method for fiber optic data storage management based on big data aims to optimize the performance of the fiber optic storage system by intelligently managing and dynamically adjusting aspects such as historical data access records, heat distribution, storage resource allocation, and index optimization. The following is a specific explanation of the functions of each step: S1: The goal of this step is to perform spatio-temporal analysis on the historical access records of the fiber optic storage system to identify high-frequency and low-frequency data blocks in the system. Long Short-Term Memory (LSTM) networks are used to extract temporal features from the access records to more accurately predict the heat changes of the data. Combining with the time decay function, the timeliness of data access can be better considered, causing the heat values of historical access data to decay, thereby generating a heat distribution map reflecting the data access frequency. These maps can help identify which data are high-frequency access data (hot data) and which are low-frequency access data (cold data), providing a basis for subsequent storage resource optimization. S2: In this step, combined with the heat distribution map, the Deep Reinforcement Learning (DRL) algorithm is used to optimize the storage resource allocation. The bandwidth utilization rate, node load rate, and predicted access volume are used as state space parameters and input into the model. The model will generate resource allocation strategies based on these input parameters, including an optical path priority matrix and a replica distribution topology instruction. The optical path priority matrix determines the transmission priority of data streams between different paths, and the replica distribution topology instruction optimizes the storage location and migration strategy of data replicas. The sliding window mechanism can monitor the changes in data streams, promptly identify high-frequency data, and migrate it to low-load nodes, thereby improving the overall load balance and storage efficiency of the system. S3: This step focuses on data storage and index optimization. By according to the heat distribution map and the replica distribution topology instruction, a spatio-temporal joint index and a vector similarity index are constructed for the management of structured data and unstructured data respectively. The spatio-temporal joint index can effectively associate the time and space characteristics of data, optimizing data retrieval and storage efficiency; the vector similarity index is optimized by calculating the similarity between data, especially suitable for the storage and query of unstructured data (such as images, texts, etc.). In addition, combined with the data heat value, the prefetch strategy can be optimized to ensure that high-frequency data maintains priority in the system, thereby generating a high-frequency data index distribution map in the cache and further improving the response speed and storage efficiency of data access. S4: The function of this step is to calculate a fiber optic storage efficiency index reflecting the efficiency of the fiber optic storage system by comprehensively considering data access latency, index hit rate, and replica migration frequency. When this index is lower than the set threshold, it indicates that the current system has low storage efficiency and may have performance bottlenecks. In this case, the system will automatically generate a set of structure optimization instructions and feedback them to the data scheduling module and the index optimization module to guide subsequent resource adjustment and optimization strategies. These instruction sets can involve operations such as fiber optic storage path adjustment, cache optimization, and index level update, thereby improving the performance of the overall storage system.
[0047] In summary, the present application has at least the following effects:
[0048] A fiber optic data storage management system and method based on big data, by combining spatio-temporal feature analysis, heat distribution maps, and replica distribution topology instructions, realizes the precise positioning and optimized storage of data, reduces unnecessary data migration and redundant storage, thereby improving the overall storage efficiency of the fiber optic storage system. Through an optimized storage resource allocation strategy, dynamic load balancing, and data migration based on deep reinforcement learning, it ensures that high-frequency data is preferentially stored in low-load nodes, reduces data access latency, and improves the system response speed. By constructing spatio-temporal joint indexes and vector similarity indexes, the storage and query processes are optimized for structured and unstructured data respectively, significantly improving the data retrieval efficiency, especially the access speed of unstructured data. Through the application of a sliding window mechanism, historical performance data analysis, and a time decay function, it can dynamically adjust the cache weight, data prefetch strategy, and fiber optic storage path optimization to ensure that the system automatically self-optimizes according to real-time data flow fluctuations. When the fiber optic storage efficiency index is lower than a preset threshold, the system can generate a structure optimization instruction set and feedback it to the data scheduling module and index optimization module to achieve targeted optimization and ensure the stability and efficient operation of the system. By combining a multi-objective decision tree and a deep reinforcement learning algorithm, the system can adjust the storage resource allocation strategy and data migration plan according to real-time situations, making it highly flexible and scalable to adapt to changing data flows and access requirements. By dynamically adjusting the replica distribution topology instructions, it optimizes the storage location and migration strategy of replicas, avoids over-frequent replica migration and unnecessary data redundancy, thereby improving the resource utilization rate and reliability of the fiber optic storage system.
[0049] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0050] The present invention is described with reference to the flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate for realizing in the process Figure 1one or more processes and / or blocks Figure 1 a device for the functions specified in one or more blocks
[0051] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks
[0052] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks
[0053] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention
[0054] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the equivalent technologies of the claims of the present invention, the present invention is also intended to include these modifications and variations
Claims
1. A fiber optic data storage and management system based on big data, characterized in that, It includes the following modules: data access analysis module, data scheduling module, index optimization module, and storage efficiency evaluation module; The data access analysis module is used to extract the spatio-temporal features of the historical access records of the optical fiber storage system through the LSTM network, calculate the data heat value in combination with the time decay function, and generate a heat distribution map identifying high-frequency and low-frequency data blocks; The data scheduling module is used to construct a storage resource allocation model through the deep reinforcement learning algorithm according to the heat distribution map, define the bandwidth utilization rate, node load rate, and predicted access volume as state space parameters, output the optical path priority matrix and the replica distribution topology instruction, and monitor the data flow fluctuation through the sliding window mechanism to trigger the migration of high-frequency data to low-load nodes; The index optimization module is used to construct a spatio-temporal joint index for structured data and a vector similarity index for unstructured data based on the heat distribution map and the replica distribution topology instruction, and optimize the prefetch policy according to the data heat value to generate a high-frequency data cache index distribution map; The storage efficiency evaluation module is used to comprehensively consider the data access latency, index hit rate, and replica migration frequency to obtain the optical fiber storage efficiency index. When the optical fiber storage efficiency index is lower than the preset threshold, a structure optimization instruction set is generated and fed back to the data scheduling module and the index optimization module; The specific process of extracting the spatio-temporal features of the historical access records of the optical fiber storage system through the LSTM network and calculating the data heat value in combination with the time decay function is as follows: Perform spatio-temporal encoding on the historical access records, extract the timestamp, geographical coordinates, and concurrent request volume, and construct a spatio-temporal feature vector; Model the time dimension features through a bidirectional LSTM network to capture the periodic pattern of data access, and at the same time extract the space dimension features through a convolutional network; Input the spatio-temporal features into the spatio-temporal attention mechanism to dynamically allocate time decay weights and generate a spatio-temporal correlation weight matrix; Input the spatio-temporal correlation weight matrix into the time decay function, dynamically decay the influence of historical data according to the time distance, and perform normalization processing on the decayed data to generate a globally comparable data heat value.
2. The fiber optic data storage management system based on big data according to claim 1, characterized in that: The specific process of generating a heat distribution map identifying high-frequency and low-frequency data blocks is as follows: Perform geographical grid division on the storage space based on the data heat value, and map the access location to a hierarchical space grid through geographical hash encoding; Identify high-frequency access grid clusters through the density clustering algorithm, and merge adjacent high-density grids to form hot spots; Separate low-frequency cold data regions through the spatio-temporal joint clustering algorithm, and dynamically adjust the block boundary in combination with the access interval stability; Mark the high-frequency blocks as red hot spots and the low-frequency blocks as blue cold regions to generate a visual heat distribution map.
3. A fiber optic data storage management system based on big data according to claim 2, characterized in that: The specific process of constructing a storage resource allocation model through the deep reinforcement learning algorithm according to the heat distribution map is as follows: Construct a hierarchical state space, and encode the block heat value, optical fiber network topology structure, and load data in the heat distribution map into a multi-dimensional state vector; Design a multi-objective reward function, with minimizing transmission delay, balancing node load, and maximizing throughput as the joint optimization objectives, and increasing the delay penalty weight for high-frequency blocks; The model is trained by the Deep Deterministic Policy Gradient algorithm, and the action decision-making network generates a resource allocation strategy to construct a storage resource allocation model.
4. A fiber optic data storage management system based on big data according to claim 3, characterized in that: Define the bandwidth utilization rate, node load rate, and predicted access volume as state space parameters. The specific process of outputting the optical path priority matrix and replica distribution topology instructions is as follows: Jointly encode the bandwidth utilization rate, node load rate, and predicted access volume to construct a dynamically perceiving state space matrix; Calculate the fiber path priority through the optical path game model, dynamically adjust the path weight in combination with the real-time load, and generate the optical path priority matrix; Design a replica distribution algorithm with spatio-temporal constraints, generate replica topology instructions in combination with the heat value distribution and node geographical distance, and dynamically correct the instruction parameters through a sliding window verification mechanism.
5. A fiber optic data storage management system based on big data according to claim 4, characterized in that: Based on the heat distribution map and replica distribution topology instructions, the specific process of constructing a spatio-temporal joint index for structured data and a vector similarity index for unstructured data is as follows: Parse the high-frequency block coordinates and heat values in the heat map, and generate spatio-temporal key values in combination with the replica topology instructions; Construct a multi-level B+ tree index based on the spatio-temporal key values, extract the semantic features of unstructured data, generate a weighted vector space in combination with the heat values, construct a vector index through the distributed HNSW algorithm, and preferentially store high-heat data to low-latency nodes; Dynamically associate the spatio-temporal joint index and the vector similarity index to construct a cross-modal retrieval path.
6. The fiber optic data storage management system based on big data according to claim 5, characterized in that: Optimize the prefetching strategy according to the data heat value. The specific process of generating a high-frequency data cache index distribution map is as follows: Divide the cache levels according to the data heat value and define the rules for preloading high-frequency data into memory; Design a dynamic prefetching window to predict future hot data and generate a prefetching task queue; Dynamically adjust the cache weight through the spatio-temporal attention mechanism, and map the cache location and validity period to a visual distribution map.
7. A fiber optic data storage management system based on big data according to claim 6, characterized in that: Comprehensively consider the data access latency, index hit rate, and replica migration frequency. The specific process of obtaining the fiber storage efficiency index is as follows: Comprehensively adjust the weight ratio of access latency, index hit rate, and replica migration frequency through the Analytic Hierarchy Process; Introduce a time decay factor, perform exponential smoothing on the historical performance data, and generate the fiber storage efficiency index through linear weighted fusion.
8. A fiber optic data storage management system based on big data according to claim 7, characterized in that: When the fiber storage efficiency index is lower than the preset threshold, the specific process of generating a structure optimization instruction set and feedbacking it to the data scheduling module and index optimization module is as follows: Establish a multi-objective decision tree, and select optimization strategies according to the branches of the fiber storage efficiency index, including: Excessive latency: Trigger the optical path priority reset instruction and the edge cache expansion instruction; Low hit rate: Generate the index level migration instruction and the prefetching strategy update instruction; Frequent migration: Start the replica distribution topology reconstruction instruction and the load balancing strategy.
9. A fiber optic data storage management method based on big data, applied to a fiber optic data storage management system according to any one of claims 1-8, characterized in that Include the following steps: S1. Extract the spatio-temporal features of the historical access records of the fiber storage system through the LSTM network, calculate the data heat value in combination with the time decay function, and generate a heat distribution map identifying high-frequency and low-frequency data blocks; S2. According to the heat distribution atlas, construct a storage resource allocation model through a deep reinforcement learning algorithm. Define the bandwidth utilization rate, node load rate, and predicted access volume as state space parameters, output the optical path priority matrix and replica distribution topology instructions, and monitor the data flow fluctuation through a sliding window mechanism to trigger the migration of high-frequency data to low-load nodes; S3. Based on the heat distribution atlas and replica distribution topology instructions, construct a spatio-temporal joint index for structured data and a vector similarity index for unstructured data, and optimize the prefetching strategy according to the data heat value to generate a high-frequency data cache index distribution map; S4. Synthesize the data access latency, index hit rate, and replica migration frequency to obtain the fiber storage efficiency index. When the fiber storage efficiency index is lower than the preset threshold, generate a structure optimization instruction set and feedback it to the data scheduling module and index optimization module.
Citation Information
Patent Citations
Data dynamic storage method and device, equipment and storage medium
CN118915978A
Task scheduling optimization analysis method and system based on big data
CN119883580A