Industrial data asset distributed storage method, medium and system

By building a resource consumption calculation model and storage optimization neural network, extracting multi-dimensional characteristics of industrial data assets, intelligent scheduling and adaptive allocation of storage resources in the distributed storage system of industrial data assets is realized, the problems of storage hotspots and resource waste in existing systems are solved, and the overall storage efficiency and adaptability of the system are improved.

CN120196680AActive Publication Date: 2025-06-24BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510277291.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-24
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Existing industrial data asset distributed storage systems are difficult to realize intelligent scheduling of storage resources, resulting in problems such as storage hotspots, resource waste and performance bottlenecks.

Method used

By building a resource consumption calculation model and a storage optimization neural network, extracting multi-dimensional features of industrial raw data flow, building a data feature mapping matrix and resource scheduling matrix, and using the collaborative optimization of deep optimization convolution kernel and resource scheduling matrix, intelligent scheduling and adaptive allocation of storage resources can be achieved.

Benefits of technology

It realizes intelligent extraction of storage characteristics of industrial data assets and dynamic optimization of storage resources, improves the overall storage efficiency of the system, enhances the ability to adapt to changes in data access patterns, and solves problems such as storage hotspots and load imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196680A_ABST
    Figure CN120196680A_ABST
Patent Text Reader

Abstract

The invention provides an industrial data asset distributed storage method, medium and system, and belongs to the technical field of electrical digital data processing.The method comprises the steps that firstly, manufacturing equipment operation duration data, maintenance period data, task quantity data and data generation frequency data are collected to serve as industrial original data streams, and a resource consumption calculation model is established; data features are extracted, a feature mapping matrix is constructed, resource scheduling optimization is realized through a storage optimization neural network, storage unit allocation is performed based on a deep optimization convolution kernel, dynamic evaluation and optimization adjustment are performed on data block storage positions, and finally distributed storage of industrial original data streams is realized. Through real-time calculation of the storage performance indexes and the load factors, a complete performance evaluation and optimization feedback mechanism is established, balanced distribution and efficient utilization of the storage resources are ensured, and the technical problem that in the prior art, an industrial data asset distributed storage system is difficult to achieve intelligent scheduling of the storage resources is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electric digital data processing. Specifically, it relates to a distributed storage method, medium and system for industrial data assets. Background Art

[0002] Industrial data asset management is an important part of industrial digital transformation. With the rapid development of industrial Internet of Things and intelligent manufacturing, the scale and complexity of industrial data show an exponential growth trend. Traditional industrial data storage solutions mainly adopt a centralized storage architecture, and the hardware configuration of storage servers is extended to meet the data growth demand. With the continuous expansion of data scale, the limitations of the centralized storage architecture in aspects such as data access efficiency, system reliability and resource utilization become increasingly prominent, and distributed storage technology gradually becomes the mainstream solution for industrial data asset management.

[0003] Current mainstream distributed storage systems usually adopt static data sharding strategies and fixed resource allocation schemes, and the system lacks the ability to perceive and dynamically adjust the utilization of storage resources in real time during operation. This static resource scheduling method causes the system to easily have problems such as storage hotspots, resource waste and performance bottlenecks when facing dynamic data access patterns and load distributions. At the same time, existing distributed storage systems often adopt empirical rule designs in data layout optimization, and it is difficult to adapt to the characteristics of high timeliness, strong correlation and suddenness of industrial data, which affects the overall storage efficiency of the system.

[0004] In the prior art, although some resource scheduling methods based on load balancing have been proposed, these methods mainly focus on some local performance indicators of the system and lack global optimization of the overall operation state of the storage system. At the same time, when dealing with feature extraction and storage optimization of industrial data assets, existing methods fail to make full use of the advantages of deep learning technology and are difficult to achieve intelligent scheduling and adaptive optimization of storage resources, resulting in low resource utilization efficiency during the actual operation of the system, that is, there is a technical problem that it is difficult to achieve intelligent scheduling of storage resources in the distributed storage system for industrial data assets in the prior art. Summary of the Invention

[0005] In view of this, the present invention provides a distributed storage method, medium and system for industrial data assets, which can solve the technical problem that it is difficult to achieve intelligent scheduling of storage resources in the distributed storage system for industrial data assets in the prior art.

[0006] The present invention is implemented as follows: A method for distributed storage of industrial data assets provided by the first aspect of the present invention includes the following steps: Collecting the operation duration data of manufacturing equipment, the maintenance cycle data of manufacturing equipment, the production task quantity data, and the data generation frequency data as industrial raw data streams; Establishing a resource consumption calculation model according to the industrial raw data streams; Using the resource consumption calculation model to extract the data time series feature values, data space feature values, and data association feature values of the industrial raw data streams, and constructing a data feature mapping matrix; Constructing a resource scheduling matrix based on the data feature mapping matrix, and inputting the resource scheduling matrix into a storage optimization neural network; Obtaining an optimized scheduling convolution kernel according to the resource scheduling convolution kernel of the storage optimization neural network, wherein the resource weight coefficient of the optimized scheduling convolution kernel is obtained based on the ratio of the corresponding weight of the resource scheduling convolution kernel to the reference weight coefficient; Applying the optimized scheduling convolution kernel to the resource scheduling matrix to obtain a resource feature distribution map; Obtaining a deep optimized convolution kernel in the deep network of the storage optimization neural network that has a mapping relationship with the resource scheduling convolution kernel; When the ratio of the resource weight coefficient of the resource scheduling convolution kernel to the reference weight coefficient is greater than the storage capacity threshold, performing proportional reduction processing on the corresponding weight of the resource scheduling convolution kernel.

[0007] Among them, the resource consumption calculation model normalizes the operation duration data of the manufacturing equipment, the maintenance cycle data of the manufacturing equipment, the production task quantity data, and the data generation frequency data, and then calculates the data storage demand index through a polynomial fitting equation. The data storage demand index is used to characterize the storage resource consumption degree of the industrial raw data stream.

[0008] Among them, the storage optimization neural network is composed of an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a first fully connected layer, a second fully connected layer, and an output layer. The training data set of the storage optimization neural network is obtained by collecting the data block storage locations, storage space utilization rates, data access latency parameters, storage node load parameters, and data redundancy parameters of the industrial field for 365 days. The storage optimization neural network is trained using the backpropagation algorithm, and the network parameters are optimized and adjusted through the training data set.

[0009] Among them, feature extraction operations are performed on the resource feature distribution map to obtain the spatial division parameters of the storage basic unit, and the data access latency parameters of the storage basic unit are obtained. The feature extraction operations include weighted cumulative operations on the resource weight coefficient and the feature values of the resource feature distribution map.

[0010] Among them, the distribution parameters of the storage optimization unit are determined based on the deep optimization convolution kernel, and the distribution parameters include the storage node load parameter and the data redundancy parameter; the storage optimization unit is subjected to distributed node allocation, and a storage address mapping table is established based on the storage space utilization rate and the data access latency parameter.

[0011] Among them, data sharding operations are performed according to the storage address mapping table, and the storage locations of data blocks are determined based on the storage node load parameter and the data redundancy parameter.

[0012] Among them, the deep network of the storage optimization neural network is used to dynamically evaluate the storage locations of the data blocks, and a storage performance index is generated based on the storage space utilization rate, the data access latency parameter, the storage node load parameter, and the data redundancy parameter.

[0013] Among them, the storage load factor is calculated based on the storage performance index. When the storage load factor is less than the storage load threshold, capacity expansion processing is performed on the storage locations of the data blocks; the storage locations of the data blocks are optimized and adjusted according to the storage load factor, and the distributed storage of the industrial raw data stream is completed.

[0014] The second aspect of the present invention provides a computer-readable storage medium. Program instructions are stored in the computer-readable storage medium. When the program instructions run on a computer, they are used to execute the above-mentioned industrial data asset distributed storage method.

[0015] The third aspect of the present invention provides an industrial data asset distributed storage system, which includes the above-mentioned computer-readable storage medium. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is provided inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is provided inside the system.

[0016] Compared with the prior art, the present invention provides an industrial data asset distributed storage method, medium, and system. The industrial data asset distributed storage method proposed by the present invention realizes the intelligent extraction of the storage characteristics of industrial data assets and the dynamic optimization of storage resources by constructing a resource consumption calculation model and a storage optimization neural network. The method first establishes a data feature mapping matrix based on the industrial raw data stream, and uses a multi-dimensional feature expression method to comprehensively describe the temporal, spatial, and correlation characteristics of the data. Then, through the collaborative optimization of the deep optimization convolution kernel and the resource scheduling matrix, the intelligent scheduling and adaptive allocation of storage resources are realized.

[0017] The method of the present invention organically combines deep learning technology with a distributed storage system by introducing a storage optimization neural network, overcoming the limitations of static rule design in traditional methods. Based on the feature mapping mechanism of the deep network, the system can accurately grasp the dynamic change characteristics of industrial data assets and perform real-time optimization and adjustment of storage resources accordingly. At the same time, by establishing an evaluation system for storage performance indicators and load factors, global monitoring and optimization of the operating status of the storage system are realized, effectively solving problems such as storage hotspots and load imbalance.

[0018] The present invention successfully solves the technical problems of intelligent scheduling of storage resources and load balancing in a distributed storage system for industrial data assets. Through a resource optimization mechanism driven by deep learning, efficient utilization of storage resources and continuous optimization of system performance are achieved. This method not only improves the overall storage efficiency of the system but also enhances the system's adaptability to changes in data access patterns, solving the technical problem in the prior art that it is difficult to achieve intelligent scheduling of storage resources in a distributed storage system for industrial data assets. Description of the Drawings

[0019] Figure 1 It is a flowchart of the method of the present invention.

[0020] Figure 2 It is a graph showing the change trend of equipment operation data within 31 days in Example 2.

[0021] Figure 3 It is a graph showing the distribution of storage demand indexes of 50 devices in Example 2.

[0022] Figure 4 It is a graph showing the analysis results of the contribution rates of various eigenvalues in Example 2.

[0023] Figure 5 It is a graph showing the comparison of various performance indicators before and after system optimization in Example 2. Detailed Embodiments

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention.

[0025] As Figure 1 shown, it is a flowchart of a distributed storage method for industrial data assets provided in the first aspect of the present invention. This method includes the following steps:

[0026] S01. Collect operation duration data of manufacturing equipment, maintenance cycle data of manufacturing equipment, production task quantity data, and data generation frequency data as industrial original data streams;

[0027] S02. Establish a resource consumption calculation model based on the industrial raw data stream. The resource consumption calculation model generates a data storage requirement index based on the manufacturing equipment operation duration data, the manufacturing equipment maintenance cycle data, the production task quantity data, and the data generation frequency data;

[0028] S03. Use the resource consumption calculation model to extract the data time series feature values, data space feature values, and data association feature values of the industrial raw data stream, and construct a data feature mapping matrix based on the data time series feature values, the data space feature values, and the data association feature values;

[0029] S04. Construct a resource scheduling matrix based on the data feature mapping matrix, and input the resource scheduling matrix into a storage optimization neural network;

[0030] S05. Obtain an optimized scheduling convolution kernel according to the resource scheduling convolution kernel of the storage optimization neural network, where the resource weight coefficient of the optimized scheduling convolution kernel is obtained based on the ratio of the corresponding weight of the resource scheduling convolution kernel to the reference weight coefficient in the resource scheduling convolution kernel;

[0031] S06. Apply the optimized scheduling convolution kernel to the resource scheduling matrix to obtain a resource feature distribution map, and determine the storage space utilization rate of the data storage node based on the resource feature distribution map;

[0032] S07. Perform feature extraction operations on the resource feature distribution map to obtain the space division parameters of the storage basic unit, and obtain the data access delay parameters of the storage basic unit, where the feature extraction operations include weighted cumulative operations on the resource weight coefficient and the feature values of the resource feature distribution map;

[0033] S08. Obtain a deep optimization convolution kernel in the deep network of the storage optimization neural network that has a mapping relationship with the resource scheduling convolution kernel;

[0034] S09. When the ratio of the resource weight coefficient of the resource scheduling convolution kernel to the reference weight coefficient is greater than the storage capacity threshold, perform proportional reduction processing on the corresponding weight of the resource scheduling convolution kernel;

[0035] S10. Determine the distribution parameters of the storage optimization unit based on the deep optimization convolution kernel, where the distribution parameters include storage node load parameters and data redundancy parameters;

[0036] S11. Perform distributed node allocation on the storage optimization unit, and establish a storage address mapping table based on the storage space utilization rate and the data access delay parameters;

[0037] S12. Perform data sharding operations according to the storage address mapping table, and determine the storage locations of data blocks based on the storage node load parameters and the data redundancy parameters;

[0038] S13. Dynamically evaluate the storage locations of the data blocks using the deep network of the storage optimization neural network, and generate storage performance metrics based on the storage space utilization rate, the data access latency parameters, the storage node load parameters, and the data redundancy parameters;

[0039] S14. Calculate the storage load factor based on the storage performance metrics. When the storage load factor is less than the storage load threshold, perform capacity expansion processing on the storage locations of the data blocks;

[0040] S15. Optimally adjust the storage locations of the data blocks according to the storage load factor to complete the distributed storage of industrial raw data streams.

[0041] The resource consumption calculation model normalizes the manufacturing equipment operation duration data, the manufacturing equipment maintenance cycle data, the production task quantity data, and the data generation frequency data, and then calculates the data storage requirement index through a polynomial fitting equation. The data storage requirement index is used to characterize the storage resource consumption degree of industrial raw data streams.

[0042] The storage optimization neural network consists of an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a first fully connected layer, a second fully connected layer, and an output layer. The training dataset of the storage optimization neural network is obtained by collecting the storage locations of data blocks, the storage space utilization rate, the data access latency parameters, the storage node load parameters, and the data redundancy parameters of the industrial site for 365 days. The storage optimization neural network is trained using the backpropagation algorithm, and the network parameters are optimally adjusted through the training dataset.

[0043] The storage performance metrics are calculated through a weighted average equation based on the storage space utilization rate, the data access latency parameters, the storage node load parameters, and the data redundancy parameters. The storage performance metrics are used to evaluate the overall operation efficiency of the distributed storage system.

[0044] The storage load factor is calculated through a standard deviation equation based on the storage space utilization rate, the data access latency parameters, the storage node load parameters, and the data redundancy parameters of the storage block storage locations. The storage load factor is used to characterize the load balancing degree among the nodes of the distributed storage system.

[0045] The specific implementation manners of the above steps are described in detail below. The specific implementation manner of step S01 is to collect device operation data through a data collection module deployed on the manufacturing device. This module includes a time series database for storing device operation duration data, and records the device startup time, shutdown time, and cumulative operation duration in real time through a preset data collection frequency; at the same time, configure a maintenance management system to record the time intervals and completion status of maintenance activities such as regular maintenance and repair of the device; obtain production task data such as the number of production task dispatch orders and task completion status in the production execution system; based on the data sampling period of the time series database, obtain the generation frequencies of various types of data, and input these data as the original data stream into the subsequent processing links. The purpose of this step is to establish the basic data source of industrial data assets.

[0046] The specific implementation manner of step S02 is to first perform normalization processing on various types of data in the industrial original data stream, and uniformly map the value range to between 0 and 1. Specifically, the maximum-minimum normalization method is used, that is, for any data x, the normalized value is the data minus the minimum value of the data sequence divided by the difference between the maximum value and the minimum value of the data sequence; then substitute the normalized device operation duration data, maintenance cycle data, task quantity data, and data generation frequency data into a polynomial fitting equation. This equation adopts a cubic polynomial form, and the polynomial coefficients are determined by the least squares method to calculate the data storage requirement index. The larger this index is, the higher the storage resource consumption degree is. The purpose of this step is to quantitatively evaluate the demand degree of industrial data assets for storage resources.

[0047] The specific implementation manner of step S03 is to first perform time series analysis on the industrial original data stream using a resource consumption calculation model, calculate the time correlation of the data using the sliding window method, set the window size to 24 hours and the step size to 1 hour to obtain a time series feature value reflecting the time distribution characteristics of the data; secondly, perform spatial feature extraction, use the principal component analysis method to perform dimensionality reduction processing on the data, and retain the principal components with a cumulative contribution rate reaching 85% as the spatial feature values; thirdly, calculate the correlation features between the data, use the Pearson correlation coefficient method to analyze the correlation between data items to obtain the correlation feature values; finally, organize these three types of feature values into a three-dimensional feature mapping matrix. The purpose of this step is to characterize the characteristic attributes of industrial data assets from multiple dimensions.

[0048] The specific implementation manner of step S04 is to construct a resource scheduling matrix based on the data feature mapping matrix. Specifically, perform standardization processing on the feature mapping matrix using the z-score standardization method, that is, subtract the mean from the original data and then divide by the standard deviation; then apply the singular value decomposition algorithm to the standardized matrix to extract the main eigenvectors to construct the resource scheduling matrix; finally, input this matrix into a pre-trained storage optimization neural network. The purpose of this step is to provide the basic data structure for subsequent storage optimization.

[0049] The specific implementation of step S05 is to first obtain the resource scheduling convolution kernel in the storage optimization neural network. The initial weights of this convolution kernel are randomly initialized through Gaussian distribution. Then, compare this convolution kernel with a preset reference weight coefficient, which is determined through statistical analysis of historical data and set to 0.5. Calculate the ratio of each weight in the resource scheduling convolution kernel to the reference weight coefficient, and use this ratio as the resource weight coefficient of the optimized scheduling convolution kernel. The purpose of this step is to construct an optimized convolution kernel suitable for the current storage requirements.

[0050] The specific implementation of step S06 is to apply the optimized scheduling convolution kernel to the resource scheduling matrix, perform convolution operations in the form of a sliding window with a window size of 3×3 and a step size of 1 to obtain a feature map. Then, perform a pooling operation on the feature map, use the max-pooling method to extract local features, and obtain a resource feature distribution map. Calculate the storage space usage of each storage node based on this distribution map to obtain the storage space utilization rate. The purpose of this step is to obtain the distribution status of storage resources.

[0051] The specific implementation of step S07 is to extract features from the resource feature distribution map. First, calculate the weighted sum of the resource weight coefficient and the feature value using the linear weighting method, and the weights are set according to the importance of the features, with larger weights assigned to features with higher importance. Then, determine the spatial partitioning parameters of the storage basic unit based on the weighted result, including unit size, distribution density, etc. At the same time, obtain the data access latency of the storage basic unit through network probing methods and record the time interval from sending a data request to receiving a response. The purpose of this step is to determine the basic configuration parameters of the storage system.

[0052] The specific implementation of step S08 is to find the deep convolution kernel with a feature mapping relationship with the resource scheduling convolution kernel in the deep network of the storage optimization neural network. First, calculate the feature similarity between the two convolution kernels using the cosine similarity method. Then, select the deep convolution kernel with the highest similarity as the deep optimized convolution kernel. The purpose of this step is to establish the correspondence between shallow features and deep features.

[0053] The specific implementation of step S09 is to compare the ratio of the resource weight coefficient of the resource scheduling convolution kernel to the reference weight coefficient. When this ratio exceeds the storage capacity threshold, which is set to 1.5, the weights of the convolution kernel need to be adjusted. Specifically, use the linear proportion reduction method to reduce the weights exceeding the threshold proportionally to meet the storage capacity requirements. The purpose of this step is to prevent over-allocation of storage resources.

[0054] The specific implementation of step S10 is to determine the distribution parameters of the storage optimization unit based on the deep optimization convolution kernel. First, calculate the load parameters of the storage nodes, including CPU usage, memory usage, network bandwidth usage, etc., and comprehensively consider various indicators using the weighted average method. Secondly, calculate the data redundancy parameter, which is the ratio of the number of data copies to the number of storage nodes. This step aims to provide a basis for storage optimization configuration.

[0055] The specific implementation of step S11 is to perform distributed node allocation for the storage optimization unit. Use the consistent hashing algorithm to map the storage space to the hash ring to ensure that only adjacent nodes are affected when nodes are added or deleted. Then, establish a storage address mapping table based on the storage space utilization rate and access latency parameters, and record the correspondence between the data identifier and the physical storage location in the form of key-value pairs. The purpose of this step is to achieve the balanced allocation of storage resources.

[0056] The specific implementation of step S12 is to perform sharding processing on the data according to the storage address mapping table. Use the range sharding method to divide continuous data into data blocks of equal size. Then, select appropriate storage nodes based on the storage node load parameters and data redundancy parameters, and allocate the data blocks to the nodes with lower load and meeting the redundancy requirements. This step aims to achieve the distributed storage of data.

[0057] The specific implementation of step S13 is to use the deep network of the storage optimization neural network to evaluate the storage location of the data blocks. First, collect the real-time operating status data of each node. Then, input the storage space utilization rate, access latency parameters, node load parameters, and data redundancy parameters into the deep network, and use the forward propagation algorithm to calculate the storage performance index. This index comprehensively considers various parameters through the weighted average method, and the weights are set according to the importance of the parameters. The purpose of this step is to evaluate the operation effect of the storage system.

[0058] The specific implementation of step S14 is to calculate the storage load factor based on the storage performance index. Use the standard deviation method to calculate the dispersion degree of the performance indexes of each node. The smaller the standard deviation, the more balanced the load. When the storage load factor is less than the storage load threshold, capacity expansion is required. The storage load threshold is set to 0.3. Specifically, it is achieved by adding storage nodes or expanding the capacity of existing nodes. This step aims to ensure the load balance of the storage system.

[0059] The specific implementation of step S15 is to optimize the storage location of the data blocks according to the storage load factor. When the load factor of a certain node is too high, migrate some data blocks to the nodes with lower load. The migration process adopts the incremental migration strategy to ensure that the system service is not interrupted. Through cyclic optimization and adjustment, the balanced storage of the industrial raw data stream is finally achieved. The purpose of this step is to maintain the stable operation of the storage system.

[0060] The data storage requirement index in the resource consumption calculation model is specifically expressed as follows:

[0061] DSI = α1T n + α2M n + α3P n + α4F n + β(T n M n ) 2 + γ(P n F n ) 2 + δ;

[0062] In the formula, DSI is the data storage requirement index; T n is the normalized device operation duration data; M n is the normalized device maintenance cycle data; P n is the normalized production task quantity data; F n is the normalized data generation frequency data; α1, α2, α3, α4 are linear term coefficients; β, γ are quadratic term coefficients; δ is the error correction term.

[0063] The data feature mapping matrix is specifically expressed as follows:

[0064]

[0065] In the formula, t ij is the j-th time series feature value of the i-th data item; s ij is the j-th spatial feature value of the i-th data item; r ij is the j-th correlation feature value of the i-th data item; m is the number of data items; n is the feature dimension.

[0066] The resource scheduling matrix is specifically expressed as follows:

[0067]

[0068] In the formula, μ is the mean of the feature mapping matrix; σ is the standard deviation of the feature mapping matrix.

[0069] The resource weight coefficient of the optimized scheduling convolution kernel is specifically expressed as follows:

[0070]

[0071] In the formula, w opt is the optimized weight coefficient; w sch is the weight of the resource scheduling convolution kernel; w base is the reference weight coefficient; λ is the attenuation coefficient, used to control the smoothness of weight adjustment.

[0072] The resource feature distribution map is specifically represented as follows:

[0073] F dist = Pool(Conv(M sch , K opt ));

[0074] In the formula, Conv represents the convolution operation; K opt is the optimized scheduling convolution kernel; Pool represents the max pooling operation.

[0075] The storage space utilization rate is specifically represented as follows:

[0076]

[0077] In the formula, U is the storage space utilization rate; S used is the used storage space; S total is the total storage space; T access is the data access time; η is the time decay factor.

[0078] The storage performance index is specifically represented as follows:

[0079]

[0080] In the formula, PI is the storage performance index; L node is the node load parameter; R data is the data redundancy parameter; ω1, ω2, ω3, ω4 are the weight coefficients.

[0081] The storage load factor is specifically represented as follows:

[0082]

[0083] In the formula, LF is the storage load factor; N is the number of storage nodes; PI i is the storage performance index of the i-th node; is the average storage performance index of all nodes.

[0084] Description of the parameter acquisition method:

[0085] 1. The normalized data is obtained through the maximum-minimum normalization method:

[0086] 2. The linear term coefficients α1, α2, α3, α4 and the quadratic term coefficients β, γ are obtained by least squares fitting;

[0087] 3. The value range of the error correction term δ is from 0.01 to 0.1;

[0088] 4. The value range of the attenuation coefficient λ is from 0.1 to 1.0;

[0089] 5. The value range of the time decay factor η is from 0.001 to 0.01;

[0090] 6. The weight coefficients ω1, ω2, ω3, ω4 are determined by the analytic hierarchy process and satisfy

[0091] Explanation of the equation construction principle:

[0092] 1. The data storage requirement exponential equation adopts a polynomial form, including a linear term and a quadratic term. The linear term represents the independent influence of each factor, the quadratic term represents the interaction between factors, and the error correction term is used to compensate for the model error;

[0093] 2. The data feature mapping matrix adopts a three-dimensional structure to uniformly represent temporal, spatial, and correlation features for subsequent processing;

[0094] 3. The resource scheduling matrix eliminates the influence of dimensions through standardization processing to improve the comparability of features;

[0095] 4. The weight coefficient equation of the optimized scheduling convolution kernel introduces an exponential decay term to ensure the smoothness and stability of weight adjustment;

[0096] 5. The storage space utilization equation considers the influence of access time and introduces an exponential term to represent the non-linear influence of time on utilization;

[0097] 6. The storage performance index equation adopts the form of a weighted sum, comprehensively considering multiple performance parameters, where the reciprocal form is used for access time and node load, indicating that the smaller these parameters are, the better;

[0098] 7. The storage load factor equation adopts the form of standard deviation and is used to measure the degree of balance of system load.

[0099] The following details the process of deriving some calculation processes.

[0100] Derivation process of the data storage requirement index DSI equation:

[0101] First, construct the basic linear model: DSI base = α1T n + α2M n + α3P n + α4F n ;

[0102] Considering the mutual influence between the device operation duration and the maintenance cycle, introduce the interaction term: (T n M n ) 2 ;

[0103] Considering the correlation between production tasks and data generation frequency, an interaction term is introduced: (P n F n ) 2 ;

[0104] Finally, an error correction term δ is added to obtain the complete equation.

[0105] The coefficients α1, α2, α3, α4 are obtained through training with historical data. The specific steps are as follows:

[0106] 1. Collect historical data for 6 months to establish a training set;

[0107] 2. Use the gradient descent method to minimize the mean square error;

[0108] 3. Determine the optimal coefficient values through cross-validation.

[0109] Derivation explanation of the data feature mapping matrix M map :

[0110] Calculation formula for the time series eigenvalue t ij :

[0111]

[0112] In the formula, x k is the data value at the k-th hour; w k is the time weight, satisfying

[0113] The spatial eigenvalue s ij is obtained through principal component analysis:

[0114] S = XV;

[0115] In the formula, X is the original data matrix; V is the eigenvector matrix;

[0116] s ij takes the first few principal components with the largest contribution rates in the S matrix.

[0117] Calculation formula for the correlation eigenvalue r ij :

[0118]

[0119] In the formula, x k , y k are the data pairs for which the correlation is to be calculated; is the corresponding mean value.

[0120] Standardization process explanation of the resource scheduling matrix M sch :

[0121] Calculation of the mean value μ:

[0122] Calculation of the standard deviation σ:

[0123] Optimize the weight coefficient w of the scheduling convolution kernel opt Derivation process:

[0124] 1. Initialize the weight coefficient:

[0125] 2. Set the reference weight: w base = 0.5;

[0126] 3. Introduce the distance penalty term: ||w sch - w base ||;

[0127] 4. Add exponential decay adjustment:

[0128] 5. Combine to obtain the final weight coefficient expression.

[0129] Resource feature distribution map F dist Supplementary explanation for the calculation:

[0130] Convolution operation:

[0131] Max pooling: Pool(x) = max{x 11 ,x 12 ,x 21 ,x 22}, and the pooling window size is 2×2.

[0132] Optimization process of the storage space utilization rate U equation:

[0133] 1. Basic utilization rate:

[0134] 2. Consider the influence of access time:

[0135] 3. Combine to obtain the final expression.

[0136] Access time T access Obtained through network probing, specifically by sending probe packets and recording the round-trip time.

[0137] Construction process of the storage performance index PI equation:

[0138] 1. Identify key performance parameters: storage space utilization rate, access time, node load, data redundancy;

[0139] 2. Use the reciprocal form for access time and node load:

[0140] 3. Introduce weight coefficients: ω1, ω2, ω3, ω4;

[0141] 4. Determine the weight values through the analytic hierarchy process.

[0142] Supplementary explanation for the derivation of the storage load factor LF equation:

[0143] 1. Calculate the average performance index:

[0144] 2. Calculate the variance of the performance index;

[0145] 3. Take the square root to obtain the standard deviation as the load factor.

[0146] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions run on a computer, they are used to execute the above-mentioned industrial data asset distributed storage method.

[0147] The third aspect of the present invention provides an industrial data asset distributed storage system, which includes the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.

[0148] Specifically, the principle of the present invention is: The technical principle of the present invention is based on the collaborative optimization idea of deep learning and distributed systems. By constructing a multi-level feature extraction and resource optimization mechanism, intelligent scheduling of storage resources is achieved. First, a resource consumption calculation model is used to extract features from industrial raw data streams. This model comprehensively considers the impacts of multiple dimensions such as equipment operation, maintenance cycle, task volume, and data frequency through a polynomial fitting equation, and can accurately reflect the demand characteristics of data for storage resources.

[0149] On the basis of feature extraction, a deep neural network is used for storage optimization. The network structure includes multiple convolutional layers and fully connected layers, and can automatically learn the deep feature representations of data. By optimizing the design of convolutional kernels, precise control of storage resource allocation is achieved. At the same time, a weight adjustment mechanism is introduced to ensure the smoothness and stability of resource scheduling. During the operation of the system, a complete performance evaluation and optimization feedback mechanism is established through the real-time calculation of storage performance indicators and load factors, ensuring the balanced allocation and efficient utilization of storage resources.

[0150] All technical links of the present invention have a strict logical relationship. From data feature extraction to resource optimization scheduling, and then to performance evaluation and optimization feedback, a complete technical closed-loop is formed. This solution makes full use of the advantages of deep learning in feature learning and optimization decision-making, and combines the actual needs of distributed systems, with strong theoretical basis and practical feasibility.

[0151] A specific embodiment 1 of the present invention is provided below. The specific implementation manners of each step in this embodiment 1 are described in detail as follows.

[0152] The specific implementation manner of step S01 is to collect device operation data through a data acquisition module deployed on manufacturing equipment. This module includes a time series database for storing device operation duration data, and records the device startup time, shutdown time, and cumulative operation duration in real time through a preset data acquisition frequency. The acquisition frequency is set to once every 1 minute; at the same time, a maintenance management system is configured to record the time intervals and completion status of maintenance activities such as regular maintenance and repair of the device. The acquisition frequency of maintenance data is set to be recorded immediately after each maintenance activity is completed; production task data such as the number of production task dispatch orders and task completion status is obtained in the production execution system. The acquisition frequency of task data is set to once every hour; the generation frequency of various types of data is obtained based on the data sampling period of the time series database. In the process of data acquisition, data preprocessing technologies are adopted, including data cleaning, outlier detection, and data completion. Among them, median filtering method is used for data cleaning to remove noise, 3σ criterion is used for outlier detection to identify abnormal data points, and linear interpolation method is used for data completion to process missing values. The purpose of this step is to establish the basic data source of industrial data assets.

[0153] The specific implementation manner of step S02 is to first perform normalization processing on various types of data in the industrial raw data stream, and uniformly map the numerical range to between 0 and 1. Specifically, it is achieved through the maximum-minimum normalization method. The normalization processing uses the following formula: In the formula, X n is the normalized data value, X is the original data value, X min is the minimum value of the data sequence, X max is the maximum value of the data sequence; then a data storage requirement index calculation model is constructed: DSI = α1T n +α2M n +α3P n +α4F n +β(T n Mn ) 2 +γ(P n F n ) 2 +δ. In the formula, DSI is the data storage requirement index, T nis the normalized device operation duration data, M n is the normalized device maintenance cycle data, P n is the normalized production task quantity data, F n is the normalized data generation frequency data. α1, α2, α3, α4 are linear term coefficients determined by the least squares method. β and γ are quadratic term coefficients optimized by the gradient descent method. δ is the error correction term with a value range of 0.01 to 0.1. This model uses a polynomial fitting equation, including linear terms and quadratic terms. The linear terms represent the independent effects of various factors, the quadratic terms represent the interaction between factors, and the error correction term is used to compensate for model errors. During the model training process, 6 months of historical data is used as the training set, and the 10-fold cross-validation method is used to evaluate the model performance. The finally obtained model can accurately evaluate the demand degree of industrial data assets for storage resources.

[0154] The specific implementation of step S03 is to first perform time series analysis on the industrial raw data stream using the resource consumption calculation model, and calculate the time series feature values using the sliding window method. The specific calculation formula is: In the formula, x k is the data value at the k-th hour, w k is the time weight, satisfying Secondly, spatial feature extraction is performed, and the spatial feature values are calculated using the principal component analysis method, which is specifically realized through eigenvalue decomposition: S = XV. In the formula, X is the original data matrix, and V is the eigenvector matrix; Thirdly, the correlation feature values between data are calculated using the Pearson correlation coefficient method, and the calculation formula is: In the formula, x k , y k is the data pair for which the correlation is to be calculated, is the corresponding mean value; Finally, these three types of feature values are organized into a three-dimensional feature mapping matrix: In the formula, t ij is the j-th time series feature value of the i-th data item, s ij is the j-th spatial feature value of the i-th data item, r ij is the j-th correlation feature value of the i-th data item, m is the number of data items, and n is the feature dimension. During the feature extraction process, a 24-hour sliding window is used for time series features, the principal components with a cumulative contribution rate reaching 85% are retained for spatial features, and the data is preprocessed by standardization when calculating correlation features.

[0155] The specific implementation of step S04 is to construct a resource scheduling matrix based on the data feature mapping matrix, and eliminate the influence of dimensions through standardization. The standardization formula is: In the formula, μ is the mean value of the feature mapping matrix, and the calculation formula is: σ is the standard deviation of the feature mapping matrix, and its calculation formula is: The standardized matrix can improve the comparability of features and eliminate the influence of different dimensions. Then, the resource scheduling matrix is input into a pre-trained storage optimization neural network. This network adopts a convolutional neural network structure, including an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a first fully connected layer, a second fully connected layer, and an output layer. The network training uses the backpropagation algorithm, and the training dataset is obtained by collecting data from the industrial site for 365 days. This step provides the basic data structure for subsequent storage optimization.

[0156] The specific implementation of step S05 is to first obtain the resource scheduling convolution kernel in the storage optimization neural network. The weight coefficients of the convolution kernel are randomly initialized by Gaussian distribution: Then, this convolution kernel is compared with the benchmark weight coefficient, which is determined by statistical analysis of historical data and is set to 0.5. Calculate the resource weight coefficient of the optimized scheduling convolution kernel. The specific calculation formula is: In the formula, w opt is the optimized weight coefficient, w sch is the weight of the resource scheduling convolution kernel, w base is the benchmark weight coefficient, and λ is the attenuation coefficient, with a value range of 0.1 to 1.0, which is used to control the smoothness of weight adjustment. The distance penalty term ||w sch - w base || and exponential decay adjustment are introduced in the calculation process of the weight coefficient to ensure the smoothness and stability of weight adjustment. The purpose of this step is to construct an optimized convolution kernel suitable for the current storage requirements.

[0157] The specific implementation of step S06 is to apply the optimized scheduling convolution kernel to the resource scheduling matrix and perform convolution operations in the form of a sliding window. The convolution calculation formula is: The window size is 3×3, and the stride is 1; then, perform a max-pooling operation on the convolution result: Pool(x) = max{x 11 , x 12 , x 21 , x 22}, the pooling window size is 2×2, and the resource feature distribution map F dit = Pool(Conv(M sch , K opt )) is obtained; calculate the storage space utilization rate based on this distribution map. The calculation formula is: In the formula, U is the storage space utilization rate, S used is the used storage space, S total is the total storage space, and T accessLet \(t\) be the data access time, \(\eta\) be the time decay factor, and the value range of \(\eta\) is from 0.001 to 0.01. The calculation of the storage space utilization rate takes into account the influence of the access time, and an exponential term is introduced to represent the non-linear influence of time on the utilization rate. This step aims to obtain the distribution status of the storage resources.

[0158] The specific implementation of step S07 is to perform feature extraction operations on the resource feature distribution map, calculate the weighted cumulative value of the resource weight coefficient and the feature value, adopt a linear weighting method, and the weights are set according to the importance of the features. Larger weights are assigned to features with higher importance; determine the space division parameters of the storage basic unit according to the weighted result, including the unit size and the distribution density; at the same time, adopt a network detection method to obtain the data access delay parameters of the storage basic unit, specifically by sending detection packets and recording the round-trip time. The access delay test of each storage unit is repeated 10 times and the average value is taken. During the feature extraction process, different weight coefficients are used for different types of features, and the calculation formula is: In the formula, \(FE\) is the feature extraction result, \(w\) i is the weight coefficient of the \(i\)-th feature, \(f\) i is the \(i\)-th feature value, and \(n\) is the number of features. The purpose of this step is to determine the basic configuration parameters of the storage system.

[0159] The specific implementation of step S08 is to find the deep convolutional kernel with a feature mapping relationship with the resource scheduling convolutional kernel in the deep network of the storage optimization neural network. First, calculate the cosine similarity between the convolutional kernels: In the formula, \(K1\) and \(K2\) are the convolutional kernels to be compared, \(k\) 1i , \(k\) 2i are the corresponding weight values; then select the deep convolutional kernel with the highest similarity as the deep optimization convolutional kernel, and the similarity threshold is set to 0.8, that is, only the convolutional kernels with a similarity greater than 0.8 are considered to have a mapping relationship. During the selection process of the deep convolutional kernel, the abstract level of the features is also considered. Deeper convolutional kernels can extract higher-level feature representations. This step aims to establish the correspondence between shallow features and deep features.

[0160] The specific implementation of step S09 is that when the ratio of the resource weight coefficient of the resource scheduling convolutional kernel to the reference weight coefficient exceeds the storage capacity threshold, the convolutional kernel weight needs to be adjusted, and the storage capacity threshold is set to 1.5. The weight adjustment adopts a linear proportional reduction method, and the reduction calculation formula is: In the formula, \(w\) new is the adjusted weight value, \(w\) old is the original weight value, \(T\) threshold is the storage capacity threshold, \(w\) ratiois the original weight ratio. A smoothing factor is introduced during the weight adjustment process to avoid sudden weight changes: w smooth = αw new + (1 - α)w old , where α is the smoothing coefficient, with a value range of 0.1 to 0.3, and w smooth is the weight value after smoothing. This step ensures that storage resources are not over-allocated through weight reduction, while the smoothing process guarantees the stability of weight adjustment.

[0161] The specific implementation of step S10 is to determine the distribution parameters of the storage optimization unit based on the deep optimization convolution kernel. First, calculate the load parameters of the storage nodes using the weighted average method: L node = ω1C usage + ω2M usage + ω3B usage , where L node is the node load parameter, C usage is the CPU usage rate, M usage is the memory usage rate, B usage is the network bandwidth usage rate, and ω1, ω2, ω3 are weight coefficients determined through the analytic hierarchy process; secondly, calculate the data redundancy parameter: where N replica is the number of data replicas, and N node is the number of storage nodes. Considering the load balance between nodes, a variance term is introduced: where V load is the load variance, L i is the load value of the i-th node, is the average load value. This step aims to provide a basis for the storage optimization configuration.

[0162] The specific implementation of step S11 is to perform distributed node allocation for the storage optimization unit. The consistent hashing algorithm is used to map the storage space onto the hash ring. Virtual node technology is adopted for hash calculation, and each physical node corresponds to multiple virtual nodes. The hash value calculation formula for virtual nodes is: h v = hash(h p + i), where h v is the virtual node hash value, h p is the physical node hash value, and i is the virtual node index. On this basis, a storage address mapping table is established by combining the storage space utilization rate and the access latency parameter. The scoring function for the mapping relationship is: where β1, β2 are weight coefficients, U is the storage space utilization rate, and T access is the access latency. The purpose of this step is to achieve the balanced allocation of storage resources.

[0163] The specific implementation of step S12 is to perform data sharding according to the storage address mapping table. The range sharding method is adopted, and the shard size is dynamically adjusted. The adjustment formula is: S chunk = S base (1 + γ·L ratio ), where S chunk is the shard size, S base is the basic shard size, which is set to 64MB, γ is the adjustment coefficient, and its value range is from 0 to 0.5, and L ratio is the load ratio. The selection of the data block storage location is based on a comprehensive scoring mechanism. The scoring calculation formula is: where λ1 and λ2 are weight coefficients, L node is the node load parameter, and R data is the data redundancy parameter. This step aims to achieve distributed storage of data and ensure storage efficiency through dynamic sharding and the scoring mechanism.

[0164] The specific implementation of step S13 is to use the deep network of the storage optimization neural network to dynamically evaluate the data block storage location. First, collect the real-time operating status data of the storage nodes, including storage space utilization rate, access latency parameter, node load parameter, and data redundancy parameter. The evaluation process uses the forward propagation algorithm to calculate the storage performance index. The index calculation formula is: where PI is the storage performance index, U is the storage space utilization rate, T access is the access latency parameter, L node is the node load parameter, R data is the data redundancy parameter, ω1, ω2, ω3, and ω4 are weight coefficients, and they satisfy The weight coefficients are determined by the analytic hierarchy process, and a judgment matrix is constructed: The weight values are calculated by the eigenvalue method. The purpose of this step is to evaluate the operation effect of the storage system.

[0165] The specific implementation of step S14 is to calculate the storage load factor based on the storage performance index, and use the standard deviation method to calculate the dispersion degree of the performance indexes of each node. The calculation formula is: where LF is the storage load factor, N is the number of storage nodes, PI i is the storage performance index of the i-th node, is the average storage performance index of all nodes. When the storage load factor is less than the storage load threshold, capacity expansion is required. The storage load threshold is set to 0.3. The capacity expansion adopts a dynamic threshold adjustment strategy. The threshold adjustment formula is: where T new is the new threshold, and T oldis the original threshold value, δ is the adjustment coefficient, and its value range is from 0.1 to 0.3. LF base is the reference load factor, which is set to 0.5. This step aims to ensure the load balance of the storage system.

[0166] The specific implementation of step S15 is to optimize the storage location of data blocks according to the storage load factor. When the node load factor is too high, data migration is started. The migration priority calculation formula is: In the formula, P migrate is the migration priority, LF i is the current node load factor, LF avg is the average load factor, S free is the remaining storage space of the target node, S total is the total storage space, μ1 and μ2 are weight coefficients. The incremental migration strategy is adopted during the migration process. The calculation formula for the amount of data migrated each time is: In the formula, S migrate is the amount of data migrated, S block is the data block size, LF target is the target load factor, P migrate is the migration time, and κ is the time decay coefficient. This step realizes the load optimization of the storage system through dynamic data migration, and finally achieves the balanced storage of industrial raw data streams.

[0167] To better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention: A manufacturing enterprise has 50 numerically controlled machine tools. Before implementing the method of the present invention, a traditional centralized storage system is used to manage industrial data. As the amount of data increases, the system performance gradually deteriorates, the storage resource utilization is unbalanced, and the data access latency increases significantly. The research team decides to use the method of the present invention to transform the system. The specific implementation process is as follows.

[0168] First, the research team collects the operation data of manufacturing equipment. The collection period is 1 month. The statistical results of equipment operation duration data, maintenance cycle data, production task quantity data, and data generation frequency data are shown in Table 1:

[0169] Table 1 Statistical table of equipment operation data

[0170] Data type Minimum value Maximum value Average value Standard deviation Running duration (hours / day) 16.5 23.8 20.4 1.8 Maintenance period (days) 25 45 35 5.2 Number of production tasks (pieces / day) 85 168 126 22.5 Data generation frequency (pieces / minute) 120 360 240 60

[0171] Figure 2Shows the changing trends of the device operation data within 31 days, including the dynamic changes of three indicators: daily operation duration, maintenance cycle, and production task volume. The collected data is normalized to obtain the normalized data, and then the data storage requirement index is calculated according to the resource consumption calculation model. After fitting by the least squares method, the model parameters are determined: α1 = 0.35, α2 = 0.25, α3 = 0.20, α4 = 0.20, β = 0.15, γ = 0.12, δ = 0.05. The distribution of the data storage requirement index of each device is shown in Table 2 as follows:

[0172] Table 2 Distribution Table of Data Storage Requirement Index

[0173]

[0174]

[0175] Figure 3 Shows the distribution of the storage requirement index of 50 devices, and intuitively displays the storage requirement intensity of different devices through the depth of color. Based on the results of data feature analysis, the research team constructs a data feature mapping matrix, performs time series analysis using a 24-hour sliding window, and sets the time weight decay coefficient to 0.042. The spatial features are extracted by the principal component analysis method, and the principal components with a cumulative contribution rate reaching 85% are selected. The main eigenvalue obtained is shown in Table 3 as follows:

[0176] Table 3 Analysis Table of Data Eigenvalues

[0177] Feature type Feature value Contribution rate Time series feature 1 0.384 0.35 Time series feature 2 0.256 0.23 Spatial feature 1 0.215 0.18 Spatial feature 2 0.168 0.15 Correlation feature 1 0.125 0.09

[0178] Figure 4 Shows the analysis results of the contribution rate of each eigenvalue. The bar chart represents the size of each eigenvalue, and the line chart represents the cumulative contribution rate. The feature mapping matrix is input into the storage optimization neural network, and the network structure includes 3 convolutional layers and 2 fully connected layers. The training data set contains 365 days of historical data, the batch size is set to 64, the learning rate is 0.001, and the number of training epochs is 200 rounds. The performance indicators of the trained network on the test set are shown in Table 4 as follows:

[0179] Table 4 Neural Network Performance Evaluation Table

[0180] Evaluation metric Training set Test set Accuracy 0.945 0.923 Recall 0.932 0.908 F1 score 0.938 0.915 Loss value 0.058 0.075

[0181] Based on the trained neural network, calculate the resource weight coefficient of the optimized scheduling convolution kernel, set the decay coefficient λ to 0.5, and the benchmark weight coefficient to 0.5. Apply the optimized scheduling convolution kernel to the resource scheduling matrix, and the obtained storage resource distribution is shown in Table 5 as follows:

[0182] Table 5 Storage Resource Distribution Table

[0183] Storage node Space utilization rate Access latency (ms) Load parameter Redundancy Nodes 1 - 5 0.75 25 0.68 2.0 Nodes 6 - 10 0.68 28 0.62 2.0 Nodes 11 - 15 0.62 30 0.55 1.5

[0184] The weight coefficients of the storage performance metrics are determined by the analytic hierarchy process: ω1 = 0.3, ω2 = 0.25, ω3 = 0.25, ω4 = 0.2. The performance evaluation results after the system has run for one month are shown in Table 6 as follows:

[0185] Table 6 System Performance Evaluation Table

[0186] Performance metric Before optimization After optimization Average response time (ms) 85 28 Storage space utilization rate 0.45 0.68 Load balancing degree 0.58 0.85 System throughput (MB / s) 256 512

[0187] During the operation of the system, when the detected storage load factor is less than the threshold of 0.3, the capacity expansion mechanism is automatically triggered. Figure 5 It shows the comparison of various performance metrics before and after system optimization, including average response time, storage space utilization rate, load balancing degree, and system throughput. The system has undergone 3 capacity expansions and 15 data block migration operations within one month. The relevant parameters of the migration process are shown in Table 7:

[0188] Table 7 Data Migration Parameter Table

[0189] Migration batch Migrated data volume (GB) Migration time (min) Impact on business time (s) Batch 1 256 45 0 Batch 2 384 62 0 Batch 3 512 85 0

[0190] Compared with the traditional distributed storage scheme, the scheme of the present invention has significant advantages: the traditional scheme adopts a fixed data sharding strategy and a static resource allocation method, and when facing dynamically changing industrial data, problems such as storage hotspots and resource waste often occur. While the present invention realizes the intelligent scheduling of storage resources through deep learning technology, and the system can adaptively adjust the resource allocation strategy according to data characteristics. The experimental results show that the scheme of the present invention has achieved significant improvements in aspects such as storage space utilization rate, system response time, and load balancing degree. In addition, the scheme of the present invention has also achieved zero service interruption during the data migration process, ensuring the continuous availability of the system. Through the resource optimization mechanism driven by deep learning, the present invention has successfully solved the problems of intelligent scheduling and load balancing in the distributed storage system of industrial data assets.

[0191] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Table 8 below.

[0192] Table 8 Variable Explanation Table

[0193]

[0194]

[0195] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. A distributed storage method for industrial data assets, characterized in that: The following steps are involved: Collect manufacturing equipment operation time data, manufacturing equipment maintenance cycle data, production task quantity data and data generation frequency data as industrial raw data streams; Establish a resource consumption calculation model based on the industrial raw data stream; use the resource consumption calculation model to extract the data time series eigenvalues, data space eigenvalues ​​and data association eigenvalues ​​of the industrial raw data stream, and construct a data feature mapping matrix; Constructing a resource scheduling matrix based on the data feature mapping matrix, and inputting the resource scheduling matrix into the storage optimization neural network; obtaining an optimized scheduling convolution kernel according to the resource scheduling convolution kernel of the storage optimization neural network, wherein the resource weight coefficient of the optimized scheduling convolution kernel is obtained based on the ratio of the corresponding weight of the resource scheduling convolution kernel to the reference weight coefficient; Applying the optimized scheduling convolution kernel to the resource scheduling matrix to obtain a resource feature distribution map; Obtaining a deep optimized convolution kernel in a deep network of the storage optimized neural network that has a mapping relationship with the resource scheduling convolution kernel; When the ratio of the resource weight coefficient of the resource scheduling convolution kernel to the reference weight coefficient is greater than a storage capacity threshold, a proportional reduction processing is performed on the corresponding weight of the resource scheduling convolution kernel.

2. The distributed storage method for industrial data assets according to claim 1, characterized in that: The resource consumption calculation model normalizes the manufacturing equipment operation time data, the manufacturing equipment maintenance cycle data, the production task quantity data and the data generation frequency data, and calculates the data storage demand index through a polynomial fitting equation. The data storage demand index is used to characterize the storage resource consumption level of the industrial raw data stream.

3. The distributed storage method for industrial data assets according to claim 1, characterized in that: The storage optimization neural network consists of an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a first fully connected layer, a second fully connected layer and an output layer. The training data set of the storage optimization neural network is obtained by collecting the data block storage location, storage space utilization, data access delay parameters, storage node load parameters and data redundancy parameters of the industrial site for 365 days. The storage optimization neural network is trained using a back propagation algorithm, and the network parameters are optimized and adjusted using the training data set.

4. The distributed storage method for industrial data assets according to claim 1, characterized in that: A feature extraction operation is performed on the resource feature distribution map to obtain a space division parameter of a storage basic unit, and a data access delay parameter of the storage basic unit is obtained, wherein the feature extraction operation includes a weighted accumulation operation on the resource weight coefficient and the feature value of the resource feature distribution map.

5. The distributed storage method for industrial data assets according to claim 1, characterized in that: Determine the distribution parameters of the storage optimization unit based on the deep optimization convolution kernel, wherein the distribution parameters include storage node load parameters and data redundancy parameters; Distributed node allocation is performed on the storage optimization unit, and a storage address mapping table is established based on the storage space utilization and the data access delay parameter.

6. The distributed storage method for industrial data assets according to claim 5, characterized in that: A data slicing operation is performed according to the storage address mapping table, and a data block storage location is determined based on the storage node load parameter and the data redundancy parameter.

7. The distributed storage method for industrial data assets according to claim 6, characterized in that: The deep network of the storage optimization neural network is used to dynamically evaluate the data block storage location, and a storage performance indicator is generated based on the storage space utilization, the data access delay parameter, the storage node load parameter and the data redundancy parameter.

8. The industrial data asset distributed storage method according to claim 7, characterized in that: A storage load factor is calculated based on the storage performance indicator. When the storage load factor is less than a storage load threshold, the capacity of the data block storage location is expanded. The data block storage location is optimized and adjusted according to the storage load factor to complete the distributed storage of the industrial raw data stream.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the industrial data asset distributed storage method according to any one of claims 1 to 8.

10. An industrial data asset distributed storage system, characterized in that: The system comprises the computer-readable storage medium as claimed in claim 9, wherein the system is any one of a computer, a server, and a single-chip microcomputer, the computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

Citation Information

Patent Citations

  • AI-based big data distributed computing task automatic optimization method and system

    CN119576507A

  • Method for dividing processing capabilities of artificial intelligence between devices and servers in network environment

    US20220207327A1

Cited By

  • Lightweight time sequence prediction method and device fused with dynamic forgetting in edge scene

    CN120781199A

  • A light-weight time series prediction method and device fusing dynamic forgetting in edge scenes

    CN120781199B