Data storage method for credential terminal
By establishing data circulation and genetic models on Xinchuang terminals, analyzing data flow direction and importance, determining the number of backups and scattering storage, the security and traceability problems in Xinchuang terminal data storage are solved, and efficient and secure data storage is achieved.
Patent Information
- Application Number
- CN202510742072.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing technology has insufficient security and weak traceability in the data storage of Xinchuang terminals, which poses a risk of data leakage, and redundant backups may lead to significant losses in the leakage of all important data.
The idea of genetic algorithm is used to establish data circulation models and genetic models. By analyzing the flow direction and importance of data, we determine the number of backups of the data set, and store data in different backup units to maximize the optimization and reduce data storage costs.
It improves data traceability and storage efficiency, reduces the possibility of data leakage, and optimizes data storage costs.
Smart Images

Figure CN120256207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information technology application innovation. More specifically, the present invention relates to a data storage method for information technology application innovation terminals. Background Art
[0002] The information technology application innovation industry is a key area for achieving the independent control of information technology. With the development of operating systems, software, and hardware related to the information technology application innovation industry, enterprises have started to use information technology application innovation terminals as tools for enterprise operation and production. For example, data on workshop production and equipment operation are monitored through information technology application innovation terminals. At the same time, network security issues are increasing. Therefore, when using information technology application innovation terminals for data storage, it is necessary to ensure the security of the data, thereby promoting the healthy development of the information technology application innovation industry and the stable progress of enterprise production.
[0003] In order to ensure the security and effective storage of stored data, hierarchical storage and redundant backup methods are commonly used at present. In related technologies, for example, a Chinese patent document with the authorization announcement number CN103558991B discloses a data hierarchical storage processing method, device, and storage device, which disclose dividing data into cold data and hot data by detecting the data access frequency of migration units, so that only hot data is stored on high-level disks, controlling the consumption of data storage resources. For example, a Chinese patent document with the authorization announcement number CN11714952B discloses a server data backup and recovery system and method, which disclose flexibly adjusting the backup plan according to factors such as data importance and update frequency, which can improve the data backup efficiency and reduce resource occupancy.
[0004] However, in the process of data storage and backup, dividing and storing data according to access frequency and data importance lacks an analysis of the relationship between data, resulting in a slow speed when tracing important data or hot data. In addition, although increasing the number of backups will increase the disaster tolerance ability, it may also increase the possibility of data leakage. When important data or hot data is backed up to one place, it may lead to the complete leakage of important data or hot data, causing heavy losses. Summary of the Invention
[0005] To solve the above technical problems of insufficient security and weak data tracing ability during data storage, the present invention provides a data storage method for information technology application innovation terminals, including: Obtain several data sets of Xinchuang terminals. The data sets are collections of data of a single item at different time periods. Obtain the metadata of all data in each data set. The metadata includes the detection time, data source, data flow direction, and data size of the data. Based on the data source and data flow direction of each data, establish a data circulation model. Based on the genetic algorithm idea and the data circulation model, establish a data genetic model. According to the genetic performance of each data in the data genetic model, obtain the importance of each data set. Based on the retrieval records of each data and the linear fitting method, obtain the heat change model of each data set. According to the importance of the data set and the heat change model of the data set, obtain the number of copies that each data set should be backed up. According to the number of copies that each data set should be backed up, establish a backup node sequence. Randomly place all nodes of the backup node sequence in all backup units to obtain several backup schemes. According to the genetic performance of the nodes in all backup units in the data genetic model in each backup scheme, obtain the goodness of each backup scheme. Store the data with the backup scheme with the maximum goodness.
[0006] The present invention analyzes the flow direction of various data of an enterprise through a data circulation model, thereby being able to improve the traceability of data and the efficiency of extracting stored data. The present invention further analyzes the flow direction between data by adopting the genetic algorithm idea and establishes a data genetic model, thereby being able to obtain the genetic performance and importance of data, and thus being able to more accurately determine the number of copies required for data backup and reduce the data storage cost.
[0007] Preferably, the establishment of the data circulation model includes: taking any data as a node, connecting any two nodes according to the data source and data flow direction relationship existing between any two nodes; recording the distance between any two nodes as 1, taking the nodes with an in-degree of 0 in the data circulation graph as starting points respectively and performing a depth-first search algorithm to obtain the shortest distance from each starting point to each node, which is recorded as the level of each node. When the level of a node is not unique, the minimum value of the level of the node is recorded as the level of the node, and all nodes are connected according to the level to obtain the data circulation model.
[0008] Preferably, the connection of any two nodes includes: for the i-th node and the c-th node, if the data corresponding to the c-th node is the data source of the data corresponding to the i-th node, then connect the i-th node and the c-th node, and the direction points to the i-th node; if the data corresponding to the c-th node is the data flow direction of the data corresponding to the i-th node, then connect the i-th node and the c-th node, and the direction points to the c-th node.
[0009] Preferably, the establishment of the data genetic model includes: regarding each data as a 01 string of the same size as the data size, denoted as the corresponding chromosome of each data; obtaining the genetic probability sequence of each data set based on the data flow of each data set reflected by the data flow model; obtaining the selection operator from node to node and the mutation operator from data set to data set based on the data sizes of the data corresponding to different nodes and the data flow model; the data genetic model is expressed as that the corresponding chromosome of the data corresponding to the node at the e-th level is genetically inherited according to the genetic probability sequence of the corresponding data set to obtain the node at the e+1-th level, and the corresponding chromosome of the data corresponding to the node at the e+1-th level is represented by the product of the corresponding chromosome of the data corresponding to the node at the e-th level, the selection operator, and the mutation operator.
[0010] Preferably, the obtaining of the genetic probability sequence of each data set includes: regarding any data set as the target data set, obtaining the frequency distribution sequence of the data flow directions of all the data in the target data set, denoted as the genetic sequence of the target data set, and dividing each value of the genetic sequence of the target data set by the number of data in the target data set to obtain the genetic probability sequence of the target data set.
[0011] The present invention judges the data flow direction from a relatively macroscopic perspective by obtaining the genetic probability of the data set, thereby avoiding the influence of adding or deleting data on the data genetic model between data and increasing the applicable range of the data genetic model.
[0012] Preferably, the obtaining of the selection operator from node to node includes: If the b-th node is the data source of the a-th node, the selection operator of the a-th node for the b-th node satisfies the expression: ; In the formula, represents the selection operator of the a-th node for the b-th node; 、 represent the data sizes of the data corresponding to the a-th node and the b-th node; 、 represent the data size sets of the data sets corresponding to the data corresponding to the a-th node and the b-th node; represents the maximum value function; represents the normalization function.
[0013] The present invention analyzes the specific influence between data through the selection operator. When the selection operator between data is small, although there is data flow, the changes in the lower-level data have little impact on the higher-level data.
[0014] Preferably, the obtaining of the mutation operator from data set to data set includes: If the b-th node is the data source of the a-th node, the mutation operator of the corresponding data set of the a-th node with respect to the corresponding data set of the b-th node satisfies the expression: ; In the formula, represents the mutation operator of the corresponding data set of the a-th node with respect to the corresponding data set of the b-th node; and represent the data size sets of the corresponding data sets of the a-th node and the b-th node; represents the maximum value function; represents the edit distance function; represents the normalization function.
[0015] The present invention measures the relationship between data sets through a mutation operator. When the mutation operator is large, even if the selection operator between data is large, the impact of changes in low-level data on high-level data is small.
[0016] Preferably, obtaining the importance of each data set includes: taking each node at the first level as a starting point, traversing the data flow model using a depth-first search algorithm to obtain all possible paths; for the node sequence of the m-th possible path of the b-th node and the a-th node , multiply the selection operator of any two adjacent nodes in by the mutation operator of the corresponding data sets of the two adjacent nodes, and multiply them cumulatively to obtain the importance of the b-th node to the a-th node in . Add the importance of the b-th node to the a-th node in all possible paths of the b-th node and the a-th node to obtain the importance of the b-th node to the a-th node; for the d-th data set, obtain the importance of the nodes corresponding to the data contained in the d-th data set to all other nodes, and sum them to obtain the importance of the d-th data set.
[0017] Preferably, the number of copies to be made for each data set satisfies the expression: ; In the formula, represents the number of copies to be made for the d-th data set; and represent the minimum and maximum values of the pre-set copy number range; represents the importance of the d-th data set; represents the slope of the heat change model of the d-th data set; represents the ceiling function.
[0018] Based on the importance of the data sets and the heat change model of the data sets, the present invention determines the number of backups for each data set, so as to make the backup data more effective and reduce the data storage cost.
[0019] Preferably, obtaining the goodness of each backup scheme includes: Taking the negative correlation normalization result of the importance of the b-th node to the a-th node as the distance between the b-th node and the a-th node; The goodness of any backup scheme satisfies the expression: ; In the formula, represents the goodness of the t-th backup scheme; represents the number of backup units; represents the distance between the u-th node and the v-th node in the s-th backup unit of the t-th backup scheme; represents the number of nodes in the s-th backup unit of the t-th backup scheme.
[0020] The beneficial effects of the present invention are as follows: (1) The present invention uses the genetic algorithm idea to analyze the relationship between the data of the information technology innovation terminal, improves the traceability of the data, can measure the impact caused by each data, and ensures the analysis of the relationship between the data; (2) The present invention selects the backup scheme with the maximum goodness, avoids storing strongly associated data in the same storage unit, resulting in greater losses, and reduces the possibility of data leakage. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart schematically showing a data storage method for an information technology innovation terminal in the present invention; Figure 2 is a schematic diagram schematically showing a data flow diagram; Figure 3 is a schematic diagram schematically showing a data flow model; Figure 4 is a schematic diagram schematically showing the nodes of the data flow model; Figure 5 is a schematic diagram schematically showing the data inheritance process. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] An embodiment of the present invention discloses a data storage method for an information technology innovation terminal. Referring to Figure 1 , it includes steps S1 - step S4: S1: Obtain a number of data sets of the information technology innovation terminal. The data set is a set of data of an item at different time periods, and obtain the metadata of all the data in each data set.
[0023] It should be noted that taking a manufacturing enterprise as an example, the departments involving data mainly include production, logistics, warehousing, sales, quality, equipment, supply chain, etc. In order to analyze the relationship between data among departments and store and back up the data of the enterprise's domestic innovation terminals at the overall level, it is necessary to obtain various data more comprehensively.
[0024] Specifically, set a time period, and collect various data of the domestic innovation terminals according to the time period in the domestic innovation environment to obtain a number of data sets. Any data set is a collection of data of a certain item at different time periods. Establish a metadata dictionary for the enterprise, and obtain the metadata of all data in each data set.
[0025] The domestic innovation environment is a domestic innovation data collection, transmission, and storage environment composed of domestic innovation chips, domestic innovation encryption systems, domestic innovation sensors, etc. The metadata dictionary is a standardized management tool for uniformly naming metadata, providing a basis for data collectors or data creators to fill in the metadata of data. The metadata includes generation time, data source, data flow, retrieval record, data size, etc. The retrieval record is the domestic innovation terminal that retrieves data from the pre-stored database and the corresponding retrieval time. It should be noted that taking the product quality data in the production process as an example, the metadata of the product quality data includes inspection time, data source, data flow, data size, etc. The data source is the various parameters of the product, and the data flow is the subsequent data output by departments such as quality and supply chain with reference to the product quality data.
[0026] So far, a number of data of the domestic innovation terminals and the metadata of each item of data have been obtained.
[0027] S2: Based on the data source and data flow of each data, establish a data circulation model.
[0028] It should be noted that taking a manufacturing enterprise as an example, there are various data circulation methods among departments. For example, the supply chain and the procurement department provide raw material data to the production department, the production department provides product data to the warehousing department, the sales department provides order data to the warehousing department, and the warehousing department completes order delivery. Therefore, the present invention establishes a data circulation model, so as to clarify the logical relationship between data and provide a basis for hierarchical storage and redundant backup of data.
[0029] Specifically, take any data as a node, and connect any two nodes. The connection process is as follows: for the i-th node and the c-th node, if the data corresponding to the c-th node is the data source of the data corresponding to the i-th node, then connect the i-th node and the c-th node, and the direction points to the i-th node; if the data corresponding to the c-th node is the data flow of the data corresponding to the i-th node, then connect the i-th node and the c-th node, and the direction points to the c-th node.
[0030] Complete the connection of all nodes to obtain a data flow graph, as shown in Figure 2 It is a schematic diagram of the data flow graph.
[0031] It should be noted that since the flow of nodes is relatively complex and the flow direction of a node is not unique, the data flow graph of an enterprise will form a complex network. To manage the nodes more clearly, the present invention classifies each node in the data flow graph.
[0032] Preferably, the distance between any two nodes is recorded as 1, and the nodes with an in-degree of 0 in the data flow graph are used as starting points to perform a depth-first search algorithm respectively to obtain the shortest distances from each starting point to each node, which are recorded as the levels of each node. When the levels of a node are not unique, the minimum value of the levels of the node is recorded as the level of the node. All nodes are connected according to the levels to obtain a data flow model, as shown in Figure 3 It is a schematic diagram of the data flow model.
[0033] So far, the data flow model has been obtained.
[0034] S3: Based on the genetic algorithm idea and the data flow model, establish a data genetic model; according to the genetic performance of each data in the data genetic model, obtain the importance of each data set; based on the retrieval records of each data and the linear fitting method, obtain the heat change model of each data set; according to the importance of the data set and the heat change model of the data set, obtain the number of copies to be backed up for each data set.
[0035] It should be noted that for the data corresponding to the nodes with lower levels, if there is a missing value, it will affect the data corresponding to all the nodes in its flow direction, thus causing a greater impact. Therefore, the data corresponding to the nodes with lower levels is more important. However, the data corresponding to the nodes with lower levels often has a large amount of data. For example, the various parameters of products in the production process. In subsequent production, the specific data of a product is often not retrieved again, but the product qualification rate extracted from the various parameters of the product is used to optimize the operation of the production machine and the use of raw materials. Therefore, although the data corresponding to the nodes with lower levels is of high importance, the retrieval frequency is low. If a high redundancy is used for backup, a large amount of data storage costs will be generated. Therefore, the present invention adaptively determines the redundancy backup amount for each data in combination with the importance of the data and the usage of the data.
[0036] It should be noted that, as shown in Figure 4 It is a schematic diagram of the nodes of the data flow model. Node a and node b are both nodes at the 3rd level. However, the data flow of the data corresponding to node b is more, affecting the remaining data at the same level. Therefore, the importance of node b is higher. Therefore, the importance of each data should also be determined in combination with the influence range of the data.
[0037] It should be noted that the data flow process is not only a simple selection from low-level data, but also involves a large amount of data processing, such as denoising operations. Therefore, the data flow process can be regarded as a process of data selection and variation. When the data sources are not unique, there is also crossover. Therefore, data can be regarded as chromosomes, and the data flow process is the genetic process of chromosomes. Therefore, in this invention, the influence of the data flow is analyzed through the idea of genetic algorithms.
[0038] Specifically, a data genetic model is established: Each data is regarded as a 01 string of the same size as the data, denoted as the corresponding chromosome of each data.
[0039] Any data set is regarded as the target data set, and the frequency distribution sequence of the data flows of all data in the target data set is obtained, denoted as the genetic sequence of the target data set. Divide each value in the genetic sequence of the target data set by the number of data in the target data set to obtain the genetic probability sequence of the target data set. It should be noted that, for example, if the data volume of the target data set is 10, and the data flows of 2 data are to the d-th data set, and the data flows of 3 data are to the (d + 1)-th data set, then the genetic sequence of the target data set is , where the serial number corresponding to 2 is d, and the serial number corresponding to 3 is d + 1.
[0040] If the b-th node is the data source of the a-th node, then the selection operator of the a-th node for the b-th node satisfies the expression: ; In the formula, represents the selection operator of the a-th node for the b-th node; , represent the data sizes of the data corresponding to the a-th node and the b-th node; , represent the data size sets of the data sets corresponding to the data corresponding to the a-th node and the b-th node; represents the maximum value function; represents the normalization function.
[0041] In the formula, , represent the relative data sizes of the data corresponding to the a-th node and the b-th node in the corresponding data sets, represents the difference between the relative data sizes of the data corresponding to the a-th node and the b-th node in the corresponding data sets, representing the relative data size difference between the data corresponding to the a-th node and the b-th node. The smaller this value is, the less part of the data corresponding to the a-th node is selected from the data corresponding to the b-th node, thus indicating that the selection operator of the a-th node for the b-th node is smaller.
[0042] It should be noted that due to data mutation, the selection operator cannot fully reflect data circulation. Therefore, the data size sets of the corresponding data sets of the a-th node and the corresponding data sets of the b-th node are analyzed. If the changes in the data sizes of the two data sets are relatively consistent, the data mutation operators of the two data sets are smaller, and thus the selection operator is more accurate.
[0043] The mutation operator of the corresponding data set of the a-th node with respect to the corresponding data set of the b-th node satisfies the expression: ; In the formula, represents the mutation operator of the corresponding data set of the a-th node with respect to the corresponding data set of the b-th node; , represent the data size sets of the corresponding data sets of the a-th node and the b-th node corresponding data; represents the maximum value function; represents the edit distance function; represents the normalization function.
[0044] In the formula, represents the edit distance of the relative data size sets of the corresponding data sets of the a-th node and the b-th node corresponding data. The larger this value is, the weaker the consistency of the data size changes between the two data sets, and thus the larger the mutation operator of the corresponding data set of the a-th node with respect to the corresponding data set of the b-th node.
[0045] The data genetic model is expressed as: The corresponding chromosome of the data of the e-th layer node is genetically inherited according to the genetic probability sequence of the corresponding data set to obtain the node of the e + 1-th layer, and the corresponding chromosome of the data of the e + 1-th layer node is represented by the product of the corresponding chromosome of the data of the e-th layer node and the selection operator and the mutation operator.
[0046] Thus, a data genetic model is established.
[0047] It should be noted that as Figure 5 is a schematic diagram of the data genetic process. In the data genetic model, the corresponding chromosomes of the lower-level data are continuously genetically inherited through selection and mutation to form the corresponding chromosomes of the higher-level data. Then, the more the genetic part of the corresponding chromosome of the data is, the wider the influence range of the data and the stronger the importance of the data.
[0048] Preferably, according to the genetic performance of each data in the data genetic model, the importance of each data set is obtained: Taking each node at the first level as a starting point, traverse the data circulation model using the depth-first search algorithm to obtain all possible paths; obtain all possible paths between any two nodes.
[0049] It should be noted that all possible paths between the a-th node and the b-th node reflect the data inheritance between the a-th node and the b-th node. Among the possible paths from the a-th node to the b-th node, the larger the selection operator and the smaller the mutation operator, the stronger the genetic performance of the b-th node to the a-th node, and the more important the b-th node is. When all the data in the dataset to which the data corresponding to the b-th node belongs have strong importance, it indicates that the dataset to which the data corresponding to the b-th node belongs has stronger importance.
[0050] For the node sequence of the m-th possible path between the b-th node and the a-th node , multiply the selection operators of any two adjacent nodes in by the mutation operators of the corresponding datasets of the two adjacent nodes, and multiply them cumulatively to obtain the importance of the b-th node to the a-th node in
[0051] For the d-th dataset, obtain the importance of the nodes corresponding to the data contained in the d-th dataset to all the other nodes, and sum them up to obtain the importance of the d-th dataset.
[0052] Thus far, the importance of each dataset has been obtained.
[0053] Preferably, based on the retrieval records of each data and the linear fitting method, obtain the heat change model of each dataset: Obtain the retrieval times of each data in each time period; subtract the generation time of each data from the start time of each time period as the relative retrieval time of each data; for all the data in the d-th dataset, establish a coordinate system with the relative retrieval time as the horizontal axis and the retrieval times of the data in the corresponding time period as the vertical axis, plot the scatter diagram of the relative retrieval time - retrieval times of all the data in the d-th dataset, and use the least squares method for linear fitting; take the obtained fitting line as the heat change model of the d-th dataset.
[0054] Thus far, the heat change model of each dataset has been obtained.
[0055] It should be noted that in the heat change model of the dataset, as the relative retrieval time increases and the retrieval times decrease, it indicates that the number of backups required for the dataset is less. At the same time, the weaker the importance of the dataset, the less the number of backups required for the data.
[0056] Preferably, according to the importance of the data sets and the heat change model of the data sets, obtain the number of backups that should be made for each data set: ; In the formula, represents the number of backups that should be made for the d-th data set; , represent the minimum and maximum values of the preset backup quantity range; represents the importance of the d-th data set; represents the slope of the heat change model of the d-th data set; represents the ceiling function. It should be noted that the preset backup quantity range is set by the implementer according to the actual implementation situation. For example, , can be set to 2 and 10.
[0057] Thus, the number of backups that should be made for each data set is obtained.
[0058] S4: Combine the number of backups that should be made for each data set and the data inheritance model to store the data.
[0059] It should be noted that in order to avoid batch loss of data caused by storing data with strong genetic performance together, data with strong genetic performance can be stored in different backup units according to the data inheritance model, so as to ensure the safety of the backup units.
[0060] Specifically, establish a backup node sequence. It should be noted that taking nodes a, b, and c as an example, if the backup quantities of the data corresponding to nodes a, b, and c are 2, 3, and 4, then the backup node sequence of nodes a, b, and c is .
[0061] Use the negative correlation normalization result of the importance of the b-th node to the a-th node as the distance between the b-th node and the a-th node.
[0062] It should be noted that in order to maximize the distance between the data of all the overall nodes of all backup units, nodes with smaller distances need to be backed up in different backup units.
[0063] Randomly place all the nodes in the backup node sequence in all backup units to obtain several backup schemes. The goodness of any backup scheme satisfies the expression: ; In the formula, represents the goodness of the t-th backup scheme; represents the number of backup units; represents the distance between the u-th node and the v-th node in the s-th backup unit of the t-th backup scheme; Represents the number of nodes of the s-th backup unit of the t-th backup scheme.
[0064] Wherein, Represents adding the distances of all nodes in the s-th backup unit of the t-th backup scheme as the overall difference of the s-th backup unit of the t-th backup scheme. Represents the average overall difference of all backup units of the t-th backup scheme. The larger this value is, the greater the distances of all nodes are, indicating that the t-th backup scheme has a higher goodness.
[0065] Obtain the backup scheme with the maximum goodness and back up the data corresponding to all nodes to complete the data storage of the information and communication technology (ICT) innovation terminal.
[0066] Thus, the data storage of the ICT innovation terminal is completed.
[0067] Although this specification has shown and described multiple embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, modifications, and alternative ways will occur to those skilled in the art without departing from the spirit and scope of the present invention.
Claims
1. A data storage method for information technology application innovation terminals, characterized in that, Including: Obtain several data sets of the information technology application innovation terminal. The data set is a set of data of an item at different time periods. Obtain the metadata of all data in each data set. The metadata includes the detection time, data source, data flow direction, and data size of the data. Based on the data source and data flow direction of each data, establish a data circulation model. Based on the genetic algorithm idea and the data circulation model, establish a data genetic model; according to the genetic performance of each data in the data genetic model, obtain the importance of each data set; based on the retrieval records of each data and the linear fitting method, obtain the heat change model of each data set; according to the importance of the data set and the heat change model of the data set, obtain the number of copies to be made for each data set. According to the number of copies to be made for each data set, establish a backup node sequence; randomly place all nodes of the backup node sequence in all backup units to obtain several backup plans; according to the genetic performance of the nodes in all backup units in the data genetic model in each backup plan, obtain the goodness of each backup plan; store the data with the backup plan with the maximum goodness.
2. The data storage method for the Xinchuang terminal according to claim 1, wherein The establishment of the data circulation model includes: Regard any data as a node, and connect any two nodes according to the data source and data flow direction relationship existing between any two nodes. Record the distance between any two nodes as 1, and use the nodes with an in-degree of 0 in the data circulation graph as starting points to perform a depth-first search algorithm respectively to obtain the shortest distance from each starting point to each node, which is recorded as the level of each node. When the level of the node is not unique, record the minimum value of the level of the node as the level of the node, and connect all nodes according to the level to obtain the data circulation model.
3. The data storage method for the IT application innovation terminal according to claim 2, wherein The connection of any two nodes includes: For the i-th node and the c-th node, if the data corresponding to the c-th node is the data source of the data corresponding to the i-th node, then connect the i-th node and the c-th node, and the direction points to the i-th node; if the data corresponding to the c-th node is the data flow direction of the data corresponding to the i-th node, then connect the i-th node and the c-th node, and the direction points to the c-th node.
4. A data storage method for an information and communication technology (ICT) innovation terminal according to claim 1, wherein The establishment of the data genetic model includes: Regard each data as a 01 string of the same size as the data size, which is recorded as the corresponding chromosome of each data; based on the data circulation of each data set reflected by the data circulation model, obtain the genetic probability sequence of each data set; based on the data size of the data corresponding to different nodes and the data circulation model, obtain the selection operator from node to node, and obtain the mutation operator from data set to data set. The data genetic model is expressed as that the corresponding chromosome of the data corresponding to the node at the e-th level is inherited according to the genetic probability sequence of the corresponding data set to obtain the node at the e + 1-th level, and the corresponding chromosome of the data corresponding to the node at the e + 1-th level is represented by the product of the corresponding chromosome of the data corresponding to the node at the e-th level and the selection operator and the mutation operator.
5. A data storage method for an information technology application innovation terminal according to claim 4, wherein, The obtaining of the genetic probability sequence of each data set includes: Regarding any data set as the target data set, obtain the frequency distribution sequence of the data flow directions of all the data in the target data set, which is denoted as the genetic sequence of the target data set. Divide each value in the genetic sequence of the target data set by the number of data in the target data set to obtain the genetic probability sequence of the target data set.
6. A data storage method for an information technology application innovation terminal according to claim 4, characterized in that, The selection operator for obtaining the node-to-node includes: If the b-th node is the data source of the a-th node, the selection operator of the a-th node for the b-th node satisfies the expression: ; In the formula, represents the selection operator of the a-th node for the b-th node; , represent the data sizes of the data corresponding to the a-th node and the b-th node; , represent the set of data sizes of the corresponding data sets corresponding to the data of the a-th node and the b-th node; represents the maximum value function; represents the normalization function.
7. A data storage method for an information technology application innovation terminal according to claim 4, characterized in that, The mutation operator for obtaining the data set-to-data set includes: If the b-th node is the data source of the a-th node, the mutation operator of the corresponding data set of the a-th node for the corresponding data set of the b-th node satisfies the expression: ; Wherein, represents the mutation operator of the corresponding data set corresponding to the data of the a-th node with respect to the corresponding data set corresponding to the data of the b-th node; , represent the data size sets of the corresponding data sets corresponding to the data of the a-th node and the b-th node; represents the maximum value function; represents the edit distance function; represents the normalization function.
8. The data storage method for the information technology application innovation terminal according to claim 1, wherein, The obtaining of the importance of each data set includes: Taking each node at the first level as a starting point, traverse the data flow model using the depth-first search algorithm to obtain all possible paths; for the node sequence of the m-th possible path from the b-th node to the a-th node , multiply the selection operator of any two adjacent nodes by the mutation operator of the corresponding data sets of the two adjacent nodes, and multiply them cumulatively to obtain the importance of the b-th node to the a-th node in , add up the importance of the b-th node to the a-th node in all possible paths from the b-th node to the a-th node to obtain the importance of the b-th node to the a-th node; for the d-th data set, obtain the importance of the nodes corresponding to the data contained in the d-th data set to all other nodes, and sum them up to obtain the importance of the d-th data set.
9. A data storage method for a domestic information technology innovation terminal according to claim 1, characterized in that, The number of copies to be made for each data set satisfies the expression: ; In the formula, represents the number of copies to be made for the d-th data set; , represent the minimum and maximum values of the range of the pre-set number of copies; represents the importance of the d-th data set; represents the slope of the heat change model of the d-th data set; represents the ceiling function.
10. A data storage method for an information technology application innovation terminal according to claim 8, characterized in that, The obtaining of the goodness of each backup scheme includes: Take the negative correlation normalization result of the importance of the b-th node for the a-th node as the distance between the b-th node and the a-th node; The goodness of any backup scheme satisfies the expression: ; In the formula, represents the goodness of the t-th backup scheme; represents the number of backup units; represents the distance between the u-th node and the v-th node in the s-th backup unit of the t-th backup scheme; represents the number of nodes in the s-th backup unit of the t-th backup scheme.
Citation Information
Patent Citations
Distributed data storage optimization method and device
CN117093131A
Method, device and equipment for generating flow chart based on flow engine, medium and product
CN118447127A
Big data-based relational database backup recovery method and system
CN119166428A
Grouping quality management system and method based on dynamic data source
CN119624261A
Cloud big data storage management method
CN119781690A