Computing power service data set storage method

By introducing dynamic mathematical models and machine learning algorithms into the computing power service data set storage system, real-time quantitative analysis and management of data access frequency, node load and timeliness, the static and local optimization problems of data scheduling strategies in traditional methods are solved, and the system's resource utilization and response speed are improved.

CN120162223AActive Publication Date: 2025-06-17JIANGSU FUTURE URBAN PUBLIC SPACE DEV & OPERATION CO LTD

Patent Information

Application Number
CN202510638890.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

Traditional static logical partitioning, fixed scheduling and simple cache strategies cannot effectively adapt to the dynamic changes in data access frequency and the load fluctuations of each storage node, resulting in the problem of local optimal or even information silos in the scheduling, classification and migration operations between storage media.

Method used

The overall data management method is adopted that integrates normalized partitioning functions, weighted node load models, data classification functions, prediction scheduling models and migration functions. The data access frequency is normalized through mathematical formulas, and the load of each storage node is evaluated in real time, a hot, temperature and cold data classification model is constructed, and a machine learning algorithm is used to predict functions to guide cache scheduling and data migration decisions.

Benefits of technology

It realizes innovative integration and information sharing between various technical links, breaks through the static and local optimization problems in traditional data scheduling strategies, improves the resource utilization and response speed of the computing power service data set storage system, and realizes the intelligence, refinement and adaptation of data storage and access management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162223A_ABST
    Figure CN120162223A_ABST
Patent Text Reader

Abstract

The invention discloses a computing power service data set storage method, which relates to electric digital data processing, and comprises the following steps: classifying data to be stored, and extracting specific characteristics of each data item; determining a data writing sequence, and executing data writing operation according to the sequence; independently calculating a hash value for each data item after data writing is completed, and transmitting each data item and the corresponding hash value to a predetermined number of storage nodes for data consistency verification; monitoring the access information of each storage node in real time, dividing logic partitions, and scheduling the data access request according to the load condition of each storage node; and data categories are divided according to the data access frequency and the data timeliness, and data caching operation is executed on the data. According to the method, the limitation of a traditional static scheduling and hierarchical management method is broken through, intelligentization, refinement and self-adaption of data storage and access management are realized, and the resource utilization rate and the response speed of a computing power service data set storage system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and particularly to a method for storing a computing power service data set. Background Art

[0002] In the current fields of computing power services and big data applications, the demand for the generation and storage of massive data has increased rapidly. However, traditional static logical partitioning, fixed scheduling, and simple caching strategies cannot fully adapt to the challenges brought by the dynamic changes in data access frequencies and the load fluctuations of each storage node. At the same time, there is a lack of an effective collaborative management mechanism among independent technologies, resulting in problems such as local optimality or even information islands in the scheduling, classification, and migration operations of data between storage media. The present invention realizes real-time quantitative analysis of data access frequencies, timeliness, and node loads by introducing a dynamic mathematical model and machine learning algorithms, breaking the technical bottleneck of the mutual independence of each link in the traditional method as a whole. Summary of the Invention

[0003] In view of the above existing problems, the present invention is proposed.

[0004] Therefore, the present invention proposes an overall data management method that integrates a normalization partitioning function, a weighted node load model, a data classification function, a predictive scheduling model, and a migration function. This method uses mathematical formulas to perform normalization partitioning on data access frequencies, evaluates the load of each storage node in real time through weighted calculation, constructs hot, warm, and cold data classification models based on data access records and timeliness, introduces a predictive function based on time series and machine learning to guide cache scheduling, and uses the migration function to automatically determine the migration decision of data between different storage media, realizing innovative integration and information sharing among various technical links, thus breaking through the static and local optimization problems existing in traditional data scheduling strategies.

[0005] To solve the above technical problems, the present invention provides the following technical solution. A method for storing a computing power service data set includes: Classify the data to be stored according to data types and update frequencies, and extract specific features of each data item; determine the data writing order based on the extracted features, historical writing records, and the current system state, and perform data writing operations in accordance with the order; after the data writing is completed, calculate the hash value for each data item separately, and transmit each data item and its corresponding hash value to a predetermined number of storage nodes for data consistency verification; monitor the access information of each storage node in real time, divide all data into a fixed number of logical partitions according to the collected access frequency, response latency, and resource utilization data, and perform scheduling on data access requests according to the load conditions of each storage node; divide data categories according to data access frequencies and data timeliness, and perform data caching operations on the data.

[0006] As a preferred solution of the computing power service data set storage method described in the present invention, wherein: extracting specific features of each data item includes dividing data into two clear categories of high-speed updated data and low-speed updated data according to the type and update frequency of the data, extracting specific attributes of each data item for the high-speed updated data and the low-speed updated data respectively, and recording the attribute values of each data item.

[0007] As a preferred solution of the computing power service data set storage method described in the present invention, wherein: determining the data writing order includes, extracting the specific attributes that make up the feature vector for each data item where represents the th attribute of the th data item and is the total number of attributes, and calculating the weighted sum of attributes using a pre-set weight vector ; is the weight corresponding to the th attribute, is the weighted sum of the attributes of the data item , and through the formula: ; obtaining the basic score of the data item, introducing a historical writing record function for quantifying the average writing delay or throughput recorded for the th data item in previous operations, and defining the current system state quantity to reflect the overall load cache occupancy rate and response delay at the data writing moment, and constructing a comprehensive scoring model, where represents the historical writing record function of the th data item, and its value represents the historical writing delay or throughput, is the th data item's comprehensive score, and through the formula: ; calculating the comprehensive score of the th data item, where and are pre-set coefficients, respectively measuring the influence of data attribute scores, historical records, and the current system state on the writing order, and the comprehensive score is used as a prediction of the writing time, that is, let: ; after the data writing operation is completed, record the actual writing time of each data item and construct a loss function to measure the difference between the predicted value and the actual value and the violation of the order between data items, and define the loss function as: ; wherein, represents the write time predicted according to the comprehensive score, is a balance parameter between 0 and 1 for weighing between prediction error and sequential consistency, and the set is defined as all pairs of data items that need to satisfy the predetermined write order constraint, where the order constraint requires that for each pair must satisfy conditions; The gradient descent method is used to perform online update on the weight parameters in the model, and its update formula is: ; wherein, is the updated weight, is the current weight, is the overall loss function, represents the learning rate, and the update is performed in real time after each data write operation.

[0008] As a preferred solution of the computing power service data set storage method described in the present invention, wherein: the separately calculating the hash value includes calculating the hash value for each data item by using a fixed algorithm, attaching the calculated hash value to the corresponding data item, and transmitting each data item attached with the hash value to a preset number of storage nodes for subsequent consistency verification.

[0009] As a preferred solution of the computing power service data set storage method described in the present invention, wherein: the data consistency verification includes comparing each received data item and its attached hash value in each storage node, using a distributed consensus algorithm to perform consistency judgment on the hash values of each data item in all storage nodes, performing conflict detection on the data item when it is found that the hash value of a certain data item does not match, and performing data synchronization operation on the data item according to the primary and standby synchronization rules.

[0010] As a preferred solution of the computing power service data set storage method described in the present invention, wherein: the real-time monitoring of the access information of each storage node includes that the steps of real-time monitoring of the data access information of each storage node include collecting the data access frequency, response delay and resource utilization rate on each storage node, transmitting the collected access information to the central scheduling module, and recording the transmitted information in the central scheduling module according to a preset data collection period; Among them, the logical partitioning of all data according to the monitoring data recorded in the central scheduling module includes: counting the data access frequency, dividing the data into a fixed number of logical partitions according to the statistical results, determining the scheduling order of data access requests based on the current load conditions of each storage node, and allocating the data access requests to the corresponding storage nodes according to the scheduling order.

[0011] As a preferred solution of the computing power service data set storage method described in the present invention, among them: the execution of scheduling for data access requests includes: counting the access records of all data items, dividing all data into hot data, warm data and cold data respectively according to the statistical results, storing the divided hot data in a high-speed storage medium, storing the divided cold data in a low-speed storage medium, and performing a predetermined intermediate storage operation on the divided warm data.

[0012] As a preferred solution of the computing power service data set storage method described in the present invention, among them: the execution of data caching operations includes that the step of periodically evaluating the data access status in each storage medium includes counting the data access frequency in each storage medium within a fixed evaluation period, determining the specific migration plan of data between different storage media according to the statistical results, and performing the migration operation of data from the low-speed storage medium to the high-speed storage medium, so that the data migration operation and the data caching operation are continuous in the operation process.

[0013] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the computing power service data set storage method are implemented.

[0014] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the computing power service data set storage method are implemented.

[0015] Advantages of the present invention: By constructing a complete set of dynamic scheduling systems based on mathematical models and machine learning algorithms, the present invention realizes the real-time dynamic allocation of data access requests and storage resources among storage nodes. Its technical solution uses a normalized partitioning function to statistically divide data access frequencies, and combines a weighting function to evaluate indicators such as CPU usage rate, memory occupancy rate, and disk I / O of each node in real time to determine the optimal scheduling order. At the same time, a hot, warm, and cold data classification model is constructed using access records and timeliness data, enabling data to be automatically allocated to high-speed, low-speed, or medium-speed storage media according to access frequencies, and achieving precise prediction of data caching timing and caching capacity through a prediction function based on time series analysis and machine learning, thereby automatically triggering data migration operations between different storage media. This technology innovatively breaks through the limitations of traditional static scheduling and hierarchical management methods, realizing the intelligent, refined, and adaptive management of data storage and access, and overall improving the resource utilization rate and response speed of the computing power service data set storage system. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for description in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other accompanying drawings without creative efforts based on these drawings.

[0017] Figure 1 Schematic flowchart of a method for storing a computing power service data set provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the above objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0019] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention, but the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of the present invention, so the present invention is not limited by the specific embodiments disclosed below.

[0020] Embodiment 1 Refer to Figure 1 , this embodiment provides a method for storing a computing power service data set, including: S1: Classify the data to be stored according to the data type and update frequency, and extract the specific features of each data item.

[0021] The extraction of the specific features of each data item includes dividing the data into two categories, high-speed updated data and low-speed updated data, according to the data type and update frequency, extracting the specific attributes of each data item for the high-speed updated data and the low-speed updated data respectively, and recording the attribute values of each data item.

[0022] S2: Determine the data writing order based on the extracted features, historical writing records, and the current system state, and perform the data writing operation in accordance with the order.

[0023] Set a group of parameters, which include data format, data source, data generation time interval, and a predefined time sliding window. The length of the time window is determined according to the actual performance of the data processing system and uses a fixed duration as the standard. This duration is determined through previous experiments and is used to preliminarily screen the data when the data enters the processing flow, so as to count the number of updates of each data item within a fixed time. This statistical process uses the moving average calculation method and the standard deviation calculation method to quantitatively analyze the data update frequency. When counting, set the moving average window to the length of the predefined time sliding window and calculate the standard deviation using the sample data. The calculation result obtains a dynamic threshold through a formula. The dynamic threshold is determined by adding a fixed multiple of the standard deviation to the moving average value, where the fixed multiple takes the value 3. This value is determined based on the historical data statistics results to achieve a quantitative division of the data update frequency. Divide the data into high-speed updated data with the number of updates higher than the threshold and low-speed updated data with the number of updates lower than the threshold according to the calculated dynamic threshold. After each type of data is divided, a rule-based parsing algorithm is used for processing. This algorithm uses a preset regular expression template and fixed parsing rules to extract specific attributes from each data item. These specific attributes are the unique identifier, generation timestamp, data size, data type, and data source information of the data item. The extracted attribute values are stored in a relational database in the form of key-value pairs according to the preset data record structure. An automatic verification module is introduced during storage. This module detects the extraction results according to the predefined data format and numerical range rules, marks the data items that do not meet the format requirements, and removes them according to the filtering conditions.

[0024] Meanwhile, a detailed processing log is generated for the processing status of each data item. The log records the number of data item updates, the calculated dynamic threshold, the data classification result, and the specific attribute values extracted. All log information is saved in a standardized format in the central database for subsequent parameter review and system debugging. In addition, a plugin interface is designed and implemented. This interface clearly defines the interface call sequence and data exchange format, allowing users to perform additional attribute parsing operations on data items according to predefined rules in specific application scenarios, and integrating the additionally extracted attributes into the data record in the same format as the main record. All operation steps are executed in sequence according to the fixed process of data preprocessing, update frequency statistics, dynamic threshold calculation, data classification, attribute parsing, automatic verification, data recording, and plugin extension. Detailed records are made of each key parameter and processing information. The entire data processing flow is implemented in a continuous operation manner within the data system, and a regular data backup mechanism is used to ensure the complete preservation and real-time update of all data and log information.

[0025] The determination of the data writing order includes extracting the specific attributes that make up the feature vector for each data item where represents the th attribute of the th data item and is the total number of attributes. Using the pre-set weight vector calculate the weighted sum of attributes is the weight corresponding to the th attribute, is the weighted sum of the attributes of data item . Through the formula: ; obtain the basic score of the data item. Introduce the historical write record function to quantify the average write delay or throughput recorded for the th data item in previous operations, and define the current system status quantity to reflect the overall load cache occupancy rate and response latency of the system at the data write moment, and construct a comprehensive scoring model, where represents the historical write record function of the th data item, and its value represents the historical write delay or throughput, is the th data item's comprehensive score. Through the formula: ; calculate the comprehensive score of the th data item, where and is a preset coefficient that measures the influence of data attribute scores, historical records, and the current system state on the write order respectively. The comprehensive score is used as the predicted write time, i.e., let: ; After the data write operation is completed, record the actual write time of each data item And construct a loss function to measure the difference between the predicted value and the actual value and the violation of the order between data items. Define the loss function as: ; Among them, represents the write time predicted according to the comprehensive score, is a balance parameter between 0 and 1 used to trade off between prediction error and order consistency. The set is defined as all pairs of data items that need to satisfy the predetermined write order constraint, where the order constraint requires that for each pair must satisfy condition; Use the gradient descent method to update the weight parameters in the model online. Its update formula is: ; Among them, is the overall loss function, represents the learning rate, is the updated weight, is the current weight. The update is performed in real time after each data write operation so that the model can continuously adjust and optimize the prediction results according to the latest actual write data in subsequent operations, thus forming a closed-loop feedback machine learning mechanism. This mechanism integrates multiple links such as data attribute extraction, historical performance recording, real-time system state acquisition, comprehensive score calculation, write order prediction, actual write time recording, loss function calculation, and parameter update. Its innovation lies in constructing a comprehensive score by linearly combining the static characteristics, historical behaviors, and current system environment of data items, and further introducing order constraints into the loss function, thereby ensuring the consistency of the write order while guaranteeing the prediction accuracy. This mathematical model combines static attributes, historical data, and dynamic system states, determines the data write order using the comprehensive score, and adjusts the model parameters through real-time feedback, thus realizing a creative online learning data write sorting mechanism based on the improvement of the existing technology.

[0026] S3: After the data writing is completed, calculate the hash value for each data item separately, and transfer each data item and its corresponding hash value to a predetermined number of storage nodes for data consistency verification.

[0027] In the step of generating the hash value of each data item after data writing is completed, an improved hash generation mechanism is introduced. This mechanism comprehensively utilizes the characteristics of the data item itself, the dynamic system state, and the verifiable delay function to break through the limitations of traditional fixed hash methods in precomputation and parallel attack prevention. The specific implementation steps are as follows: For each data item Call a secure pseudo-random generation function Obtain a unique random number ; At the same time, use a deterministic extraction function Extract a fixed signature value from the data item ; After obtaining the current normalized timestamp , the following combination function is used to and and perform weighted fusion: ; Among them, is the weight coefficient for balancing the random number, represents the bitwise exclusive OR operation; Introduce a neural network mapping function Process the current system state to obtain a fixed-length bit string output, and then concatenate the above combination result with through the concatenation operator to form an input string: ; Use a verifiable delay function to calculate the input string , and its output is: ; This function is designed to have a fixed calculation delay under the condition of no parallel acceleration, thus introducing a time constraint in the generated output; Concatenate the original data item with the delayed output through the concatenation operator to form a new data string, and use a fixed encryption hash function to calculate the hash value: ; The generated hash value is the improved hash value, and then is appended to to form an additional data item: ; Send the additional data item to a predetermined number of storage nodes for subsequent data consistency verification.

[0028] The data consistency check includes comparing each received data item and its attached hash value within each storage node, using a distributed consensus algorithm to determine the consistency of the hash values of each data item in all storage nodes, performing conflict detection on a data item when a mismatch in its hash value is found, and performing a data synchronization operation on the data item according to the primary-backup synchronization rule.

[0029] Specifically, during the process of performing data consistency checks between storage nodes, each storage node receives a data item and its attached hash value and then uses a fixed algorithm to verify whether the stored locally matches the attached value. Then, each node aggregates the hash values it calculates to form (where represents the set of all storage nodes), and uses a distributed consensus mechanism to calculate the consensus hash value of the data item (e.g., through majority voting or a mode algorithm), and its calculation formula is: ; Among them, represents the value with the highest frequency of occurrence in the set. Then, the average comparison distance is introduced to quantify the difference in hash values between nodes, and its calculation formula is: ; Among them, represents the Hamming distance between hash values, is the total number of storage nodes; when exceeds the preset threshold , the conflict detection process is triggered. This process selects a primary node (which can be determined according to node historical reliability or fixed allocation), and calculates the distance between each node and the primary node : ; If for a certain node there is (where is the preset synchronization threshold), then a data synchronization operation is performed to update the node to the data of the primary node, that is: ; Among them, represents the hash value distance between node l and the primary node , respectively represent node l and the master node the data item in the content of represent the hash value calculated or received by storage node l for the data item the hash value represents the hash value corresponding to the data item in the master node P.

[0030] The whole process integrates a consensus mechanism based on taking the mode of hash values, an inconsistency metric of the average comparison distance, and a primary-backup synchronization rule, and realizes a closed-loop process of data item consistency verification among storage nodes through online conflict detection and data synchronization operations.

[0031] S4: Monitor the access information of each storage node in real time, divide all data into a fixed number of logical partitions according to the collected access frequency, response delay, and resource utilization data, and schedule data access requests according to the load conditions of each storage node.

[0032] Real-time monitoring of the data access information of each storage node includes collecting the data access frequency, response delay, and resource utilization on each storage node, transmitting the collected access information to the central scheduling module, and recording the transmitted information in the central scheduling module according to a preset data collection period; the step of logically partitioning all data according to the monitoring data recorded in the central scheduling module includes counting the data access frequency, dividing all data into a fixed number of logical partitions according to the statistical results, and clarifying the scheduling order of data access requests according to the current load conditions of each storage node, and allocating the data access requests to the corresponding storage nodes according to this order.

[0033] Specifically, after receiving the monitoring data of each storage node, the central scheduling module first counts the data access frequency of each node according to the preset collection period, and divides all the data into K logical partitions according to the statistical results according to the predefined number of logical partitions K (for example, K = 10). At the same time, after collecting the load metrics including CPU usage rate, memory occupancy rate, and disk I / O, etc. in each storage node, these load information are transmitted to the central scheduling module, and the data packets transmitted by all nodes are recorded and summarized in the central scheduling module according to the preset data collection period. Then, the data access requests of each storage node are sorted according to the recorded load information to determine the scheduling order, and the data access requests are allocated to the corresponding storage nodes in the predefined data format according to this order; the central scheduling module also counts the historical access records of all data items, and divides all the data into hot data, warm data, and cold data according to the statistical results. Among them, the hot data is stored on the high-speed storage medium according to the predefined rules, the cold data is stored on the low-speed storage medium, and the warm data performs the predefined intermediate storage operations, and determines the cache scheduling timing and cache capacity of the data items with higher access frequency according to the historical access records, loads the data to be cached from the storage medium into the cache, and performs consistency verification on the loaded data and the original stored data during the loading process; the central scheduling module also counts the data access frequency in each storage medium within a fixed evaluation period, determines the specific migration plan of the data between different storage media according to the statistical results, and performs the migration operation of the data from the low-speed storage medium to the high-speed storage medium, so that the data migration operation and the data caching operation are continuous in the whole operation process.

[0034] S5: Divide the data into categories according to the data access frequency and data timeliness, and perform data caching operations on the data.

[0035] Furthermore, after the monitoring data is collected, the central scheduling module calculates the access frequency for each data item using the formula: ; to divide the data into a fixed number of logical partitions, where represents the number of accesses of a single data item within the statistical period, represents the maximum access frequency within the statistical period, is the total number of predefined logical partitions, represents the ceiling function; The load metrics (including CPU usage rate, memory occupancy rate, and disk I / O utilization rate) collected by each storage node through the monitoring agent are respectively denoted as and , and through the weighted function: ; Calculate the load score of each node l, where and are preset weight coefficients; the central scheduling module sorts each node according to the load score, and allocates the data access requests to the corresponding storage nodes in the predefined data format according to the sorting result, so as to determine the scheduling order.

[0036] Meanwhile, the central scheduling module statistically analyzes the historical access records of all data items, and records the access frequency of each data item and the data timeliness parameter (such as the time interval since the last access), and according to the classification function: ; Divide the data items into hot data, warm data and cold data, where and are the hot data and cold data thresholds of the access frequency respectively, and are the timeliness thresholds respectively; according to the data classification result, the hot data is stored in the high-speed storage medium, the cold data is stored in the low-speed storage medium, and the warm data is managed by intermediate storage operations, further refining the data management strategy.

[0037] For data items with a high access frequency, the system introduces a prediction function based on the historical access pattern (which can be implemented through a time series model or a machine learning algorithm), and this function is used to predict the future access trend, so as to determine the scheduling time and the required cache capacity of the data cache; once it is predicted that the data access frequency is about to increase, the system will load the data to be cached from the storage medium to the high-speed cache according to the predefined protocol, and perform consistency verification on the data during the loading process to ensure that the data content has not deviated during the loading process.

[0038] In addition, within a fixed evaluation period, the central scheduling module statistically analyzes the data access frequency in each storage medium, and according to the statistical data, uses the migration function: ; Judge whether it is necessary to perform a data migration operation, where is the preset migration threshold; if the data access frequency reaches or exceeds , the system will automatically trigger a data migration operation to migrate the data from the low-speed storage medium to the high-speed storage medium, and maintain the consistency of the data during the migration process to ensure the continuity of caching and migration of the data throughout the operation process.

[0039] Specifically, in this embodiment, first, through the normalization formula Divide the data into a fixed number of logical partitions according to the access frequency, and then use the load function Perform a weighted score on the current load of each storage node, and determine the scheduling order of data access requests based on the scoring results; subsequently, through the classification function Divide all the data into three categories: hot, warm, and cold, and formulate storage policies respectively; for data items with high access frequency, through the prediction function Determine the cache scheduling timing and cache capacity, and perform data migration operations when necessary. The decision basis is determined by the migration function to achieve the dynamic management and scheduling of data among different storage media, breaking through the technical bottlenecks existing in traditional methods when dealing with dynamic data environments.

[0040] Embodiment 2 The second embodiment of the present invention, which is different from the previous embodiment in that: If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes contributions to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.

[0041] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0042] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the procedures Figure 1 or blocks Figure 1 specified in one or more of the procedures and / or blocks.

[0043] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the procedures Figure 1 or blocks Figure 1 specified in one or more of the procedures and / or blocks.

[0044] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0045] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A computing service data set storage method, characterized by: include, Classify the data to be stored according to data type and update frequency, and extract specific features of each data item; Determine the data writing sequence based on the extracted features, historical writing records and current system status, and perform the data writing operation in sequence; After the data is written, a hash value is calculated for each data item separately, and each data item and its corresponding hash value are transmitted to a predetermined number of storage nodes for data consistency verification; Monitor the access information of each storage node in real time, divide all data into a fixed number of logical partitions based on the collected access frequency, response latency and resource utilization data, and schedule data access requests based on the load of each storage node; Data is classified according to data access frequency and data timeliness, and data caching operations are performed on the data.

2. A computing service data set storage method according to claim 1, characterized in that: The extracting of specific features of each data item includes dividing the data into two clear categories of high-speed update data and low-speed update data according to the type and update frequency of the data, extracting specific attributes of each data item for the high-speed update data and the low-speed update data respectively, and recording the attribute value of each data item.

3. A computing service data set storage method as claimed in claim 2, characterized in that: Determining the data writing order includes: Extract the constituent feature vectors for each data item The specific attributes of Indicates Data items attributes and is the total number of attributes, using a pre-set weight vector Calculate the weighted sum of attributes, It corresponds to The weight of the attribute, For data items The weighted sum of the attributes is given by the formula: ; Get the basic score of the data item and introduce the historical write record function To quantify the The average write latency or throughput of a data item recorded in previous operations, and defines the current system state To reflect the overall load cache occupancy and response delay of the system at the time of data writing, and construct a comprehensive scoring model, where Indicates The historical write record function of each data item, whose value represents the historical write latency or throughput, For the The comprehensive score of each data item is calculated by the formula: ; Calculate the The comprehensive score of data items, and are pre-set coefficients, which respectively measure the impact of data attribute score, historical records and current system status on the write order. The comprehensive score is used to predict the write time, that is, ; After the data writing operation is completed, record the actual writing time of each data item A loss function is constructed to measure the difference between the predicted value and the actual value and the violation of the order between the data items. The loss function is defined as: ; in, represents the write time predicted based on the comprehensive score, is a balance parameter between 0 and 1, which is used to balance prediction error and sequential consistency. is defined as all pairs of data items that need to satisfy a predetermined write order constraint, where the order constraint requires that for each pair Must meet conditions; The gradient descent method is used to update the weight parameters in the model online, and the update formula is: ; in, is the updated weight, is the current weight, is the overall loss function, represents the learning rate, and the update is performed in real time after each data write operation.

4. A computing service data set storage method as claimed in claim 3, characterized in that: The separate calculation of hash values ​​includes calculating a hash value for each data item using a fixed algorithm, appending the calculated hash value to the corresponding data item, and transmitting each data item appended with the hash value to a preset number of storage nodes for subsequent consistency verification.

5. A computing service data set storage method as claimed in claim 4, characterized in that: The data consistency check includes comparing each received data item and its attached hash value in each storage node, using a distributed consensus algorithm to make consistency judgments on the hash values ​​of each data item in all storage nodes, and performing conflict detection on the data item when it is found that the hash value of a data item does not match, and performing data synchronization operations on the data item in accordance with the master-slave synchronization rules.

6. A computing service data set storage method as claimed in claim 5, characterized in that: The real-time monitoring of the access information of each storage node includes the step of real-time monitoring of the data access information of each storage node including collecting data access frequency, response delay and resource utilization rate on each storage node, transmitting the collected access information to the central scheduling module, and recording the transmitted information in the central scheduling module according to a preset data collection cycle; The logical partitioning of all data based on the monitoring data recorded in the central scheduling module includes counting the data access frequency, dividing the data into a fixed number of logical partitions according to the statistical results, and determining the scheduling order of data access requests based on the current load of each storage node, and allocating the data access requests to the corresponding storage nodes according to the scheduling order.

7. A computing service data set storage method as claimed in claim 6, characterized in that: The scheduling of data access requests includes collecting statistics on access records of all data items, dividing all data into hot data, warm data and cold data according to the statistical results, storing the divided hot data in a high-speed storage medium, storing the divided cold data in a low-speed storage medium, and performing predetermined intermediate storage operations on the divided warm data.

8. A computing service data set storage method as claimed in claim 7, characterized in that: The execution of the data caching operation includes the step of periodically evaluating the data access status in each storage medium, including counting the data access frequency in each storage medium within a fixed evaluation period, determining a specific data migration plan between different storage media based on the statistical results, and executing the data migration operation from the low-speed storage medium to the high-speed storage medium, so that the data migration operation and the data caching operation are performed continuously in the operation process.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • HBase client main and standby switching method and system based on fault perception

    CN119537484A

  • Data processing method, controller, battery management system and vehicle

    CN119917026A

  • Machine learning model scaling system with energy efficient network data transfer for power aware hardware

    US20220036123A1

  • Data storage method, apparatus and computing device

    WO2024187922A1

Cited By

  • Storage method for consistency verification of automobile extended-guarantee multi-source data

    CN121116966A