Large-scale data balanced division method
By associating data with the request area and using deep learning models for data classification and distributed storage, the problems of high latency and blockage in large-scale data processing are solved, and efficient and low-latency data storage and call experience are achieved.
Patent Information
- Application Number
- CN202510051285.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-27
AI Technical Summary
When traditional data storage systems process large-scale data, they are prone to high latency and process blockage, especially when multiple data requests occur simultaneously, which may lead to storage system crash.
By associating data with requested areas, data requests are classified using deep learning models and stored in distributed terms based on the degree of region correlation. When new data is written, it is stored to a temporary storage node and the optimal storage node is selected according to the call ratio of different locations.
It realizes high efficiency and low latency when requesting data, avoids process clogging and storage system crashes, and improves the data storage and calling experience.
Smart Images

Figure CN120045131A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data storage, and relates to a method for evenly partitioning large-scale data. Background Art
[0002] With the booming development of various applications such as mobile devices, social networks, and the Internet of Things, the data generated by human society has grown explosively. The traditional data storage method is usually disk storage, and users store all the data to be stored on the disk so that they can view the data anytime and anywhere. However, as the amount of data to be stored increases, it becomes increasingly difficult for traditional disks to meet the storage needs of users based on massive data in terms of capacity, performance, and bandwidth. Therefore, a data storage system supported by a cloud platform has emerged. A data center is deployed in the data storage system, and users can upload the data to be stored to the data center, and the data center stores the data to be stored.
[0003] In related technologies, a storage cluster for storing data is set in the data center. When a user uploads the data to be stored to the data storage system, the data center will receive the data to be stored, and the data center adds the received data to be stored to the storage cluster for storage. However, the storage location of the data cannot be associated with the request location of the data. During the data extraction process, it is easy to cause too high a delay. In the case of multiple data requests occurring simultaneously, it is also easy to cause process blockage, resulting in the collapse of the storage system. Summary of the Invention
[0004] To solve the problems of high delay during data extraction and easy process blockage when multiple data requests occur simultaneously in the background art, the present invention proposes a method for evenly partitioning large-scale data.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A method for evenly partitioning large-scale data includes:
[0007] Associating data with a request area according to the request frequencies of different data at different locations;
[0008] A deep learning model classifies the requests for data sent from different locations and the data types of the requested data, and forms a regional association degree between different data categories and different locations;
[0009] Distributed storage is performed on the location of the storage node of the data according to the regional association degree between any data and the area where the request location is located;
[0010] When new data is written, the new data is first stored in a temporary storage node. Based on the call ratio of the data in the temporary storage node at different locations as weights, a suitable location is selected to save the data to the optimal storage node.
[0011] Further, the specific method for associating the data with the request location is as follows:
[0012] First, generate a location set LOC{a1, a2, a3,..., an} for all request locations that initiate requests for a single data.
[0013] Extract any judgment location from the request location set, judge the distance between each request location and the any judgment location, and divide according to a preset distance. Those with a distance less than the preset distance between any request location and the judgment location are merged.
[0014] Repeat the judgment through multiple judgment locations, and the distance between multiple judgment locations is greater than twice the preset distance.
[0015] Perform a merging process on multiple request locations to generate a request area set πε{A1, B1, C1, D1,...., δn}.
[0016] Generate a mapping list between the data and the request area based on the minimum distance between the storage node and the request area, and complete the association operation between the data and the request location.
[0017] Further, the specific method for the deep learning model to classify the requests and the data types of the request data sent from different locations is as follows:
[0018] Classify the data requested in a certain area according to the correlation between multiple data requested by all locations in the area.
[0019] Perform cross-validation on the data requested in the area according to the deep learning model, perform cross-calculation between every two data requested in the area, and calculate the correlation between all data.
[0020] The deep learning model classifies the data according to the correlation between all data, and at the same time records the characteristic attributes in each category of data, and uses the characteristic attributes as the marked content of a certain category of data.
[0021] Compare the data not requested in the current area with the characteristic data in any category of the data requested in the current area, and classify the data not requested in the current area.
[0022] Further, the specific method for forming a location association degree between different data categories and different locations is as follows:
[0023] Match the feature attributes of the data categories requested in the current area with all locations in the current area through a deep learning model to generate the matching degree between the locations and the feature attributes;
[0024] According to the matching results between each location in the current area and the feature attributes, obtain the associated attributes between any location and the feature attributes;
[0025] Match the associated attributes of all locations in the current area with the feature attributes in the overall data to generate a list of associations between the data request locations and the data categories.
[0026] Furthermore, the method for distributed storage of data according to the regional association degree between any data and the area where the request location is located includes:
[0027] Set a regional correlation threshold. When the regional correlation degree between any data and the data request area exceeds the regional correlation threshold, mark the associated attribute between the data and the current data request area;
[0028] When the current data has an associated attribute only with a single data request area, move the data to the storage node closest to the data request area associated with the data;
[0029] When the current data has associated attributes with multiple data request areas, first judge the distances between the multiple data request areas. If the sum of the minimum distances between the multiple request areas exceeds the set value, create a new backup of the current data and store the current data and the backup of the current data in the storage node closest to the multiple request areas;
[0030] If the sum of the minimum distances between the multiple request areas does not exceed the set value, store the current data in the closest storage node, and the closest storage node has the smallest sum of distances to the center points of the multiple request areas compared to other storage nodes.
[0031] Furthermore, it also includes:
[0032] When any storage node calls the data a preset number of times, count the number of calls for each data in the current storage node. If the number of calls for any data does not meet the expected number, update the storage location of the current data.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] The present invention aggregates the data with the highest requested frequency in the area where any request location is located through unified calculation, and at the same time moves the data of this category to the data storage node closest to the request area, ensuring the rate during data requests, greatly reducing the delay during data transmission, and improving the usage experience when calling data in the data request area. Brief Description of the Drawings
[0035] Figure 1 is the operation flowchart of the present invention. Specific embodiments
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0037] Such as Figure 1 shown, the technical solution adopted by the present invention is as follows: A large-scale data balanced partitioning method, including:
[0038] Associate data with the request area according to the request frequencies of different data at different locations.
[0039] The deep learning model classifies the requests and the data types of the requested data sent from different locations, and forms a regional association degree between different data categories and different locations.
[0040] Distributively store the location of the storage node of the data according to the regional association degree between any data and the area where the request location is located.
[0041] When new data is written, first store the new data in a temporary storage node, and select a suitable location to save the data to the optimal storage node according to the data call ratio in the temporary storage node at different locations as the weight.
[0042] Suppose in fields such as autonomous driving, the timeliness requirement of data is high and the throughput is large. If the data is stored in different cities, a large amount of data is likely to cause problems such as data transmission channel congestion during the transmission process, affecting the safety of intelligent driving. Therefore, it is necessary to reasonably plan the data storage location so that the data storage can meet the timeliness requirement.
[0043] When it is necessary to partition the large-scale data cluster in the distributed storage, first associate the data with the request area according to the request frequencies of different data at different locations.
[0044] The specific method for associating the data with the request location is:
[0045] First, generate a location set LOC{a1, a2, a3,..., an} for all request locations that initiate requests to a single piece of data. Extract any judgment location from the request location set, calculate the distance between each request location and the judgment location, and divide them according to a preset distance. Merge those with a distance less than the preset distance between any request location and the judgment location. The preset distance can be set by the user. The smaller the user sets the preset distance, the more accurate the result of data storage division will be, but the computational cost will also be greater. The larger the user sets the preset distance, the more ambiguous the result of data storage division will be, and the corresponding computational cost will decrease.
[0046] Repeat the judgment through multiple judgment locations, with the distance between multiple judgment locations greater than twice the preset distance. This is to prevent a situation where a data request location happens to have the same distance from two judgment locations, resulting in both judgment locations including this request location simultaneously, or neither of the two judgment locations including this request location, causing the data request sent from this request location to fail to meet the real-time requirement.
[0047] Perform a merging process on multiple request locations to generate a request area set πε{A1, B1, C1, D1,...., δn}. Divide multiple request locations according to their spatial geographical locations to avoid overloading the computing server during simultaneous calculations without effective computing benefits. It should be noted that when the storage nodes are sufficient in number and widely distributed geographically, the large-scale data balanced division method provided by the present invention can regard the storage nodes as a cyclic set.
[0048] Generate a mapping list between the data and the request area based on the minimum distance between the storage node and the request area, and complete the association operation between the data and the request location. Based on the geographical location distance between the storage node and the request area, the request area can be corresponding to the storage node, and the request area can be associated with the nearest storage node to complete the data call of the request area at the nearest storage node.
[0049] When the minimum distances between multiple storage nodes and a single request area are the same, randomly select a single storage node as the mapping relationship object. Correspondingly, when the minimum distances between multiple request areas and a single storage node are the same, randomly select a single request area as the mapping relationship object.
[0050] After associating the storage node and the request area, the deep learning model classifies the requests for data sent from different locations and the data types of the requested data, and forms a regional association degree between different data categories and different locations.
[0051] The specific method for the deep learning model to classify requests sent from different locations and the data types of the request data is as follows:
[0052] Based on the correlation between multiple data requested by all locations in a certain area, the data requested in this area is classified. In a certain request area, an enterprise may call the market information in this area, and it also includes users calling the information of local users in the social field, etc. For such reasons, the request area can be associated with the data requested in this request area.
[0053] Based on the deep learning model, cross-validation is performed on the data requested in this area, and cross-calculation is performed between every two data requested in this area to calculate the correlation between all data. The above-mentioned association method is only part of the association information, and it is necessary to rely on the deep learning model to complete the association between all data and all request locations included in the request area.
[0054] The specific method for forming a location association degree between different data categories and different locations is as follows:
[0055] Through the deep learning model, the characteristic attributes of the data category requested in the current area are matched with all locations in the current area to generate a matching degree between the location and the characteristic attributes. A single data may have multiple attributes, and the characteristics of all data requested in this request area are uniformly calculated to obtain the characteristic attributes shared by most of the data in all the requested data.
[0056] Based on the matching results between each location in the current area and the characteristic attributes, the associated attributes between any location and the characteristic attributes are obtained. The characteristic attributes are further refined and associated according to all the request locations in this request area to obtain the characteristic attributes corresponding to the request locations.
[0057] Match the associated attributes of all locations in the current area with the characteristic attributes in the overall data to generate a data request location and data category association list.
[0058] The deep learning model classifies the data based on the correlation between all data, and at the same time records the characteristic attributes in each category of data, and uses the characteristic attributes as the marked content of a certain category of data. The deep learning model then marks the data itself to classify or process other unrequested data.
[0059] Compare the data not requested in the current area with the feature data in any category of the data requested in the current area, and classify the data not requested in the current area. After obtaining the data feature attributes required for the requested area, based on this feature attribute, judge and classify the data not requested in the requested area, which can be done by means of a decision tree or a neural network, classify the data not requested in the requested area, and obtain the data expected to be requested in this area.
[0060] In this embodiment, an example is given to judge and classify the data not requested in a certain requested area by means of a decision tree.
[0061] As an intuitive and easy-to-understand machine learning model, the core of a decision tree is to make decisions based on a tree structure. Each non-leaf node represents a decision branch on a feature attribute, and the output of each decision branch represents the judgment result of this feature attribute, while the leaf node represents the final category or decision result. This path from the root node to the leaf node corresponds to a conjunction rule, and the entire decision tree represents a set composed of multiple such rules.
[0062] The learning algorithm of a decision tree mainly focuses on how to select the optimal splitting attribute. The samples contained in the branch nodes of the decision tree belong to the same category as much as possible, that is, the "purity" of the nodes is getting higher and higher. Multiple criteria such as information gain, gain ratio, and Gini index are applied in practice.
[0063] First, extract the feature attributes from the data requested in the requested area. The feature attributes are of multiple attribute types and conform to the tree-like judgment branches of the decision tree.
[0064] Set the judgment branches of the decision tree in the form of a binary method. The binary method can avoid missing a certain feature attribute in the decision branches of the decision tree and can classify all the attributes of the data to be matched according to the feature attributes. Set the decision branches of the decision tree as different feature attributes generated by the data requested in the requested area to complete the creation of the decision tree model.
[0065] The decision tree uses the binary method to select the feature attribute that can best distinguish the data from all the feature attributes, and divides the data to be matched into two or more subsets according to the selected feature attribute, and then repeats this process for each subset until the stopping condition is reached, such as the depth of the tree reaches the limit or the purity of the subsets after splitting is high. During the decision-making process, the decision tree starts from the root node, makes judgments according to the feature rules of the nodes, and follows the path of the decision tree until it reaches the leaf node to obtain the final decision result or prediction result. The decision tree has the advantages of being easy to understand and interpret and being able to process non-linear data, and users can freely choose the matching method.
[0066] After obtaining the expected request data for the requested area, perform distributed storage on the location of the storage node of the data according to the regional correlation degree between any data and the area where the request location is located.
[0067] The method for performing distributed storage on the data according to the regional correlation degree between any data and the area where the request location is located includes:
[0068] Set a regional correlation threshold. When the regional correlation degree between any data and the data request area exceeds the regional correlation threshold, mark the associated attribute between the data and the current data request area.
[0069] When the current data has an associated attribute with only a single data request area, move the data to the storage node closest to the data request area associated with the data.
[0070] When the current data has an associated attribute with multiple data request areas, first judge the distances between the multiple data request areas. If the sum of the minimum distances between the multiple request areas exceeds the set value, create a new backup of the current data, and store the current data and the backup of the current data in the storage node closest to the multiple request areas.
[0071] If the sum of the minimum distances between the multiple request areas does not exceed the set value, store the current data in the closest storage node, and the sum of the distances between the closest storage node and the center points of the multiple request areas is the smallest compared to other storage nodes.
[0072] When new data is written, first store the new data in a temporary storage node. Select a suitable location to save the data to the optimal storage node according to the data call ratio in the temporary storage node at different locations as the weight. It should be noted that the storage device used for the temporary storage node should be able to meet the type of multiple reads and writes to avoid data loss and other situations after multiple reads and writes. At the same time, using a temporary storage node can avoid affecting the data that has been partitioned in other storage nodes and avoid destroying the original data structure.
[0073] The user can set a separate line in the information transmission line between the temporary storage node and each storage area according to their own needs. Because no storage optimization operations are performed on the temporary storage node, the data in the temporary storage node cannot meet data types with extremely high requirements for instantaneity such as autonomous driving. Inevitably, the data may not be classified when written. Therefore, if the user wants to meet data with high requirements for instantaneity like autonomous driving, they can only establish a separate high-speed information transmission channel to meet the requirements.
[0074] When any storage node calls data a preset number of times, it counts the number of calls for each piece of data in the current storage node. If the number of calls for any piece of data does not meet the expected number, it updates the current data storage location. The present invention can also monitor the data in the already partitioned storage nodes in real time to avoid excessive delay during data calls caused by incorrect partitioning.
[0075] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A large-scale data balanced partitioning method, characterized in that: Included are: Associating data with requesting regions based on how often different data are requested in different locations; The deep learning model classifies the data requests from different locations and the data types of the requested data, as well as forms regional associations between different data categories and different locations; Distribute the data storage nodes according to the regional association between any data and the area where the request is located; When new data is written, the new data is first stored in a temporary storage node. The data call ratio in the temporary storage node at different locations is used as a weight, and a suitable location is selected to save the data to the optimal storage node.
2. A large-scale data balanced partitioning method according to claim 1, characterized in that ,The specific method of associating data with the request location is: First, all the request locations for a single data are generated into a location set LOC{a1,a2,a3,...,an}, Extract any judgment location from the request location set, determine the distance between each request location and any judgment location, divide them according to the preset distance, and merge any request location with a judgment location whose distance is less than the preset distance; Repeat the judgment at multiple judgment locations, and the distance between the multiple judgment locations is greater than 2 times the preset distance; Merge multiple request locations to generate a request area set πε{A1,B1,C1,D1,....,δn}; Based on the minimum distance between the storage node and the request area, a mapping list between the data and the request area is generated to complete the association operation between the data and the request location.
3. A large-scale data balanced partitioning method according to claim 1, characterized in that ,The specific method of the deep learning model for classifying the requests from different locations and the data types of the requested data is as follows: Classify the data requested by a region according to the correlation between multiple data requested by all locations in the region; The data requested for the area are cross-validated based on the deep learning model, and cross-calculations are performed between every two data requested for the area to calculate the correlation between all data; The deep learning model classifies the data according to the correlation between all the data, and records the characteristic attributes in each category of data, and uses the characteristic attributes as the label content of a certain category of data; The data not requested in the current area is compared with the characteristic data in any category of the data requested in the current area, and the data not requested in the current area is classified.
4. A large-scale data balanced partitioning method according to claim 3, characterized in that ,The specific method of forming location association between different data categories and different locations is: The deep learning model is used to match the data category feature attributes requested by the current region with all locations in the current region, and the matching degree between the location and the feature attributes is generated; According to the matching results of each location and feature attribute in the current area, the associated attributes of any location and feature attribute are obtained; Match the associated attributes of all locations in the current area with the characteristic attributes in the overall data to generate an associated list of data request locations and data categories.
5. A large-scale data balanced partitioning method according to claim 1, characterized in that ,The method for distributing and storing data based on the regional association between any data and the area where the request location is located includes: Set a regional relevance threshold. When the regional relevance of any data to the data request area exceeds the regional relevance threshold, the data is marked with an associated attribute with the current data request area. When the current data has an associated attribute only with a single data request area, the data is moved to a storage node that is closest to the data request area associated with the data; When the current data has associated attributes with multiple data request areas, the distances between the multiple data request areas are first determined. If the minimum distances between the multiple request areas exceed the set value, a new backup of the current data is created, and the current data and the current data backup are stored in the storage nodes closest to the multiple request areas. If the sum of the minimum distances between multiple request areas does not exceed the set value, the current data is stored in the nearest storage node, and the sum of the distances between the nearest storage node and the center points of multiple request areas relative to other storage nodes is the smallest.
6. A large-scale data balanced partitioning method according to claim 1, characterized in that , also includes: When any storage node calls the preset number of data, it counts the number of calls for each data in the current storage node. If the number of calls for any data does not meet the expected number, the current data storage location is updated.