Data cluster processing method, data cluster processing equipment and storage medium

By obtaining the metadata of the Gaussian cluster, calculating space asset information, building an exception governance model, and cleaning up redundant data to improve the resource utilization rate of large-scale clusters, the problem of low resource utilization rate of clusters is solved, and reasonable allocation and balanced management of resources are achieved.

CN120523604APending Publication Date: 2025-08-22CHINA MERCHANTS BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510693340.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The resource utilization rate in large-scale clusters is low, the inefficient model is not cleaned, the cluster nodes cannot be fully loaded, and resource management fails to automatically trigger resource release based on model activity or performance indicators, resulting in low cluster resource utilization.

Method used

By obtaining the metadata of the Gaussian cluster, calculating the spatial asset information, building an exception governance model, determining the spatial redundancy governance items, updating the redundant data in response to the exception handling instructions, identifying and cleaning up the redundant data to recycle space resources.

Benefits of technology

The data utilization rate of cluster space is improved, the rationality of resource occupation and allocation balance are optimized, and diagnostic and auxiliary recollection group space is provided to the person in charge of the data area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523604A_ABST
    Figure CN120523604A_ABST
Patent Text Reader

Abstract

The invention discloses a data cluster processing method, data cluster processing equipment and a storage medium, relates to the technical field of data cluster processing, and discloses a data cluster processing method which comprises the steps that metadata of a Gaussian cluster is acquired, and space asset information of the Gaussian cluster is calculated according to the metadata; according to a historical cluster processing rule corresponding to the metadata, constructing an exception governance model; determining a spatial redundancy governance item of the spatial asset information according to a spatial redundancy governance rule of the anomaly governance model; and updating redundant data of the space asset information in response to an exception handling instruction corresponding to the space redundancy governance item. On the basis, the treatment items in the cluster data are identified, and redundant space resources are treated and recycled, so that the resource occupation rationality and the allocation balance in the cluster data are improved, and the data utilization rate of the cluster space is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data cluster processing, and in particular to a data cluster processing method, a data cluster processing device, and a storage medium. Background Art

[0002] In large-scale clusters, server nodes can reach tens of thousands, the amount of stored data can reach petabytes, and the number of files can reach hundreds of millions. As business and data volumes continue to grow, cluster capacity expansion, and storage and computing resources reach a certain scale, resource integration and governance of big data clusters becomes essential.

[0003] In the relevant data cluster resource management, the system failed to automatically trigger resource release based on model activity or performance indicators, resulting in a large number of inefficient models not being cleaned up, cluster nodes being fully loaded and unable to expand, resulting in low cluster resource utilization.

[0004] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a data cluster processing method, a data cluster processing device and a storage medium, aiming to solve the technical problem of low cluster resource utilization.

[0006] To achieve the above objectives, the present application proposes a data cluster processing method, which includes: Obtaining metadata of a Gaussian cluster, and calculating spatial asset information of the Gaussian cluster based on the metadata; Building an anomaly governance model based on historical cluster processing rules corresponding to the metadata; Determining the spatial redundancy governance items of the spatial asset information according to the spatial redundancy governance rules of the anomaly governance model; In response to the exception handling instruction corresponding to the space redundancy management item, the redundant data of the space asset information is updated.

[0007] In one embodiment, the step of constructing an anomaly governance model based on the historical cluster processing rules corresponding to the metadata includes: If the historical cluster processing rules include data storage and slice redundancy processing, a cluster slice and split strategy model is constructed; If the historical cluster processing rules include inefficient model processing, construct an abandoned no downstream governance model; If the historical cluster processing rules include data skew processing, build a table skew governance model; If the historical cluster processing rules include spatial exception processing, a temporary table cleanup model is constructed.

[0008] In one embodiment, the step of determining the spatial redundancy management items of the spatial asset information according to the spatial redundancy management rules of the abnormality management model includes: Obtaining the incremental usage information of the space asset information in the global zone, data zone, and management zone, as well as the cluster space trend; Determine, based on the incremental usage information and the cluster space trend, a warning data area whose space usage is higher than a preset usage rate; The space redundancy management item of the warning data area is determined according to the slice redundancy management rule, the tilted table cleaning rule, the no downstream model cleaning rule and the temporary table cleaning rule.

[0009] In one embodiment, the step of determining the space redundancy management item of the warning data area according to the slice redundancy management rule, the tilted table cleaning rule, the no downstream model cleaning rule, and the temporary table cleaning rule includes: Determine a slice redundancy splitting strategy for the warning data area according to the slice redundancy management rule; Obtaining a tilt rate of a data table in the early warning data area, and determining a tilt table cleanup strategy for the early warning data area based on the tilt rate and the tilt table cleanup rule; Determine a no-downstream model cleanup strategy based on the table-level lineage relationship of the warning data area and a no-downstream model cleanup rule; A temporary table clearing strategy is determined according to the temporary table clearing rule and the temporary table data in the warning data area.

[0010] In one embodiment, the step of determining the slice redundancy splitting strategy of the warning data area according to the slice redundancy management rule includes: Acquiring data access information of the early warning data area within a preset time period, and marking the data temperature of the early warning data area according to the data access information; The slice redundancy splitting strategy is determined according to the data temperature and the slice redundancy management rule.

[0011] In one embodiment, the step of determining the no-downstream model of the spatial asset information based on the abandoned no-downstream model and the table-level lineage relationship of the spatial asset information includes: Determine the table-level lineage relationship of the spatial asset information, and determine the model downstream dependency information corresponding to the table-level lineage relationship based on the abandoned no-downstream model; If the model downstream dependency information indicates that a downstream exists, determining the current recognition model's own dependency and cargo outbound operation information; If the recognition model has its own dependencies and there is no cargo outbound operation, it is determined that the current recognition model has no downstream model.

[0012] In one embodiment, after the step of determining the spatial redundancy governance items of the spatial asset information according to the spatial redundancy governance rules of the anomaly governance model, the data cluster processing method further includes: Obtaining remaining space information of the warning data area; Generate and output a cluster warning report of the space asset information based on the remaining space information and the space redundancy management items; Obtain an execution instruction corresponding to the cluster warning report, and set the execution instruction as the exception handling instruction.

[0013] In one embodiment, after the step of updating the redundant data of the space asset information in response to the exception handling instruction corresponding to the space redundancy management item, the data cluster processing method further includes: A redundant data processing task for obtaining the spatial asset information; Determining spatial governance benefit information of the spatial asset information according to the redundant data processing task; Output the space management income information and the space management report of the space asset information.

[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a data cluster processing device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data cluster processing method as described above.

[0015] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and stores a computer program on the storage medium. When the computer program is executed by the processor, the steps of the data cluster processing method described above are implemented.

[0016] One or more technical solutions proposed in this application have at least the following technical effects: After obtaining the metadata of the Gaussian cluster, the spatial asset information of the cluster is calculated through the metadata. At the same time, an anomaly governance model is constructed through the historical cluster processing rules of the metadata. Then, according to the spatial redundancy governance rules of the anomaly governance model, the spatial redundancy governance items that require data redundancy processing of the current spatial asset information are determined. Finally, in response to the anomaly processing instructions corresponding to the spatial redundancy governance items, the redundant data in the spatial asset information is updated. In this way, by identifying governance items and governing and recovering spatial resources, the person in charge of the data area is provided with diagnosis, assistance and recovery of cluster space, effectively evaluating the rationality of resource occupancy and distribution balance, and improving the data utilization rate of the cluster space. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 This is a schematic diagram of the cluster space governance architecture in the data cluster processing method of this application; Figure 2 This is a schematic diagram of the cluster space management process in the data cluster processing method of this application; Figure 3 A flowchart of the first embodiment of the data cluster processing method of the present application is provided; Figure 4 A flowchart of the second embodiment of the data cluster processing method of the present application is provided; Figure 5 Schematic diagram of the identification process of the data cluster processing method for this application without downstream models; Figure 6 A schematic diagram of the generation of cluster management warning reports for the processing method of the data cluster of this application; Figure 7 Schematic diagram of the spatial governance benefits of the processing method for this application data cluster; Figure 8 A schematic diagram of the cluster governance benefit report for the data cluster processing method of this application; Figure 9 Schematic diagram of the device structure of the hardware operating environment involved in the data cluster processing method in the embodiment of the present application.

[0020] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0022] In large-scale clusters, server nodes can reach tens of thousands, the amount of stored data can reach petabytes, and the number of files can reach hundreds of millions. As business and data volumes continue to grow, cluster capacity expansion, and storage and computing resources reach a certain scale, resource integration and governance of big data clusters becomes essential.

[0023] In the relevant data cluster resource management, the system failed to automatically trigger resource release based on model activity or performance indicators, resulting in a large number of inefficient models not being cleaned up, cluster nodes being fully loaded and unable to expand, resulting in low cluster resource utilization.

[0024] The Gauss cluster nodes on the big data platform are fully loaded and cannot be expanded. Cluster A has less than 20% free space. The parallel reconstruction of the old and new warehouses for big data presents challenges in ensuring stable production operations and sustaining business needs. Pain points for this type of cluster include: a lack of cost awareness in cluster space usage, a large number of inefficient models not yet cleaned up, and a failure to separate and store large tables according to the "hot data > warm data > cold data" principle. This results in a mismatch between space growth and actual business operations, leading to significant waste. Cluster nodes are fully loaded and cannot be expanded. Cluster A has less than 700TB of free space, with monthly growth of nearly 100TB. Some data areas are operating at critical capacity, and if space is full, it could cause production accidents. The large number of existing models has a significant downstream impact, and the usage range is hidden in the code, resulting in high analysis and governance costs, making developers hesitant to split and clean up data. Cluster space lacks a value measurement tool, making space growth unclear and ineffective in assessing the rationality of resource usage and the balance of its allocation.

[0025] The main solution of the embodiment of the present application is: obtaining metadata of a Gaussian cluster, and calculating spatial asset information of the Gaussian cluster based on the metadata; Building an anomaly governance model based on historical cluster processing rules corresponding to the metadata; Determining the spatial redundancy governance items of the spatial asset information according to the spatial redundancy governance rules of the anomaly governance model; In response to the exception handling instruction corresponding to the space redundancy management item, the redundant data of the space asset information is updated.

[0026] Specifically, by identifying the governance items in the Gaussian cluster and reclaiming space resources based on the governance items, we can provide diagnosis, assistance and recovery of cluster space for data area managers, thereby effectively evaluating the rationality of resource occupancy and distribution balance, and improving the data utilization rate of cluster space.

[0027] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, data cluster processing, etc. The following uses a data cluster processing device as an example to illustrate this embodiment and the following embodiments.

[0028] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0029] The cluster space governance architecture of the data cluster processing method of this application is as follows Figure 1 As shown, cluster space data processing includes metadata collection, data processing, and application visualization. Cluster applications can be visualized during the metadata collection and data processing stages. Specifically, during the cluster space governance stage, metadata is first collected, including basic information such as database space, data table space storage, cluster slicing and splitting information, table-level lineage information, table-job correspondence information, and table scripts. This is followed by global data identification and processing, including data processing at three levels: modeling, calculation, and analysis. After processing, the metadata information is converted into cluster space management platform data.

[0030] Among them, in the data processing stage, the cluster space is cleaned up of redundant data through the model layer. The abnormal management models of the model layer include slice redundancy management model, temporary table cleaning management model, abandoned no downstream management model and severe table skew management model. These management models process redundant data related to four categories of content: slice redundancy, temporary table cleaning, abandoned no downstream, and severe table skew.

[0031] Before redundant data processing is complete, cluster space inventory and incremental trends can be calculated. Cluster space anomalies and governance calculations can be performed based on the early warning governance model. After data governance is implemented, the computing layer periodically calculates the cost and benefits of cluster space governance, including cluster space usage, incremental growth, anomalies, governance, and benefits. Finally, the analysis layer performs usage analysis, including obtaining data date fields, retrieving downstream table script content based on these fields, parsing SQL syntax data, extracting date expressions, and analyzing usage time ranges to optimize data processed by the model layer based on usage time ranges. During the application visualization phase, spatial data overviews can be viewed, including overall cluster distribution and trends, data zone distribution and trends, data zone incremental trends, team distribution and trends, team incremental space, management office distribution and trends, and management office incremental space. Visualized spatial alerts include data zone and data zone space alerts, anomaly model asset alerts, and cluster management alert reports. Visualized spatial governance includes governance task details, governance whitelists, cluster governance benefit reports, and table cleanup and splitting strategies.

[0032] Furthermore, the cluster space organization process Figure 2As shown, the cluster space governance task identification process includes data storage, inefficient model handling, data skew, and spatial anomaly handling at the model level. Data storage involves splitting and archiving hot data in the online area, warm data in the near-line area, and cold data in the historical repository. Inefficient model handling involves determining the absence of downstream applications by identifying the absence of table-level lineage relationships, dependencies, and EX warehouse operations. Access records and data update methods can also be used to determine the absence of downstream models. Data skew determination is based on tablespace size and actual skew rate. For example, a tablespace ≥100GB and a skew rate ≥60%; a tablespace ≥10GB and a skew rate ≥70%; and a tablespace ≥1GB and a skew rate ≥90% are all considered high skew rates. Spatial anomaly handling includes handling temporary tables, processing laboratory data, and addressing sudden increases in single tablespaces.

[0033] After identifying cluster space governance tasks, anomalies are pushed based on the cluster space early warning mechanism. Automatic acceptance of cluster governance tasks, i.e., processing of inefficient data, is then performed. A visual output of the cluster governance benefit report is then generated. Finally, a governance policy knowledge base can be established to analyze downstream usage ranges, including data lifecycle references, split cleanup strategies, historical library archiving strategies, space usage information, and slice statistics. The identification results are then fed back to the cluster space, achieving a closed-loop feedback loop for historical data.

[0034] Based on this, the embodiment of the present application provides a method for processing data clusters, referring to Figure 3 , Figure 3 This is a flow chart of the first embodiment of the method for processing data clusters of the present application.

[0035] In this embodiment, the data cluster processing method includes steps S10 to S40: Step S10: Obtain metadata of the Gaussian cluster, and calculate spatial asset information of the Gaussian cluster based on the metadata.

[0036] A Gaussian cluster refers to a server cluster built on the Gaussian distributed architecture, utilizing a multi-node collaborative computing model and supporting dynamic resource allocation and data sharding. Metadata describes structured information about the cluster's node topology, storage capacity, data distribution, and load status, including but not limited to node IP addresses, disk usage, number of shard replicas, data access frequency, database space, data tablespace storage, cluster sharding and splitting strategies, table-level lineage relationships, table-job mappings, and table scripts. Spatial asset information is a quantitative metric derived through metadata aggregation and calculation, including total storage capacity, percentage of valid data, redundant replica distribution, and hot and cold data partition status. Data warehouse tables of the flow, snapshot, and statistical types have corresponding data dates. The data range corresponding to these dates is called a single data slice. The distribution of warehouse data tables across the cluster depends on the table's distribution key (PI). Uneven distribution of the distribution key, known as table skew, can lead to data concentration on a few nodes, resulting in data skew and inefficient access.

[0037] In this embodiment, metadata of each node is pulled in real time through the cluster management interface, and the pulled metadata is stored in the database. Subsequently, the incremental storage situation and incremental trend of the node cluster space are calculated through the metadata, and the incremental usage of cluster space assets in the global area, data area, and management team dimensions, as well as the incremental trend of the cluster storage space, are obtained. Among them, the data area refers to the data hierarchy of the data warehouse, and data at different levels and data of different applications correspond to their own data areas. In this process, based on the cleaned metadata, the current space inventory indicators of the cluster are calculated, including the total capacity, the amount of valid data, the amount of redundant data, and the distribution of hot and cold data, and the total cluster capacity C is obtained. total =500TB, effective data volume V effective =150TB, redundant data volume V redundant =50TB, cold data accounts for P cold =35%. To calculate the incremental trend of cluster storage space, a time series model can be constructed based on historical metadata to predict the incremental trend of cluster storage space. For example, if storage usage increased from 400TB to 420TB over the past 30 days, after fitting the prediction model, it is predicted that it will increase to 445TB over the next 30 days (95% confidence interval: [440TB, 450TB]). This application does not limit the specific prediction process.

[0038] This embodiment calculates the spatial asset information of the current cluster through metadata, so as to identify the items to be managed of the Gaussian cluster based on the spatial asset information, thereby managing and reclaiming spatial resources and optimizing the resource utilization of the cluster space.

[0039] Step S20: construct an anomaly governance model based on the historical cluster processing rules corresponding to the metadata.

[0040] Historical cluster processing rules are a policy library formed by processing equipment based on operational experience or log analysis. For example, they determine whether node data is classified as hot, warm, and cold to improve the alignment of space growth with actual business needs. When a temporarily generated data table is detected, it is deleted. The system also detects whether there are abandoned models with no downstream data and whether table data is skewed. The anomaly governance model is built to address at least four common issues, including multiple sub-models such as slice redundancy, temporary table cleanup, cleanup of abandoned tables with no downstream data, and cleanup of severely skewed tables.

[0041] Therefore, in this embodiment, if the historical cluster processing rules include data storage and slice redundancy processing, a cluster slice and split strategy model is constructed; if the historical cluster processing rules include inefficient model processing, a discarded no downstream governance model is constructed; if the historical cluster processing rules include data skew processing, a table skew governance model is constructed; if the historical cluster processing rules include spatial anomaly processing, a temporary table cleanup model is constructed. It is understood that the anomaly governance model includes any combination of one or more of the above sub-models.

[0042] Step S30: determining the spatial redundancy management items of the spatial asset information according to the spatial redundancy management rules of the abnormality management model.

[0043] In this embodiment, the spatial redundancy governance rule is a specific governance strategy output by the abnormal governance model, and the spatial redundancy governance rule includes at least data slicing redundancy governance rules, tilted table cleaning rules, no downstream model cleaning rules, and temporary table cleaning rules. The spatial redundancy governance items are the data that need to be processed and the processing method thereof, including at least the data slicing redundancy splitting strategy, temporary table cleaning strategy, no downstream model cleaning strategy, and temporary table cleaning strategy. For example, the spatial redundancy governance items are to determine the distribution of the data in the online area and the near-line area on the day, to determine the number of data slices in the online area, whether the cleaning mark Y / N has changed, whether the cleaning policy configuration has been changed to the standard value, whether the total space occupied by the no downstream model is 0 for two consecutive days, whether the model has been offline, and whether the temporary table has been deleted.

[0044] It should be noted that the spatial redundancy governance rules are related to the sub-model of the anomaly governance model. If the sub-model includes a cluster slicing and splitting strategy model, the spatial redundancy governance rules include slicing redundancy governance rules. If it includes a table skew governance model, the spatial redundancy governance rules include skew table cleanup rules.

[0045] Specifically, when processing spatial asset information based on spatial redundancy governance rules, early warnings can be issued in cluster data areas, existing manageable items can be pushed, and auxiliary analysis and decision-making control increments can be provided. That is, metadata is first used to calculate the incremental usage information and cluster spatial trends of the cluster data area. Based on this incremental usage information and cluster spatial trends, the space allocation and usage status of the spatial asset information is determined. Based on this space allocation and usage, early warnings and detection processing are carried out for data areas that exceed the quota or are about to reach the quota. Therefore, when the space utilization rate of a data area is high, abnormal asset processing can be carried out on that data area.

[0046] In addition, redundancy detection and cleanup can be performed directly on data in all data zones based on governance rules.

[0047] Step S40: responding to the exception handling instruction corresponding to the space redundancy management item, updating the redundant data of the space asset information.

[0048] In this embodiment, when determining the data items that need to be processed, an early warning prompt is usually given based on the spatial redundancy management item, that is, a corresponding early warning report is generated and output to provide the operator with a control increment to assist in analysis and decision-making, so as to obtain the management operation selected by the operator based on the early warning report, and use the processing instruction corresponding to the operation as an exception processing instruction.

[0049] In addition, after generating the space redundancy management item, the relevant exception handling instructions can be automatically matched through the processing method of similar redundant management items in historical data, and the redundant data of the space asset information can be updated based on the exception handling instructions.

[0050] Redundant data in space asset information refers to data that needs to be cleared. The cluster interface can be called through the distributed transaction framework to perform redundant data deletion / migration operations, thereby improving the utilization of space resources. It should be noted that the data collection and redundant processing process can be displayed in a visual way, making the storage information and configuration information of each object in the cluster transparent, and visually monitoring the cluster space usage and space increment trends from the dimensions of data area and management room, so that cost awareness can penetrate into every developer and realize the visualization and transparency of cluster space management.

[0051] This embodiment provides a method for processing data clusters. By dynamically calculating spatial asset information and building a data-driven anomaly management model, it achieves precise management of Gaussian cluster storage resources, provides accurate decision-making support for application operators, and processes redundant data based on exception handling instructions to improve cluster resource utilization.

[0052] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above introduction, and no further details will be given later. On this basis, when determining the space redundancy management items, it is necessary to first determine the area with high space utilization rate, and then perform data redundancy processing based on the area. Specifically, please refer to Figure 4 , step S30 includes steps S31 to S33: Step S31, obtaining the incremental usage information of the space asset information in the global area, data area and management area and the cluster space trend.

[0053] In this embodiment, the incremental usage information of the spatial asset information in the global area, data area and management area and the cluster spatial trend are first calculated through metadata, so as to calculate the spatial anomaly based on this information.

[0054] Step S32: determining, based on the incremental usage information and the cluster space trend, a warning data area whose space usage is higher than a preset usage rate.

[0055] After obtaining the incremental usage information and cluster space trends of multiple different data areas, space anomaly calculations are performed, and data areas that exceed the standard or are about to be full are detected, that is, warning data areas whose space utilization rate is higher than the preset utilization rate are determined.

[0056] Step S33: Determine the space redundancy management item of the warning data area according to the slice redundancy management rule, the tilted table cleaning rule, the no downstream model cleaning rule, and the temporary table cleaning rule.

[0057] In this embodiment, after the early warning data area is determined, the spatial redundancy management items in the early warning data area may be determined based on the spatial redundancy management rules.

[0058] Specifically, the spatial redundancy management item includes the slice redundancy splitting strategy of the warning data area, so the slice redundancy splitting strategy of the warning data area can be determined according to the slice redundancy management strategy. For example, please refer to Figure 5 You can first obtain data access information for the warning data area within a preset period, and then mark the data temperature of the warning data area based on the data access information. For example, according to the frequency and level of data occurrence, the data can be marked as hot, warm, and cold data, corresponding to the cluster online area, near-line area, and historical database. Then, based on the data temperature and slice redundancy governance rules, the slice redundancy splitting strategy is determined. For example, based on the table form and data slice statistics, the non-compliant governance items are detected as shown in the following table:

[0059] Optionally, spatial redundancy management includes a skewed table cleanup strategy. Therefore, the skew rate of the data table in the warning data area can be obtained. Based on the skew rate and the skewed table cleanup rules, a temporary table cleanup strategy for the warning data area can be determined. It is understood that table skew prevents data from being written to severely skewed data nodes, making it a key management item. For example, the temporary table cleanup rules are shown in the following table:

[0060] A skewed table cleanup strategy requiring data cleanup is determined based on the temporary table cleanup rule.

[0061] Optionally, the spatial redundancy management item includes a no-downstream model cleanup strategy. In this process, the no-downstream model cleanup strategy can be determined based on the table-level lineage relationship of the warning data area and the no-downstream model cleanup rules. For example, the no-downstream model (also known as inefficient model) cleanup rules are as follows: Figure 5 As shown, the table-level lineage relationship of spatial asset information (Gauss cluster full model) is first determined. Then, based on the discarded no-downstream model, the model's downstream dependency information corresponding to the marked lineage relationship is determined, including whether a downstream model exists. If the model's downstream dependency information indicates no downstream model exists, the model is judged to have no downstream model. If the model's downstream dependency information indicates a downstream model exists, further judgment is made based on the model's own dependencies and EX (Excel) outbound operation information. Therefore, if the model's downstream dependency information indicates a downstream model exists, the current recognition model's own dependencies and outbound operation information are determined. If the recognition model has its own dependencies and no outbound operation exists, the current recognition model is judged to have no downstream model. If the model has its own dependencies, it is not a no-downstream model. Similarly, if an outbound operation exists, the detected model is also determined to have no downstream model. Furthermore, whether a model has no downstream model can be determined based on the model's access status and data update status.

[0062] Optionally, the space redundancy management item includes a temporary table cleanup strategy. In this case, the temporary table data in the warning data area can be determined first, and then the temporary table cleanup strategy can be determined based on the temporary table cleanup rules and temporary table data. For example, in the temporary table cleanup strategy, the production backup temporary table data occupies a large space and needs to be pushed for cleanup on a daily basis.

[0063] For example, please refer to Figure 6When performing space usage calculations, we obtain information about the global space asset area (i.e., A / B clusters (Gaussian A and Gaussian B clusters), the data area, and the management area (i.e., the management team) for incremental usage and cluster space trends. This includes the total space, nearline space, online space, and spatial trends for the A / B clusters. For the data area, this includes spatial distribution, monthly incremental space trends, and spatial trends. For the management team, this includes spatial distribution, incremental space, and spatial trends. After obtaining these parameters, we perform spatial anomaly calculations. We first identify data areas or clusters with excessive space utilization, then calculate abnormal data assets for data with excessive space utilization. This includes calculating information such as slice redundancy, temporary table cleanup, discarded downstream models, and severe table skew, so that we can perform anomaly management based on these identified anomaly management items.

[0064] This embodiment provides a data cluster processing method. During the process of cleaning redundant data in the warning data area where the space usage rate is higher than the normal usage rate, an early warning and notification mechanism is established based on the data area and table object, providing a large-scale stock model governance solution. In the absence of incremental space allocation, it meets the requirements of new business development and the smooth parallel switching of old and new warehouses. In addition, different types of cleanup strategies for the warning data area are calculated through multiple different categories of sub-models of the abnormal governance model, deepening the rationality evaluation of the incremental model and solution guidance, providing diagnosis, assistance and recovery of cluster space for data area personnel, and improving cluster space utilization.

[0065] Based on the second embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the second embodiment can be referred to above and will not be described in detail. On this basis, after determining the spatial redundancy quality items of the spatial asset information, a spatial early warning patrol system can also be constructed to patrol the data area and database data, and after locating the manageable items of the positioning space anomaly, periodic early warning reports are pushed. Therefore, after step S30, it also includes: Step S50: Obtain remaining space information of the warning data area.

[0066] In this embodiment, please continue to refer to Figure 6 ,In the process of spatial anomaly calculation, after the anomaly governance items are calculated ,through the anomaly governance model, the remaining space information of the ,warning data area is obtained. The remaining space information ,includes the available estimation of remaining space, the actual ,remaining space or daily space increment, etc., so as to generate cluster management ,warning report data through the remaining space information.

[0067] Step S60: Generate and output a cluster warning report of the space asset information based on the remaining space information and the space redundancy management items.

[0068] In this embodiment, after detecting data assets with excessively high space utilization based on the governance model, these abnormal governance items and remaining space information are combined to generate cluster management warning report data. This includes generating warnings for excessive data slice redundancy and abnormal notifications for excessive temporary tables requiring deletion. The abnormal data information and cluster management warning report can be visualized in the display interface.

[0069] Step S70: Obtain an execution instruction corresponding to the warning information, and set the execution instruction as the exception handling instruction.

[0070] In this embodiment, after outputting the warning report in the visual interface, the operator can select the operation to be performed based on the actual content of the warning report, such as selecting redundant outputs to be cleared in the visual interface, or setting a scheduled cleanup task, etc. The control system of the processing equipment of the data cluster will respond to the user's operation information and execute the execution instructions corresponding to the operation information.

[0071] This embodiment provides a data cluster processing method, which outputs and displays cluster warning reports in a visual manner, thereby intelligently pushing warning reports, making cluster space resource allocation and management transparent and digital, and improving cluster resource utilization.

[0072] Based on the first embodiment of the present application, in the fourth embodiment of the present application, the same or similar content as the above-mentioned first embodiment can be referred to the above introduction, and no further details will be given later. On this basis, after performing cluster space anomaly and governance calculations based on the early warning governance model, the cluster space governance cost-benefit can also be periodically counted, and the space governance cost-benefit can be visualized. Therefore, after step S40, steps S70~S90 are also included: Step S70: obtaining redundant data processing tasks of the space asset information.

[0073] Step S80: Determine the spatial governance benefit information of the spatial asset information based on the redundant data processing task.

[0074] In this embodiment, the redundant data management task refers to the task of clearing redundant data. Figure 7The details of the space governance tasks include data storage, inefficient model processing, data redundancy processing and temporary table processing. After scanning the SQL database, the space redundancy processing rules are determined, including splitting strategy, cleaning strategy, no downstream, temporary table and data skew, etc. Then, redundant data is cleaned based on these strategies, including judging the distribution of the day's data in the online area and near-line area, judging the number of data slices in the online area, whether the cleaning mark Y / N has changed, whether the cleaning strategy configuration has been changed to the standard value, whether the total space occupied by the no downstream model is 0 for two consecutive days, whether the model has been offline and whether the temporary table has been deleted. The governance items automatically cleaned by the system are automatically cleared. The data table skew will have a certain calculation error based on the data cleaning, and provide decision-making references based on data skew.

[0075] Finally, after the redundant data processing task is completed, during the data storage process, the space saved in the online and near-line areas is calculated based on the two dates of task generation and completion, and compared with the total space occupied before and after determining that there is no downstream model governance. At the same time, temporary tables are automatically cleaned up and used as a decision reference. The space governance benefits of temporary tables that have not been cleaned for a long time are calculated based on the actual occupied space.

[0076] Specifically, the formula for calculating spatial governance benefit information based on redundant data processing tasks is as follows:

[0077] in, The total number of slices generated by the task, is the total number of slices completed by the task, A single slice occupies space, It is the average space occupied by a single table slice.

[0078] For example, after the Gauss cluster space management was put into production, more than 7,000 problem items requiring management were detected, including more than 6,500 objects with no downstream and occupying more than 1G of space, saving more than 350T of space. From the second month, space warning emails were sent when the data zone utilization rate was greater than or equal to 95%, and detailed items of the data zone that could be managed were pushed. Using the retail customer management team as a pilot, seven large tables were managed, saving 25T of space. The remaining space in the near-line area of ​​cluster A was around 100T, and the management of existing large tables was effective. In the second month, a cluster management benefit report was pushed to each data zone to promote cost awareness. By the second month, a cumulative management benefit of 34T was achieved.

[0079] Step S90: outputting the space management income information and the space management report of the space asset information.

[0080] For example, a spatial governance report such as Figure 8As shown, the space governance report includes the total space, space usage between the current month and the previous month in the online area and the near-line area, as well as the minutes of the current month (e.g., June), the number of new tasks in the current month, the number of new completed tasks in the current month, the new completion rate in the current month, the total number of tasks, the total number of completed tasks, and the overall completion rate, etc. At the same time, the total space saved in the current month is 7769.75G, and the overall space saved is 32.88T.

[0081] Therefore, in this embodiment, after cluster space anomalies and governance calculations are performed based on the early warning governance model, the cluster space governance cost-benefit is periodically counted, and the cluster space governance effect is visualized and output to improve the cluster's spatial redundancy processing effect.

[0082] Based on the first embodiment or the fourth embodiment of the present application, in the fifth embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction and will not be repeated hereafter. On this basis, after performing cluster space anomaly and governance calculations based on the early warning governance model, a feedback mechanism can also be set up, and a historical data retention strategy can be set up to optimize the meaningless slice governance through historical governance data to form a feedback closed-loop processing of redundant data in the cluster space, so that in the subsequent slice redundancy governance and development process, the slice retention strategy can be reasonably set with reference to specific historical data such as data date. In addition, after generating the cluster space governance report, historical data information can be obtained based on the feedback mechanism, and the subsequent redundant data governance process can be optimized based on the historical data information.

[0083] Therefore, as an optional implementation, after step S90, the data date field of the processed data table can also be obtained, and then the downstream table of the data table and the script corresponding to the downstream table can be obtained. The script is then split according to the statement block and the grammar is rationalized. The processed statement block is then subjected to grammatical analysis to generate an AST abstract syntax tree, and the Visitor mode processing AST is customized in the actual scenario of space management. The AST includes: SQLExprTableSource, which is used to obtain the table alias, splice it with the date field, and match it. The field information extracted from table A is "a. dw_dat_dt"; SQLBetweenExpr / SQLBinaryOpExpr / SQLInListExpr, by capturing the Between, operator (Op), and In keywords, extracting complex expressions, recursively disassembling them, and forming a "operator as the central axis, the left end is the target field, and the right end is the prime number characteristic expression value" paradigm. Finally, the variables in the expression are converted into a general standard variable expression containing the date according to the variable definition description in the script, and the expression date scope is determined according to the script data date. The transformation conversion code is as follows:

[0084] Taking July 16th as an example, the script's data date is T+1. If ${TX_DATE} is replaced with "20xx0715," Expression A actually represents the date range "=20xx0714." Therefore, under the influence of Table B, Table A's date scope is "20xx0714." In practice, Table A has multiple downstream tables, each with multiple corresponding scripts. The union of these tables is the actual date scope of Table A.

[0085] The above parameters are for illustration only and are not intended to limit the present application.

[0086] The present application provides a data cluster processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data cluster processing method of the first embodiment above.

[0087] Reference below Figure 9 , which shows a schematic diagram of the structure of a data cluster processing device suitable for implementing an embodiment of the present application. The data cluster processing device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), and the like, as well as fixed terminals such as digital TVs and desktop computers. Figure 9 The processing device of the data cluster shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0088] like Figure 9As shown, the data cluster's processing device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the data cluster's processing device. Processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems may be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. Communication devices 1009 may allow the processing equipment of the data cluster to communicate with other devices wirelessly or by wire to exchange data. While the diagram illustrates the processing equipment of the data cluster with various systems, it should be understood that implementation or presence of all of the illustrated systems is not required. More or fewer systems may alternatively be implemented or present.

[0089] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0090] The data cluster processing device provided in this application utilizes the data cluster processing method described in the above-mentioned embodiments to address the technical issue of low cluster resource utilization. Compared to the prior art, the data cluster processing device provided in this application achieves the same beneficial effects as the data cluster processing method described in the above-mentioned embodiments. Other technical features of the data cluster processing device are the same as those disclosed in the above-mentioned embodiments and are not further elaborated here.

[0091] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0092] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0093] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, wherein the computer-readable program instructions are used to execute the data cluster processing method in the above-mentioned embodiment.

[0094] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0095] The computer-readable storage medium may be included in the processing device of the data cluster; or may exist independently without being assembled into the processing device of the data cluster.

[0096] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by a data cluster processing device, the data cluster processing device is caused to: obtain metadata of a Gaussian cluster, and calculate spatial asset information of the Gaussian cluster based on the metadata; Building an anomaly governance model based on historical cluster processing rules corresponding to the metadata; Determining the spatial redundancy governance items of the spatial asset information according to the spatial redundancy governance rules of the anomaly governance model; In response to the exception handling instruction corresponding to the space redundancy management item, the redundant data of the space asset information is updated.

[0097] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0099] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0100] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned data cluster processing method, thereby resolving the technical issue of low cluster resource utilization. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the data cluster processing method provided in the aforementioned embodiments, and are not further elaborated here.

[0101] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A data cluster processing method, characterized in that: The data cluster processing method includes: Obtaining metadata of a Gaussian cluster, and calculating spatial asset information of the Gaussian cluster based on the metadata; Building an anomaly governance model based on historical cluster processing rules corresponding to the metadata; Determining the spatial redundancy governance items of the spatial asset information according to the spatial redundancy governance rules of the anomaly governance model; In response to the exception handling instruction corresponding to the space redundancy management item, the redundant data of the space asset information is updated.

2. The data cluster processing method according to claim 1, wherein: The step of constructing an anomaly governance model according to the historical cluster processing rules corresponding to the metadata includes: If the historical cluster processing rules include data storage and slice redundancy processing, a cluster slice and split strategy model is constructed; If the historical cluster processing rules include inefficient model processing, construct an abandoned no downstream governance model; If the historical cluster processing rules include data skew processing, build a table skew governance model; If the historical cluster processing rules include spatial exception processing, a temporary table cleanup model is constructed.

3. The data cluster processing method according to claim 1, wherein: The step of determining the spatial redundancy management items of the spatial asset information according to the spatial redundancy management rules of the abnormality management model includes: Obtaining the incremental usage information of the space asset information in the global zone, data zone, and management zone, as well as the cluster space trend; Determine, based on the incremental usage information and the cluster space trend, a warning data area whose space usage is higher than a preset usage rate; The space redundancy management item of the warning data area is determined according to the slice redundancy management rule, the tilted table cleaning rule, the no downstream model cleaning rule and the temporary table cleaning rule.

4. The data cluster processing method according to claim 3, wherein: The step of determining the space redundancy management item of the warning data area according to the slice redundancy management rule, the tilted table cleaning rule, the no downstream model cleaning rule, and the temporary table cleaning rule includes: Determine a slice redundancy splitting strategy for the warning data area according to the slice redundancy management rule; Obtaining a tilt rate of a data table in the early warning data area, and determining a tilt table cleanup strategy for the early warning data area based on the tilt rate and the tilt table cleanup rule; Determine a no-downstream model cleanup strategy based on the table-level lineage relationship of the warning data area and a no-downstream model cleanup rule; A temporary table clearing strategy is determined according to the temporary table clearing rule and the temporary table data in the warning data area.

5. The data cluster processing method according to claim 4, characterized in that: The step of determining the slice redundancy splitting strategy of the warning data area according to the slice redundancy management rule includes: Acquiring data access information of the early warning data area within a preset time period, and marking the data temperature of the early warning data area according to the data access information; The slice redundancy splitting strategy is determined according to the data temperature and the slice redundancy management rule.

6. The data cluster processing method according to claim 4, characterized in that: The step of determining the no-downstream model of the spatial asset information according to the abandoned no-downstream model and the table-level lineage relationship of the spatial asset information includes: Determine the table-level lineage relationship of the spatial asset information, and determine the model downstream dependency information corresponding to the table-level lineage relationship based on the abandoned no-downstream model; If the model downstream dependency information indicates that a downstream exists, determining the current recognition model's own dependency and cargo outbound operation information; If the recognition model has its own dependencies and there is no cargo outbound operation, it is determined that the current recognition model has no downstream model.

7. The data cluster processing method according to claim 3, characterized in that: After the step of determining the spatial redundancy management items of the spatial asset information according to the spatial redundancy management rules of the abnormality management model, the data cluster processing method further includes: Obtaining remaining space information of the warning data area; Generate and output a cluster warning report of the space asset information based on the remaining space information and the space redundancy management items; Obtain an execution instruction corresponding to the cluster warning report, and set the execution instruction as the exception handling instruction.

8. The data cluster processing method according to claim 1, wherein: After the step of responding to the exception handling instruction corresponding to the space redundancy management item and updating the redundant data of the space asset information, the data cluster processing method further includes: A redundant data processing task for obtaining the spatial asset information; Determining spatial governance benefit information of the spatial asset information according to the redundant data processing task; Output the space management income information and the space management report of the space asset information.

9. A data cluster processing device, characterized in that: The data cluster processing device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data cluster processing method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the data cluster processing method according to any one of claims 1 to 8 are implemented.