A data management method and system for reducing computing power occupancy

By applying sparse representation and principal component analysis, dynamic partition optimization, event-driven filtering and heterogeneous computing task scheduling methods in data management, the problems of high-dimensional data redundancy, static partition management, low event screening efficiency and uneven task scheduling in the existing technology are solved, and efficient data processing and resource utilization are achieved.

CN119440776BActive Publication Date: 2025-05-30SHANDONG HUITONG TECH CO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510037988.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-30
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

In data management, the existing technology has not been fully processed, resulting in feature extraction relying on fixed dimensionality reduction algorithms, making it difficult to dynamically optimize, increasing computational complexity and time consumption; the partition management strategy is static and cannot be adjusted in real time according to the access frequency changes, resulting in extended data access paths and high communication overhead across nodes; event data screening lacks an efficient filtering mechanism, resulting in long processing processes and high computing power overhead; task scheduling fails to combine computing density and dynamic changes in equipment load, resulting in uneven resource allocation and low overall computing power utilization.

Method used

Through a sparse representation algorithm based on data distribution and principal component analysis method, the eigenvectors with the top 30% contribution rate are extracted to generate sparse low-dimensional matrix; a dynamic partition optimization algorithm is used to construct a data access frequency distribution model and real-time monitoring to dynamically adjust the partition weight; in the event filtering stage, the Bloom filter and the hash function matching mechanism of the rule table are used to quickly filter key event data; in task scheduling, light tasks and intensive tasks are reasonably allocated through the calculation density analysis of task features and the dynamic regulation of the load balancing model.

Benefits of technology

It effectively reduces the data dimension, reduces computing resource consumption, improves data processing efficiency and accuracy; significantly improves the locality of data access and the rationality of storage allocation, reduces the computing overhead and communication delay caused by cross-partition access; realizes rapid screening and accurate extraction of event data, reduces the processing volume of irrelevant data; improves the efficiency and resource utilization of task scheduling, and reduces resource conflicts and task delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119440776B_ABST
    Figure CN119440776B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of computer data management, and discloses a data management method and system for reducing computing power occupancy, including S1: Based on data distribution analysis, using a sparsification representation algorithm and a principal component analysis method, the redundancy of high-dimensional data features is calculated and removed by setting L1 regularization constraint conditions, and the data correlation matrix is decomposed in combination with matrix operation methods, and the eigenvectors with the top-ranked contribution values are extracted and converted into a sparse low-dimensional matrix; a sparse low-dimensional data matrix is generated. Through the combination of sparsification representation and principal component analysis, L1 regularization is innovatively used to screen features of high-dimensional data, redundant data items are removed, and the principal component features with high correlation are extracted, effectively reducing the data dimension, reducing the consumption of computing resources, while retaining the main information characteristics of the data, improving the efficiency and accuracy of subsequent processing, and realizing the distributed scheduling of tasks and the efficient utilization of device computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data management, and particularly to a data management method and system for reducing computing power occupancy. Background Art

[0002] Computer data management is a core technical field in computer science, mainly studying how to efficiently store, organize, process, and access data. With the rapid development of technologies such as big data, artificial intelligence, and cloud computing, data management technology has gradually evolved from traditional static storage and processing to dynamic, intelligent, and distributed management, covering multiple directions such as database management systems, distributed storage, data stream processing, and data compression and encryption. Its goal is to, in the case of rapid growth of data volume, improve the data processing efficiency and the computing performance of the system by optimizing the data organization structure and access mechanism, while reducing resource occupancy, and ensuring the security, reliability, and scalability of data. It is widely applied to various high-concurrency and high-performance computing scenarios. Among them, a data management method and system for reducing computing power occupancy refers to designing an efficient data management mechanism and task allocation strategy to reduce the consumption of computing resources in the data processing process, while improving the data storage and processing efficiency. This method is applicable to scenarios that require optimizing computing power allocation and can be used in fields such as big data analysis, real-time computing, cloud service optimization, and intelligent task scheduling, which helps to reduce energy consumption, save resources, and achieve an overall improvement in system performance.

[0003] In the existing technology for data management, due to the fact that the redundancy characteristics of high-dimensional data are not fully processed, feature extraction usually relies on fixed dimensionality reduction algorithms and it is difficult to dynamically optimize according to different data distributions, resulting in too high data dimensions, increasing the complexity and time consumption of subsequent calculations. In partition management, the partition strategy is mostly static partitioning and fails to adjust in real time according to the change of access frequency, resulting in an extended data access path and a large cross-node communication overhead, affecting the data processing efficiency. In event data screening, the existing methods often lack an efficient filtering mechanism and usually compare rules by traversing all data step by step, with a long processing process and a high computing power overhead, and it is also difficult to guarantee the accuracy. In the task scheduling stage, the allocation method of tasks and devices fails to combine the dynamic changes of computational intensity and device load, easily causing uneven resource allocation, resulting in some devices being overloaded or vacant, and the overall computing power utilization rate being low. These deficiencies make the existing technology show a high degree of resource waste when facing large-scale data processing or real-time computing requirements and cannot meet the efficient and flexible computing needs. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention provides a data management method and system for reducing computing power occupancy, which solves the problems in the prior art that in data management, due to the redundant characteristics of high-dimensional data not being fully processed, feature extraction usually relies on fixed dimensionality reduction algorithms and is difficult to dynamically optimize according to different data distributions, resulting in too high data dimensions, increasing the complexity and time consumption of subsequent calculations. In partition management, the partition strategy is mostly static partitioning and fails to adjust in real time according to changes in access frequency, resulting in an extended data access path, a large cross-node communication overhead, and affecting data processing efficiency.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A data management method for reducing computing power occupancy, comprising the following steps:

[0006] S1: Based on data distribution analysis, using a sparsification representation algorithm and a principal component analysis method, setting L1 regularization constraint conditions, performing operations to eliminate the redundancy of high-dimensional data features, decomposing the data correlation matrix through matrix operations, extracting the eigenvectors with the top 30% contribution rates, and converting them into a sparse low-dimensional matrix; generating a sparsified low-dimensional data matrix;

[0007] S2: Based on the sparsified low-dimensional data matrix, adopting a dynamic partition optimization algorithm, screening the data sets with high access frequencies by constructing a data access frequency distribution model, allocating the data to different storage partitions according to weights in combination with the distributed node task table, dynamically adjusting the partition weights based on real-time monitoring and updating the allocation results; generating an optimized dynamic data partition table;

[0008] S3: Based on the optimized dynamic data partition table, adopting an event-driven filtering algorithm, screening the key event data sets that meet the conditions by constructing a Bloom filter, judging each item of the data label according to the defined rule table, and extracting and storing the data marked with the key event identifier for the data that hits the conditions; generating a key event data stream;

[0009] S4: Based on the key event data stream, adopting a heterogeneous computing task scheduling algorithm, classifying and marking the task eigenvectors according to the computational intensity, allocating the light tasks to the CPU parallel queue for processing according to the task load analysis results, allocating the high-intensity tasks to the GPU batch queue for processing, and adjusting the distribution of computing tasks on each device in combination with the dynamic load model; generating an efficiently processed computing result set.

[0010] Preferably, based on data distribution analysis, using a sparsification representation algorithm and a principal component analysis method, performing operations to eliminate the redundancy of high-dimensional data features by setting L1 regularization constraint conditions, decomposing the data correlation matrix through matrix operations, extracting the eigenvectors with higher contribution values and converting them into a sparse low-dimensional matrix; the specific steps for generating a sparsified low-dimensional data matrix are:

[0011] S101: Based on the high-dimensional dataset, using the L1 regularization constraint algorithm, by constructing an objective function, taking the sparsity of the eigenvalue as a constraint condition, calculating each feature in the data matrix column by column, eliminating the feature terms with weight values tending to zero under the constraint condition, and generating a preliminary sparse feature matrix;

[0012] S102: Based on the preliminary sparse feature matrix, using the covariance matrix calculation method, constructing a feature correlation matrix according to the covariance values between features, obtaining the eigenvector matrix through eigenvalue decomposition, and extracting the eigenvectors with the top-ranked weights as the principal components to generate an eigenvector matrix;

[0013] S103: Based on the eigenvector matrix, through the matrix projection method, projecting the high-dimensional data into the principal component space, reducing the feature dimension according to the principal component contribution rate, and eliminating the feature dimensions with principal component contribution rates less than the preset threshold to generate a principal component low-dimensional matrix;

[0014] S104: Based on the principal component low-dimensional matrix, using the sparse matrix mapping method to optimize the matrix sparsity, retaining the sparse features with weights in the top 30% in the principal components, outputting the final sparse matrix result, and generating a sparsified low-dimensional data matrix.

[0015] Preferably, based on the sparsified low-dimensional data matrix, using the dynamic partitioning optimization algorithm, screening the high-frequency accessed datasets by constructing a data access frequency distribution model, allocating the data to different storage partitions according to the weights in combination with the distributed node task table, dynamically adjusting the partition weights based on real-time monitoring and updating the allocation results; the specific steps for generating the optimized dynamic data partition table are as follows:

[0016] S201: Based on the sparsified low-dimensional data matrix, construct a data access frequency model, according to the data access log records, count the access times and time intervals of each data block, calculate the access frequency and generate the data block access frequency ranking to generate a data access frequency model;

[0017] S202: Based on the data access frequency model, using the K-means clustering algorithm, allocate the data blocks with access frequencies in the top 30% to the priority partition, allocate the data blocks with access frequencies in the bottom 30% to the secondary partition, mark and record the initial state of the partition to generate an initial data partition table;

[0018] S203: Based on the initial data partition table, through the real-time access monitoring method, capture the dynamic access situation of the data blocks during the operation process, adjust the partition state according to the latest access frequency, update the data partition distribution, and generate a dynamically adjusted partition table;

[0019] S204: Based on the dynamically adjusted partition table, the adjusted data partition results are written into the distributed storage nodes through the distributed node storage management method, the partition optimization storage is completed, and the optimized dynamic data partition table is generated.

[0020] Preferably, based on the optimized dynamic data partition table, an event-driven filtering algorithm is used to construct a Bloom filter to filter the key event data set that meets the conditions, and the data tags are judged item by item according to the defined rule table. The key event identifiers are marked for the data that hit the conditions and then extracted and stored; the specific steps of generating the key event data stream are:

[0021] S301: Based on the optimized dynamic data partition table, a Bloom filter is constructed, a rule table of key events is defined, rule conditions are converted into a hash function of the Bloom filter, and potential event features of the data block are calculated using the formula: ; Mark data blocks as potential events by calculating item by item; Calculate the probability of potential events , generating potential event data after Bloom filtering; among them, Representative The potential event probability of each data block, Representative The weighted sum of the key eigenvalues ​​in the data blocks, Represents the mean of the weighted sum of the key feature values ​​of all data blocks, Represents the total number of data blocks, Represents the sum of squares of the weighted sum of the eigenvalues ​​of all data blocks and the mean value, Representative The regular hash value of each data block;

[0022] S302: Based on the potential event data after Bloom filtering, the event feature values ​​in the data blocks are retrieved one by one through the key event matching method to determine whether the definition conditions of the rule table are met, and the data blocks that meet the conditions are screened to generate key event data blocks;

[0023] S303: Based on the key event data block, an event priority labeling method is used to classify and label the key event data block, generate event priority labels according to event importance, and generate key event data with labeled priorities;

[0024] S304: Based on the key event data with marked priorities, the data is classified and stored in the designated storage node according to the priorities through the data extraction and storage algorithm, and the extraction and storage of the key event data are completed to generate the key event data.

[0025] Preferably, based on the critical event data stream, a heterogeneous computing task scheduling algorithm is adopted. By classifying and marking the computational intensity of the task feature vectors, light tasks are assigned to the CPU parallel queue for processing according to the task load analysis results, and high-intensity tasks are assigned to the GPU batch queue for processing. The computational task distribution of each device is adjusted in combination with the dynamic load model. The specific steps for generating an efficiently processed computational result set are as follows:

[0026] S401: Based on the critical event data stream, classify the task feature vectors. By calculating the computational complexity and data dependency of the computational tasks, mark the tasks as lightweight tasks or intensive computational tasks, and generate the classified and marked task feature vectors;

[0027] S402: Based on the classified and marked task feature vectors, through the load balancing analysis model, calculate the current load and available computing power of each computing device, and assign the tasks to the appropriate computing devices in combination with the task characteristics to generate a task-device allocation table;

[0028] S403: Based on the task-device allocation table, through the parallel task scheduling algorithm, assign lightweight tasks to the CPU parallel queue and intensive computational tasks to the GPU batch computing queue to complete the task scheduling and generate the allocated computational task queue;

[0029] S404: Based on the allocated computational task queue, through the device execution and feedback mechanism, calculate each task item by item to output the processing results, and generate the final task processing result set to generate an efficiently processed computational result set.

[0030] A data management system for reducing computing power occupancy includes the following modules: a feature extraction module, a dynamic partitioning module, an event filtering module, and a task scheduling module;

[0031] The feature extraction module, based on the high-dimensional data set, adopts the L1 regularization constraint algorithm to calculate the feature sparsity column by column to eliminate the feature items with weights tending to zero, combines the covariance matrix to calculate the correlation between features to extract feature vectors, uses the matrix projection method to transform the high-dimensional data into the principal component space and eliminates the feature dimensions with contribution rates less than the preset value, optimizes the matrix sparsity and then outputs the main features to generate a sparsified low-dimensional data matrix;

[0032] The feature extraction module includes a sparsification processing sub-module, a feature vector extraction sub-module, and a principal component transformation sub-module;

[0033] The dynamic partitioning module is based on the sparse low-dimensional data matrix. It builds a data access frequency model by analyzing the access logs, uses the K-means clustering algorithm to allocate high-frequency access data to priority partitions and generate a partition table. It combines real-time monitoring to capture dynamic data access and adjust the partition status. It writes the adjusted partitions to the storage nodes through the distributed node management method to generate an optimized dynamic data partition table.

[0034] The dynamic partitioning module includes a frequency modeling submodule, a partition clustering submodule, and a partition dynamic adjustment submodule;

[0035] The event filtering module builds a Bloom filter based on the optimized dynamic data partition table to filter potential events. It marks potential events by calculating the data feature values ​​and matching them with the hash function of the event rules one by one. It uses the key event matching method to filter the data blocks that meet the conditions and mark the key events with priority tags. Finally, it uses the extraction and storage algorithm to store the event data to the target node according to the priority, generating a key event data stream.

[0036] The event filtering module includes a Bloom filtering submodule, an event feature matching submodule, and a priority marking submodule;

[0037] Task scheduling module: Based on the key event data stream, the task density and data dependency are calculated in combination with the task feature vector to classify the tasks. The available computing power of the computing device is calculated through the load balancing analysis model to dynamically generate a task allocation table. The parallel task scheduling algorithm is used to allocate lightweight tasks to the CPU parallel queue and high-density tasks to the GPU batch queue. After completing the task, the final result is output to generate a set of calculation results for efficient processing.

[0038] The task scheduling module includes a task feature classification submodule, a task load balancing submodule, and a parallel queue scheduling submodule.

[0039] Preferably, the sparse processing submodule, based on a high-dimensional data set, adopts an L1 regularization constraint algorithm, constructs an objective function, takes the sparsity of the eigenvalue as a constraint condition, calculates the feature weights in the data matrix column by column, removes feature items whose weights tend to zero, and generates a preliminary sparse feature matrix;

[0040] The feature vector extraction submodule uses the covariance matrix calculation method based on the preliminary sparse feature matrix to calculate the covariance values ​​between features to construct the feature correlation matrix, extracts the feature vector matrix through matrix decomposition, and selects the feature vectors with the highest weight ranking to generate the feature vector matrix;

[0041] The principal component transformation sub-module projects high-dimensional data into the principal component space based on the eigenvector matrix through matrix projection methods, screens the feature dimensions according to the principal component contribution rate, eliminates the feature dimensions with a contribution rate less than the preset threshold, optimizes the matrix sparsity, and then outputs the result to generate a sparse low-dimensional data matrix.

[0042] Preferably, the frequency modeling sub-module analyzes the data access log based on the sparse low-dimensional data matrix, calculates the access frequency by statistically counting the access times and time intervals of each data block, generates an access frequency sorting model, and generates a data access frequency model.

[0043] The partition clustering sub-module adopts the K-means clustering algorithm based on the data access frequency model. Through the clustering calculation of the data block access frequency, it assigns the data blocks with access frequencies in the top 30% to the priority partition and the data blocks with access frequencies in the bottom 30% to the secondary partition, generating an initial partition table and an initial data partition table.

[0044] The partition dynamic adjustment sub-module captures the dynamic access situation of data blocks through real-time monitoring based on the initial data partition table, updates the access frequency data, adjusts the partition status according to the latest access frequency, and optimizes the partition structure to generate an optimized dynamic data partition table.

[0045] Preferably, the Bloom filter sub-module constructs a Bloom filter based on the optimized dynamic data partition table. By defining a rule table for key events, it calculates the eigenvalue hash function of each data block and marks the potential event data blocks that meet the rules as potential events, generating Bloom-filtered potential event data.

[0046] The event feature matching sub-module adopts the key event matching method based on the Bloom-filtered potential event data, retrieves the event feature values one by one, and determines whether they meet the conditions of the rule table, screening out the event data blocks that meet the conditions to generate key event data blocks.

[0047] The priority annotation sub-module classifies the event priorities of data blocks based on the key event data blocks, performs label annotation according to the event importance, and generates key event data with priority annotation, generating key event data with annotated priorities.

[0048] Preferably, the task feature classification sub-module classifies the computational intensity of the task feature vector based on the key event data with annotated priorities, and marks it as a lightweight task or a dense task according to the computational complexity and data dependency of the task, generating a task feature vector with classification marks.

[0049] The task load balancing sub-module calculates the current load and available computing power of each device based on the task feature vector with classification tags, combines with the load balancing analysis model, generates the allocation table of tasks and devices, and generates the task-device allocation table.

[0050] The parallel queue scheduling sub-module, based on the task-device allocation table, adopts a parallel task scheduling algorithm to allocate lightweight tasks to the CPU parallel queue and intensive tasks to the GPU batch queue, completes task scheduling, and generates the allocated computing task queue.

[0051] The present invention provides a data management method and system for reducing computing power occupancy. It has the following beneficial effects:

[0052] 1. Through the combination of sparse representation and principal component analysis, the present invention innovatively uses L1 regularization to screen features of high-dimensional data, eliminates redundant data items, extracts principal component features with high correlation, effectively reduces the data dimension, reduces the consumption of computing resources, while retaining the main information characteristics of the data, and improves the efficiency and accuracy of subsequent processing. In the dynamic partition optimization, by constructing a data access frequency model, combining with the K-means clustering algorithm and a real-time dynamic adjustment mechanism, it significantly improves the locality of data access and the rationality of storage allocation, reduces the computing overhead and communication latency caused by cross-partition access. In the event filtering stage, by introducing a Bloom filter and a rule table hash function matching mechanism, it realizes the rapid screening of potential events, avoids repeated operations on irrelevant data, further narrows the data processing scope, and improves the accuracy and efficiency of data filtering. In task scheduling, through the calculation intensity analysis of task features and the dynamic regulation of the load balancing model, it realizes the distributed scheduling of tasks and the efficient utilization of device computing power, lightweight tasks and intensive tasks are reasonably allocated, reducing resource conflicts and computing power waste. The overall processing logic forms a multi-level resource-saving mechanism from data input to task output through innovative methods of data reduction, partition optimization, event screening, and task scheduling, achieving a remarkable effect of reducing computing power occupancy, and providing a more efficient solution for high-performance computing and real-time computing scenarios.

[0053] 2. By combining the sparsification of high-dimensional datasets with principal component analysis, the present invention eliminates redundant features, extracts principal component features with high weights, reduces the data dimension, improves the utilization rate of computing resources, and retains the core information, making subsequent data processing more efficient and accurate. By combining the data access frequency model with dynamic partition adjustment, through real-time monitoring and clustering division of access characteristics, the storage distribution and access path of data are optimized, significantly reducing the communication overhead of cross-node access and waste of storage resources. In the screening of event data, by performing fast hash calculation and rule matching on the feature values of data blocks through a Bloom filter, potential key events are screened out, avoiding resource occupation caused by full traversal, and further refining the data extraction range through event priority classification. In task scheduling, by classifying the computational intensity of tasks and dynamically analyzing load balancing, light tasks and intensive tasks are respectively assigned to different device queues, significantly improving the allocation efficiency of computing power resources and device utilization rate, and avoiding resource conflicts and task delays. The overall processing logic constructs a highly targeted and efficient optimization mechanism in multiple links of the data processing chain, enabling the system to achieve significant savings in computing power and performance improvement in large-scale data processing and real-time computing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 Schematic diagram of the main steps of the present invention;

[0055] Figure 2 Schematic diagram of the refinement of S1 of the present invention;

[0056] Figure 3 Schematic diagram of the refinement of S2 of the present invention;

[0057] Figure 4 Schematic diagram of the refinement of S3 of the present invention;

[0058] Figure 5 Schematic diagram of the refinement of S4 of the present invention;

[0059] Figure 6 System block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0061] Please refer to the attached Figure 1 -attached Figure 6, an embodiment of the present invention provides a data management method for reducing computing power occupancy, including the following steps:

[0062] S1: Based on data distribution analysis, using a sparsification representation algorithm and a principal component analysis method, the redundancy of high-dimensional data features is calculated and removed by setting L1 regularization constraint conditions, and the data correlation matrix is decomposed by combining matrix operation methods, and the feature vectors with the top-ranked contribution values are extracted and converted into a sparse low-dimensional matrix; a sparsified low-dimensional data matrix is generated;

[0063] Based on data distribution analysis, combined with the matrix construction of high-dimensional data features, the feature sparsity of the data matrix is calculated column by column through the L1 regularization constraint algorithm, the feature terms with weight values close to zero are removed, and at the same time, a covariance matrix is constructed to quantify the correlation between features. The feature vectors are extracted by combining matrix decomposition and the ranking of their overall contribution values to the data is calculated. The data is transformed into the principal component space through matrix projection and the feature vectors with larger contribution rates are screened. The sparse matrix is optimized to retain the main features and the low-weight features are removed. Finally, a sparsified low-dimensional data matrix is generated through the optimized matrix.

[0064] S2: Based on the sparsified low-dimensional data matrix, using a dynamic partitioning optimization algorithm, a data access frequency distribution model is constructed to screen the datasets with high access frequencies, and the data is allocated to different storage partitions according to the weights in combination with the distributed node task table. The partition weights are dynamically adjusted based on real-time monitoring and the allocation results are updated; an optimized dynamic data partition table is generated;

[0065] Based on the sparsified low-dimensional data matrix, a data access frequency model is constructed by analyzing the data access log records. The access times and time intervals are counted in units of data blocks and the access frequency distribution is calculated. The access frequencies are partitioned by combining the K-means clustering method. The high-access-frequency data is allocated to the priority partition and an initial partition table is generated. At the same time, the data access dynamics are monitored in real time to adjust the access frequency data, and the adjusted partition status is stored in the distributed node. The cross-partition access communication cost is reduced through partition structure optimization and written into the optimized dynamic data partition table.

[0066] S3: Based on the optimized dynamic data partition table, using an event-driven filtering algorithm, a Bloom filter is constructed to screen the key event datasets that meet the conditions, and the data labels are judged item by item according to the defined rule table. After marking the key event identifiers for the data that hits the conditions, it is extracted and stored; a key event data stream is generated;

[0067] In the above content, the characteristic values of the data blocks are hashed and potential events are screened by constructing a Bloom filter, and the probability of potential events is calculated according to the formula .

[0068] In the formula, represents the potential event probability of the th data block, represents the weighted sum of the eigenvalue of the th data block, represents the mean of the weighted sum of the eigenvalues of all data blocks, is the total number of data blocks, represents the sum of the squared variances of the weighted sum of the eigenvalues and the mean, is the th regular hash value of the data block.

[0069] S4: Based on the critical event data stream, adopt a heterogeneous computing task scheduling algorithm. By classifying and marking the computational intensity of the task feature vectors, allocate lightweight tasks to the CPU parallel queue for processing and high-intensity tasks to the GPU batch queue for processing according to the task load analysis results, and adjust the distribution of computing tasks on each device in combination with the dynamic load model; generate an efficiently processed computational result set.

[0070] Based on the critical event data stream, extract the task feature vectors and calculate the intensity and data dependency relationship of each task, classify and mark lightweight tasks and intensive tasks. Combine the current computing power of the device and the task load status, dynamically generate a task and device allocation table through the load balancing model. At the same time, call the parallel scheduling algorithm to add lightweight tasks to the CPU task queue and intensive tasks to the GPU task queue for batch processing, and complete the tasks item by item according to the device execution feedback and generate the final computational result.

[0071] Based on the data distribution analysis, adopt a sparsification representation algorithm and a principal component analysis method. By setting the L1 regularization constraint condition, perform operations to eliminate the redundancy of high-dimensional data features, decompose the data correlation matrix in combination with the matrix operation method, extract the feature vectors with the top-ranked contribution values and convert them into a sparse low-dimensional matrix; the specific steps for generating the sparsified low-dimensional data matrix are as follows:

[0072] S101: Based on the high-dimensional data set, adopt the L1 regularization constraint algorithm. By constructing an objective function and taking the sparsity of the eigenvalues as the constraint condition, calculate each feature in the data matrix column by column, eliminate the feature items with weight values tending to zero under the constraint condition, and generate a preliminary sparse feature matrix;

[0073] Based on the sparsified low-dimensional data matrix, by analyzing the data access log records, count the access times and access time intervals of each data block, calculate the data access frequency by the time series analysis method, in combination with the access frequency formula ;

[0074] where, is the access frequency of the data block , is the access count of the data block, is the time interval. Generate a data access frequency model by sorting the access frequencies, and generate a data access frequency model.

[0075] S102: Based on the preliminary sparse feature matrix, adopt the covariance matrix calculation method, construct a feature correlation matrix according to the covariance values between features, obtain the eigenvector matrix through eigenvalue decomposition, extract the eigenvectors with the top-ranked weights as the principal components, and generate an eigenvector matrix;

[0076] S103: Based on the eigenvector matrix, through the matrix projection method, project the high-dimensional data into the principal component space, reduce the feature dimensions according to the principal component contribution rate, and eliminate the feature dimensions with the principal component contribution rate less than the preset threshold to generate a principal component low-dimensional matrix;

[0077] S104: Based on the principal component low-dimensional matrix, adopt the sparse matrix mapping method, optimize the matrix sparsity, retain the sparse features with the top 30% of the principal component weights, output the final sparse matrix result, and generate a sparse low-dimensional data matrix.

[0078] In S101, eliminating the feature terms with weight values tending to zero optimizes the sparsity, and the generated preliminary sparse feature matrix provides a basis for subsequent feature correlation calculations;

[0079] In S102, call the preliminary sparse feature matrix, analyze the relationship between features through covariance matrix, construct an eigenvector matrix, and provide support for high-dimensional data projection;

[0080] In S103, based on the eigenvector matrix, screen and reduce the invalid features through the principal component contribution rate, and the generated principal component low-dimensional matrix provides a data basis for sparse optimization;

[0081] In S104, call the principal component low-dimensional matrix, generate the final sparse low-dimensional data matrix through sparse optimization, and form a complete feature optimization chain.

[0082] Based on the sparse low-dimensional data matrix, adopt the dynamic partition optimization algorithm, screen the high-frequency accessed data set by constructing a data access frequency distribution model, allocate the data to different storage partitions according to the weights in combination with the distributed node task table, and dynamically adjust the partition weights based on real-time monitoring and update the allocation results; The specific steps to generate the optimized dynamic data partition table are as follows:

[0083] S201: Based on the sparse low-dimensional data matrix, construct a data access frequency model, count the access times and time intervals of each data block according to the data access log records, calculate the access frequency and generate the data block access frequency ranking, and generate a data access frequency model;

[0084] S202: Based on the data access frequency model, use the K-means clustering algorithm to allocate the data blocks with access frequencies in the top 30% to the priority partition, and allocate the data blocks with access frequencies in the bottom 30% to the secondary partition. Mark and record the initial state of the partition to generate the initial data partition table;

[0085] S203: Based on the initial data partition table, capture the dynamic access situation of the data blocks during the operation through the real-time access monitoring method, adjust the partition state according to the latest access frequency, update the data partition distribution, and generate the partition table after dynamic adjustment;

[0086] S204: Based on the partition table after dynamic adjustment, write the adjusted data partition result into the distributed storage node through the distributed node storage management method to complete the optimized storage of the partition and generate the optimized dynamic data partition table.

[0087] In S201, calculate the access frequency model based on the sparse low-dimensional data matrix to provide reference data for subsequent partition division;

[0088] In S202, call the data access frequency model and generate the initial data partition table through K-means clustering to provide the initial partition basis for dynamic adjustment;

[0089] In S203, call the initial data partition table and update the partition state in combination with the real-time monitoring data to generate the partition table after dynamic adjustment;

[0090] In S204, call the partition table after dynamic adjustment to write the partition result into the distributed node storage to form a complete partition optimization process and generate the optimized dynamic data partition table of the final result.

[0091] Based on the optimized dynamic data partition table, adopt the event-driven filtering algorithm. By constructing a Bloom filter to screen the key event data sets that meet the conditions, judge each item of the data label according to the defined rule table, and extract and store the data marked with the key event identifier for the data that hits the conditions; The specific steps to generate the key event data stream are as follows:

[0092] S301: Based on the optimized dynamic data partition table, construct a Bloom filter, define the rule table of key events, convert the rule conditions into the hash function of the Bloom filter, calculate the potential event characteristics of the data block, and use the formula: ; Mark the data block as a potential event by calculating item by item; Calculate the potential event probability , and generate the potential event data after Bloom filtering;

[0093] Among them, represents the potential event probability of the th data block, represents the The weighted sum of the key feature values in a data block represents the mean value of the weighted sums of the key feature values of all data blocks represents the total number of data blocks represents the sum of the squares of the differences between the weighted sum of the feature values of all data blocks and the mean value represents the rule hash value of the nth data block

[0094] S302: Based on the potential event data after Bloom filtering, through the key event matching method, retrieve the event feature values in the data blocks one by one, determine whether they meet the defined conditions in the rule table, filter the data blocks that meet the conditions, and generate key event data blocks

[0095] S303: Based on the key event data blocks, adopt the event priority annotation method, classify and label the key event data blocks, generate event priority labels according to the importance of the events, and generate key event data with marked priorities

[0096] S304: Based on the key event data with marked priorities, through the data extraction and storage algorithm, classify and store the data to the specified storage nodes according to the priorities, complete the extraction and storage of the key event data, and generate key event data

[0097] In S301, based on the optimized dynamic data partition table, construct a Bloom filter and calculate the potential event probability. The potential event data after Bloom filtering generated provides an initial screening set for subsequent screening

[0098] In S302, call the potential event data after Bloom filtering, match and screen item by item in combination with the rule table conditions. The generated key event data blocks provide a screening basis for subsequent classification and annotation

[0099] In S303, based on the key event data blocks, generate key event data with marked priorities through priority classification and annotation, providing a grading basis for subsequent storage optimization

[0100] In S304, call the key event data with marked priorities, classify and store the data to the specified nodes through the storage algorithm, and finally form a key event data stream, completing the processing of the entire process of event data from screening to storage

[0101] Based on the key event data stream, adopt the heterogeneous computing task scheduling algorithm. Through the calculation intensity classification and marking of the task feature vectors, allocate the light tasks to the CPU parallel queue for processing according to the task load analysis results, allocate the high-intensity tasks to the GPU batch queue for processing, and adjust the distribution of the computing tasks of each device in combination with the dynamic load model; the specific steps to generate an efficiently processed calculation result set are as follows

[0102] S401: Classify the task feature vectors based on the critical event data stream. By calculating the computational complexity and data dependency of the tasks, mark the tasks as lightweight tasks or intensive computing tasks, and generate the classified task feature vectors with classification tags.

[0103] S402: Based on the classified task feature vectors with classification tags, calculate the current load and available computing power of each computing device through the load balancing analysis model, and allocate the tasks to the appropriate computing devices according to the task features to generate a task-device allocation table.

[0104] S403: Based on the task-device allocation table, through the parallel task scheduling algorithm, allocate the lightweight tasks to the CPU parallel queue and the intensive computing tasks to the GPU batch computing queue to complete the task scheduling and generate the allocated computing task queue.

[0105] S404: Based on the allocated computing task queue, through the device execution and feedback mechanism, calculate each task item by item to output the processing result, and generate the final task processing result set to generate an efficiently processed computing result set.

[0106] In S401, parse the task feature vectors based on the critical event data stream and calculate the density to generate the classified task feature vectors with classification tags, which provide a classification basis for task allocation.

[0107] In S402, call the classified task feature vectors and combine the load balancing analysis to generate the task-device allocation table, which provides a device matching basis for task scheduling.

[0108] In S403, based on the task-device allocation table, use the parallel task scheduling algorithm to allocate tasks to different device queues, and the generated allocated computing task queue provides a clear queue structure for task execution.

[0109] In S404, call the allocated computing task queue to complete task execution and integrate all computing results, and finally generate an efficiently processed computing result set.

[0110] A data management system for reducing computing power occupancy, characterized by including the following modules: a feature extraction module, a dynamic partitioning module, an event filtering module, and a task scheduling module.

[0111] The feature extraction module, based on the high-dimensional data set, uses the L1 regularization constraint algorithm to calculate the feature sparsity column by column to eliminate the feature items with weights tending to zero, combines the covariance matrix to calculate the correlation between features to extract the feature vectors, uses the matrix projection method to transform the high-dimensional data into the principal component space and eliminates the feature dimensions with contribution rates less than the preset value, and outputs the main features after optimizing the matrix sparsity to generate a sparsified low-dimensional data matrix.

[0112] In the feature extraction module, the sparsification processing submodule performs weight elimination to generate a preliminary sparse feature matrix, the eigenvector extraction submodule calculates feature correlation and generates a eigenvector matrix, and the principal component conversion submodule uses the matrix projection method to screen features with higher contribution rates to generate a sparse low-dimensional data matrix.

[0113] The feature extraction module includes a sparse processing submodule, a feature vector extraction submodule, and a principal component conversion submodule;

[0114] The dynamic partitioning module is based on the sparse low-dimensional data matrix. It builds a data access frequency model by analyzing the access logs, uses the K-means clustering algorithm to allocate high-frequency access data to priority partitions and generate a partition table. It combines real-time monitoring to capture dynamic data access and adjust the partition status. It writes the adjusted partitions to the storage nodes through the distributed node management method to generate an optimized dynamic data partition table.

[0115] In the dynamic partitioning module, the frequency modeling submodule analyzes the access log to generate a data access frequency model, the partition clustering submodule performs clustering calculations based on the access frequency and generates an initial data partition table, and the partition dynamic adjustment submodule monitors access in real time, dynamically optimizes the partition structure, and generates an optimized dynamic data partition table.

[0116] The dynamic partitioning module includes a frequency modeling submodule, a partition clustering submodule, and a partition dynamic adjustment submodule;

[0117] The event filtering module builds a Bloom filter based on the optimized dynamic data partition table to filter potential events. It marks potential events by calculating the data feature values ​​and matching them with the hash function of the event rules one by one. It uses the key event matching method to filter the data blocks that meet the conditions and mark the key events with priority tags. Finally, it uses the extraction and storage algorithm to store the event data to the target node according to the priority, generating a key event data stream.

[0118] In the event filtering module, the Bloom filtering submodule filters potential event data blocks to generate potential event data after Bloom filtering, the event feature matching submodule filters data blocks that meet the conditions of the rule table to generate key event data blocks, and the priority labeling submodule classifies and labels the event data priorities to generate key event data labeled with priorities.

[0119] The event filtering module includes a Bloom filtering submodule, an event feature matching submodule, and a priority marking submodule;

[0120] Task scheduling module; based on the critical event data stream, combined with the task feature vector to calculate the task density and data dependency for task classification, dynamically generate a task allocation table through the load balancing analysis model by calculating the available computing power of the device, and use the parallel task scheduling algorithm to allocate lightweight tasks to the CPU parallel queue and high-density tasks to the GPU batch queue, and output the final result after the task is completed to generate an efficiently processed calculation result set;

[0121] In the task scheduling module, the task feature classification sub-module generates a task feature vector with classification marks by calculating the task feature vector classification, the task load balancing sub-module generates a task device allocation table by combining the device load and computing power, and the parallel queue scheduling sub-module generates an allocated computing task queue through task scheduling.

[0122] The task scheduling module includes a task feature classification sub-module, a task load balancing sub-module, and a parallel queue scheduling sub-module.

[0123] Sparsification processing sub-module, based on the high-dimensional data set, using the L1 regularization constraint algorithm, by constructing an objective function, taking the sparsity of the eigenvalue as a constraint condition, calculating the feature weights in the data matrix column by column, removing the feature terms with weights tending to zero, and generating a preliminary sparse feature matrix;

[0124] Based on the high-dimensional data set, using the L1 regularization constraint algorithm, by constructing an objective function ;

[0125] Among them, is the feature weight vector, is the data matrix, is the target vector, is the regularization parameter, and the objective function is optimized and iterated by the gradient descent method, the feature weights in the data matrix are calculated and screened column by column, the features with weight values tending to zero are removed, and the features with significant importance are retained, so as to optimize the feature sparsity.

[0126] Feature vector extraction sub-module, based on the preliminary sparse feature matrix, using the covariance matrix calculation method, calculating the covariance values between features to construct a feature correlation matrix, extracting the feature vector matrix through matrix decomposition, and selecting the feature vectors with the top weight rankings to generate a feature vector matrix;

[0127] Based on the preliminary sparse feature matrix, using the covariance matrix calculation method, using the formula ;

[0128] Among them, is the covariance matrix, is each column of data in the matrix, is the column mean. A feature correlation matrix is constructed based on the covariance values between features, decomposed using the eigen - decomposition algorithm, and the eigen - vector matrix is extracted. The eigen - vectors with the top weights are selected by sorting the eigenvalues, and the principal - component eigen - vectors are retained for the next - step optimization.

[0129] The principal - component transformation sub - module, based on the eigen - vector matrix, projects the high - dimensional data into the principal - component space through matrix projection, screens the feature dimensions according to the contribution rate of the principal components, eliminates the feature dimensions with a contribution rate less than the preset threshold, optimizes the matrix sparsity, and then outputs the result to generate a sparse low - dimensional data matrix.

[0130] Based on the eigen - vector matrix, the high - dimensional data is mapped into the principal - component space through matrix projection, using the formula ;

[0131] where, is the projection result, is the initial matrix, is the eigen - vector matrix. According to the formula for calculating the contribution rate of the principal components the dimensions with a principal - component contribution rate lower than the preset threshold are eliminated to reduce redundant information, and the low - dimensional representation ability of the data is further improved through sparse optimization. After optimizing the matrix sparsity, the final sparse low - dimensional matrix is formed;

[0132] The preliminary sparse feature matrix generated by the sparsification processing sub - module provides the basic data for subsequent feature - correlation calculations, significantly reducing the interference of useless features;

[0133] The eigen - vector extraction sub - module constructs a covariance matrix using the preliminary sparse feature matrix and extracts eigen - vectors. The generated eigen - vector matrix serves as the basis for principal - component transformation;

[0134] The principal - component transformation sub - module calls the eigen - vector matrix for high - dimensional projection and sparse optimization, and finally outputs a sparse low - dimensional data matrix, forming a complete feature - processing chain.

[0135] The frequency - modeling sub - module, based on the sparse low - dimensional data matrix, analyzes the data - access logs. By counting the access times and time intervals of each data block, it calculates the access frequency and generates an access - frequency sorting model to generate a data - access frequency model;

[0136] The partition - clustering sub - module, based on the data - access frequency model, uses the K - means clustering algorithm. Through the clustering calculation of the data - block access frequencies, it assigns the data blocks with access frequencies in the top 30% to the priority partition and the data blocks with access frequencies in the bottom 30% to the secondary partition, generating an initial partition table to generate an initial data - partition table;

[0137] The partition dynamic adjustment sub-module, based on the initial data partition table, captures the dynamic access situation of data blocks through real-time monitoring, updates the access frequency data, adjusts the partition status according to the latest access frequency, optimizes the partition structure, and generates an optimized dynamic data partition table.

[0138] The frequency modeling sub-module provides a quantitative metric for the subsequent priority division of partitions through the data access frequency model generated by analyzing access logs;

[0139] The partition clustering sub-module calls the data access frequency model, and the initial data partition table generated by the K-means clustering algorithm provides an initial partition structure for dynamic adjustment;

[0140] The partition dynamic adjustment sub-module, based on the initial data partition table, combines real-time monitoring to update the access frequency and optimize the partition structure, and finally generates an optimized dynamic data partition table, realizing the optimization of dynamic partitions and the reasonable allocation of resources.

[0141] The Bloom filter sub-module, based on the optimized dynamic data partition table, constructs a Bloom filter. By defining a rule table for key events, calculates the eigenvalue hash function of each data block, marks the potential event data blocks that meet the rules as potential events, and generates the Bloom-filtered potential event data;

[0142] The event feature matching sub-module, based on the Bloom-filtered potential event data, uses the key event matching method to retrieve the event eigenvalue item by item and determines whether it meets the conditions of the rule table, filters out the event data blocks that meet the conditions, and generates the key event data blocks;

[0143] The priority annotation sub-module, based on the key event data blocks, classifies the event priorities of the data blocks, performs label annotation according to the importance of the events, and generates the key event data with priority annotation, generating the key event data with annotated priority.

[0144] The Bloom filter sub-module generates the Bloom-filtered potential event data for subsequent screening by performing hash calculation and rule matching on the eigenvalues of the data blocks in the optimized dynamic data partition table;

[0145] The event feature matching sub-module calls the Bloom-filtered potential event data and filters out the key event data blocks that meet the conditions through item-by-item comparison of the eigenvalues;

[0146] The priority annotation sub-module, based on the key event data blocks, classifies the data according to the importance of the events through calculation, generates the key event data with annotated priority, and realizes the hierarchical management and optimization of key events.

[0147] The task feature classification sub-module calculates the computational intensity classification of the task feature vector based on the critical event data with annotation priorities, marks it as a lightweight task or a dense task according to the computational complexity and data dependency relationship of the task, and generates a task feature vector with classification marks.

[0148] Based on the critical event data with annotation priorities, extract the feature vector of the task, and use the computational intensity analysis method. According to the formula ;

[0149] where is the density of the th task, is the computational complexity of the task, indicating the total amount of computational resources required to complete the task is the data dependency relationship of the task, indicating the number of accesses to other data blocks during the task execution. Classify the tasks according to the density results, mark the tasks with higher density as dense tasks, and mark the tasks with lower density as lightweight tasks.

[0150] The task load balancing sub-module calculates the current load and available computing power of each device based on the task feature vector with classification marks, combines the load balancing analysis model, generates an allocation table of tasks and devices, and generates a task-device allocation table.

[0151] Based on the task feature vector with classification marks, combine the load balancing analysis model, calculate the current load and available computing power of each device. The load calculation formula is as follows: ;

[0152] where is the load utilization rate of device , represents the total weight of the tasks already allocated to device , is the total computing power of the device. Allocate the tasks to the low-load devices with lightweight tasks first and to the high-computing-power devices with dense tasks first, and generate an allocation table of tasks and devices.

[0153] The parallel queue scheduling sub-module, based on the task-device allocation table, uses the parallel task scheduling algorithm to allocate lightweight tasks to the CPU parallel queue and dense tasks to the GPU batch queue, completes the task scheduling, and generates an allocated computational task queue.

[0154] Based on the task-device allocation table, use the parallel task scheduling algorithm to construct a task queue model, allocate lightweight tasks to the parallel task queue of the CPU, use the queue priority model to dynamically adjust the execution order of lightweight tasks, construct a batch task processing mechanism to allocate dense tasks to the batch computing queue of the GPU, and optimize the resource allocation through the scheduling strategy of maximizing the utilization of device computing power to complete the task scheduling.

[0155] The task feature classification sub-module provides a clear classification basis for the allocation and scheduling of tasks through the task feature vector of the classification label generated by computing intensity analysis;

[0156] The task load balancing sub-module calls the task feature vector of the classification label, combines the load and computing power of the device to allocate resources, and the generated task-device allocation table provides an allocation plan for the actual scheduling of tasks;

[0157] The parallel queue scheduling sub-module, based on the task-device allocation table, through the priority scheduling and queue management of tasks, finally generates an allocated computing task queue to complete the optimization of task scheduling and allocation.

[0158] In the present invention, unless otherwise clearly defined and limited, terms such as "installation", "connection", "connection", "fixation" and the like shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection, an electrical connection or communication with each other; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0159] Obviously, the above-described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. The drawings show the preferred embodiments of the present invention, but do not limit the patent scope of the present invention. The present invention can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present invention more thorough and comprehensive. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or equivalently replace some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present invention in other related technical fields is equally within the scope of the patent protection of the present invention.

Claims

1. A data management method for reducing computing power usage, characterized in that: The following steps are involved: S1: Based on data distribution analysis, the sparse representation algorithm and principal component analysis method are used to eliminate the redundancy of high-dimensional data features by setting L1 regularization constraints, decompose the data correlation matrix by combining matrix operations, extract the feature vectors with the top 30% contribution rate and convert them into sparse low-dimensional matrices; Generate a sparse low-dimensional data matrix; S2: Based on the sparse low-dimensional data matrix, a dynamic partition optimization algorithm is used to select high-frequency access data sets with more than 10 accesses per minute by building a data access frequency distribution model. The data is then allocated to different storage partitions according to weights in combination with the distributed node task table. The partition weights are dynamically adjusted based on real-time monitoring and the allocation results are updated. Generate optimized dynamic data partition table; S3: Based on the optimized dynamic data partition table, an event-driven filtering algorithm is used to construct a Bloom filter to filter key event data sets that meet the conditions. Data tags are judged item by item according to the defined rule table, and key event identifiers are marked for data that hits the conditions before extraction and storage. Generate key event data streams; S4: Based on the key event data stream, a heterogeneous computing task scheduling algorithm is used. By classifying and marking the computing density of task feature vectors, light tasks with computing complexity lower than a specific threshold are assigned to the CPU parallel queue for processing according to the task load analysis results, and high-density tasks with computing complexity higher than a specific threshold are assigned to the GPU batch queue for processing. The distribution of computing tasks on each device is adjusted in combination with the dynamic load model; a set of computing results for efficient processing is generated; Based on the optimized dynamic data partition table, an event-driven filtering algorithm is used to construct a Bloom filter to filter the key event data set that meets the conditions. The data tags are judged item by item according to the defined rule table. The key event identifiers are marked for the data that hits the conditions and then extracted and stored. The specific steps for generating the key event data stream are as follows: S301: Based on the optimized dynamic data partition table, a Bloom filter is constructed, a rule table of key events is defined, rule conditions are converted into a hash function of the Bloom filter, and potential event features of the data block are calculated using the formula: ; Marking data blocks as potential events by calculating them one by one; Calculate the probability of potential events , generate potential event data after Bloom filtering; in, Representative The potential event probability of each data block, Representative The weighted sum of the key eigenvalues ​​in the data blocks, Represents the mean of the weighted sum of the key feature values ​​of all data blocks, Represents the total number of data blocks, Represents the sum of squares of the weighted sum of the eigenvalues ​​of all data blocks and the mean value, Representative The regular hash value of a data block; S302: Based on the potential event data after Bloom filtering, the event feature values ​​in the data blocks are retrieved one by one through the key event matching method to determine whether the definition conditions of the rule table are met, and the data blocks that meet the conditions are screened to generate key event data blocks; S303: Based on the key event data block, an event priority labeling method is used to classify and label the key event data block, generate event priority labels according to event importance, and generate key event data with labeled priorities; S304: Based on the key event data with marked priorities, the data is classified and stored in the designated storage node according to the priorities through the data extraction and storage algorithm, the extraction and storage of the key event data is completed, and the key event data stream is generated.

2. A data management method for reducing computing power usage according to claim 1, characterized in that: Based on data distribution analysis, the sparse representation algorithm and principal component analysis method are used to eliminate the redundancy of high-dimensional data features by setting L1 regularization constraints, and the data correlation matrix is ​​decomposed by combining matrix operations to extract the feature vectors with the top 30% contribution rate and convert them into sparse low-dimensional matrices. The specific steps to generate a sparse low-dimensional data matrix are: S101: Based on the high-dimensional data set, the L1 regularization constraint algorithm is used to construct the objective function, take the sparsity of the eigenvalue as the constraint condition, calculate each feature in the data matrix column by column, remove the feature items with weight values ​​less than 0.01 under the constraint condition, and generate a preliminary sparse feature matrix; S102: Based on the preliminary sparse feature matrix, a covariance matrix calculation method is used to construct a feature correlation matrix according to the covariance values ​​between features, and a feature vector matrix is ​​obtained through feature decomposition. The feature vectors ranked in the top 30% of the weights are extracted as principal components to generate a feature vector matrix. S103: Based on the eigenvector matrix, the high-dimensional data is projected into the principal component space by a matrix projection method, the feature dimension is reduced according to the principal component contribution rate, the feature dimension with a principal component contribution rate less than 10% is eliminated, and a principal component low-dimensional matrix is ​​generated; S104: Based on the principal component low-dimensional matrix, a sparse matrix mapping method is used to optimize the matrix sparsity, retain the sparse features of the principal component weights in the top 30%, output the final sparse matrix result, and generate a sparse low-dimensional data matrix.

3. A data management method for reducing computing power usage according to claim 1, characterized in that: Based on the sparse low-dimensional data matrix, a dynamic partition optimization algorithm is used to construct a data access frequency distribution model to screen high-frequency access data sets with more than 10 accesses per minute. Combined with the distributed node task table, the data is allocated to different storage partitions according to weights. The partition weights are dynamically adjusted based on real-time monitoring and the allocation results are updated. The specific steps to generate the optimized dynamic data partition table are as follows: S201: constructing a data access frequency model based on a sparse low-dimensional data matrix, counting the number of accesses and time intervals of each data block according to data access log records, calculating the access frequency and generating a data block access frequency ranking, and generating a data access frequency model; S202: Based on the data access frequency model, a K-means clustering algorithm is used to allocate data blocks with access frequencies in the top 30% to priority partitions, and data blocks with access frequencies in the bottom 30% to secondary partitions, and the initial state of the partitions is marked and recorded to generate an initial data partition table; S203: Based on the initial data partition table, the dynamic access situation of the data blocks during operation is captured by a real-time access monitoring method, the partition state is adjusted according to the latest access frequency, the data partition distribution is updated, and a dynamically adjusted partition table is generated; S204: Based on the dynamically adjusted partition table, the adjusted data partition results are written into the distributed storage nodes through the distributed node storage management method, the partition optimization storage is completed, and the optimized dynamic data partition table is generated.

4. A data management method for reducing computing power usage according to claim 1, characterized in that: Based on the key event data stream, a heterogeneous computing task scheduling algorithm is adopted. By classifying and marking the computing density of task feature vectors, light tasks are assigned to CPU parallel queue processing according to the task load analysis results, and high-density tasks are assigned to GPU batch queue processing. The distribution of computing tasks on each device is adjusted in combination with the dynamic load model. The specific steps to generate a set of computing results for efficient processing are as follows: S401: Classifying task feature vectors based on key event data streams, marking tasks as lightweight tasks or intensive computing tasks by calculating the computational complexity and data dependency of the tasks, and generating task feature vectors marked by classification; S402: Based on the task feature vectors marked by classification, the current load and available computing power of each computing device are calculated through the load balancing analysis model, and the tasks are assigned to the corresponding computing devices in combination with the task features to generate a task device allocation table; S403: Based on the task device allocation table, the lightweight tasks are allocated to the CPU parallel queue and the intensive computing tasks are allocated to the GPU batch computing queue through the parallel task scheduling algorithm, the task scheduling is completed, and the allocated computing task queue is generated; S404: Based on the assigned computing task queue, through the device execution and feedback mechanism, the processing results are output after the calculations are completed item by item, and a final task processing result set is generated to generate an efficient processing computing result set.

5. A data management system for reducing computing power usage, characterized in that: Includes the following modules: feature extraction module, dynamic partitioning module, event filtering module, task scheduling module; The feature extraction module uses the L1 regularization constraint algorithm to calculate the feature sparsity column by column based on the high-dimensional data set, removes the feature items whose weights tend to zero, calculates the correlation between features in combination with the covariance matrix to extract the feature vector, and uses the matrix projection method to convert the high-dimensional data into the principal component space and remove the feature dimensions whose contribution rate is less than the preset value. After optimizing the matrix sparsity, the main features are output to generate a sparse low-dimensional data matrix; The feature extraction module includes a sparse processing submodule, a feature vector extraction submodule, and a principal component conversion submodule; The dynamic partitioning module is based on the sparse low-dimensional data matrix. It builds a data access frequency model by analyzing the access logs, uses the K-means clustering algorithm to allocate high-frequency access data to priority partitions and generate a partition table. It combines real-time monitoring to capture dynamic data access and adjust the partition status. It writes the adjusted partitions to the storage nodes through the distributed node management method to generate an optimized dynamic data partition table. The dynamic partitioning module includes a frequency modeling submodule, a partition clustering submodule, and a partition dynamic adjustment submodule; The event filtering module builds a Bloom filter based on the optimized dynamic data partition table to filter potential events. It marks potential events by calculating the data feature values ​​and matching them with the hash function of the event rules one by one. It uses the key event matching method to filter the data blocks that meet the conditions and mark the key events with priority tags. Finally, it uses the extraction and storage algorithm to store the event data to the target node according to the priority, generating a key event data stream. The event filtering module includes a Bloom filtering submodule, an event feature matching submodule, and a priority marking submodule; Task scheduling module; Based on the key event data stream, the task density and data dependency are calculated in combination with the task feature vector to classify the tasks. The available computing power of the computing device is calculated through the load balancing analysis model to dynamically generate a task allocation table. The parallel task scheduling algorithm is used to allocate lightweight tasks to the CPU parallel queue and high-density tasks to the GPU batch queue. After the task is completed, the final result is output to generate a set of calculation results for efficient processing. The task scheduling module includes a task feature classification submodule, a task load balancing submodule, and a parallel queue scheduling submodule; The Bloom filter submodule builds a Bloom filter based on the optimized dynamic data partition table. By defining the rule table of key events, it calculates the characteristic value hash function of each data block, marks the potential event data blocks that meet the rules as potential events, and generates potential event data after Bloom filtering; The event feature matching submodule uses the key event matching method based on the potential event data after Bloom filtering to retrieve the event feature values ​​one by one and determine whether they meet the conditions of the rule table, filter out the event data blocks that meet the conditions, and generate key event data blocks; The priority labeling submodule classifies the event priorities of the data blocks based on the key event data blocks, labels the events according to their importance, and generates key event data with priority labels.

6. A data management system for reducing computing power usage according to claim 5, characterized in that: The sparse processing submodule is based on high-dimensional data sets and adopts the L1 regularization constraint algorithm. By constructing the objective function, the sparsity of the eigenvalue is used as a constraint condition, the feature weights in the data matrix are calculated column by column, and the feature items with weights approaching zero are eliminated to generate a preliminary sparse feature matrix. The feature vector extraction submodule uses the covariance matrix calculation method based on the preliminary sparse feature matrix to calculate the covariance values ​​between features to construct the feature correlation matrix, extracts the feature vector matrix through matrix decomposition, and selects the feature vectors with the highest weight ranking to generate the feature vector matrix; The principal component conversion submodule, based on the eigenvector matrix, projects the high-dimensional data into the principal component space through the matrix projection method, selects the feature dimensions according to the principal component contribution rate, removes the feature dimensions whose contribution rate is less than the preset threshold, optimizes the matrix sparsity and then outputs the result to generate a sparse low-dimensional data matrix.

7. A data management system for reducing computing power usage according to claim 5, characterized in that: The frequency modeling submodule analyzes data access logs based on the sparse low-dimensional data matrix, calculates the access frequency and generates an access frequency ranking model by counting the number of accesses and time intervals of each data block, and generates a data access frequency model; The partition clustering submodule uses the K-means clustering algorithm based on the data access frequency model to cluster the access frequencies of data blocks, allocate the data blocks with the top 30% access frequencies to the priority partitions, and allocate the data blocks with the bottom 30% access frequencies to the secondary partitions, and generate the initial partition table; The partition dynamic adjustment submodule, based on the initial data partition table, captures the dynamic access status of data blocks through real-time monitoring, updates the access frequency data, adjusts the partition status and optimizes the partition structure according to the latest access frequency, and generates an optimized dynamic data partition table.

8. A data management system for reducing computing power usage according to claim 5, characterized in that: The task feature classification submodule classifies the computational density of task feature vectors based on the key event data with marked priorities, marks them as lightweight tasks or intensive tasks according to their computational complexity and data dependencies, and generates task feature vectors with classification labels. The task load balancing submodule calculates the current load and available computing power of each device based on the task feature vectors marked by classification and combined with the load balancing analysis model, generates a task and device allocation table, and generates a task and device allocation table; The parallel queue scheduling submodule uses a parallel task scheduling algorithm based on the task device allocation table to assign lightweight tasks to the CPU parallel queue and intensive tasks to the GPU batch queue to complete task scheduling and generate an assigned computing task queue.

Citation Information

Patent Citations

  • Big data task scheduling method

    CN112256418A

  • Data storage method and system

    CN117827850A

  • Computing power resource processing method

    CN118069380A