Bucket expansion regulation and control method and device based on integrated decision tree, and electronic equipment
By integrating decision trees into the bucket expansion control method, the data distribution of the financial data warehouse is monitored and dynamically adjusted in real time, which solves the imbalance problem caused by changes in data volume, optimizes resource utilization efficiency and query performance, and improves the stability and high availability of the system.
Patent Information
- Application Number
- CN202511133381.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-18
AI Technical Summary
The indexing strategies of financial data warehouses cannot adapt to changes in data volume in real time, resulting in uneven data distribution, low system resource utilization efficiency, and difficulty in achieving dynamic adaptive adjustment with existing technologies.
An expansion and control method based on ensemble decision trees is adopted. By collecting data bucket performance data and related financial system operation performance data in real time through a monitoring agent, decision tree regression algorithm is used to generate decision results, dynamically adjust bucket indexing strategy, and optimize data redistribution.
It has achieved system performance improvement and resource utilization efficiency optimization for financial data warehouses under the scenario of sudden increase in data volume, ensured data storage load balance, and improved query performance and system high availability.
Smart Images

Figure CN120973792A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of big data and distributed computing technology or other related fields. Specifically, it relates to a bucket expansion control method and device, and electronic equipment based on integrated decision trees. Background Technology
[0002] In financial data warehouse applications, the dynamic changes in data volume pose a significant challenge to storage and query performance. Traditional data warehouse indexing strategies typically employ a static bucketing mechanism, where the number of buckets and distribution rules are pre-set during system initialization. Subsequent updates are limited to manual intervention or fixed-period adjustments. This static management approach often leads to severely uneven data distribution when facing high-frequency transactions, sudden data surges, or cyclical business fluctuations in the financial sector. Specifically, hot data tends to accumulate in a few data buckets, while cold data is scattered and consumes redundant storage resources.
[0003] Due to the lack of dynamic perception of real-time data characteristics, existing technologies are unable to automatically adapt to changes in data scale and access patterns. This not only leads to a decrease in storage space utilization but also causes a deterioration in query performance. In particular, in complex analysis scenarios, it can easily lead to competition for computing node resources, resulting in a decrease in the overall system throughput.
[0004] In related technologies, solutions often rely on preset rules or offline analysis to adjust indexing strategies, which suffer from significant lag and coarse-grained defects. For example, redistributing data in batches via scheduled tasks cannot respond to sudden load changes during peak business periods, while prediction models based on historical data are difficult to adapt to the dynamic characteristics of real-time data streams. During data migration, service shutdowns or degraded operations are often required, severely impacting the continuity of financial services.
[0005] Furthermore, existing technologies typically separate data distribution optimization from system resource regulation, failing to establish a collaborative decision-making mechanism based on multi-dimensional performance indicators. This leads to a disconnect between bucket expansion operations and the underlying resource status, further exacerbating the problem of low resource utilization efficiency.
[0006] There is currently no effective solution to the above problems. Summary of the Invention
[0007] The main objective of this application is to provide a bucket expansion control method, device, and electronic device based on integrated decision trees, so as to at least solve the technical problem in the related technology that the indexing strategy of financial data warehouse cannot adapt to changes in data volume in real time, resulting in uneven data distribution and low system resource utilization efficiency.
[0008] To achieve the above objectives, according to one aspect of this application, a bucket expansion control method based on ensemble decision trees is provided. The method includes: collecting data bucket performance data of a financial data warehouse and operational performance data of the associated financial system driving the operation of the financial data warehouse through a pre-deployed real-time monitoring agent to obtain a decision basis; generating a decision result using a decision tree regression algorithm and the decision basis, wherein the decision tree regression algorithm pre-configures N decision trees constructed through random feature selection based on an ensemble framework, the N decision trees generating N-dimensional decision suggestions based on the decision basis and performing average aggregation to obtain the decision result, where N is a first preset value; determining a target bucket expansion strategy from M bucket expansion strategies with pre-determined priorities based on the decision result, where M is a second preset value; and performing a data redistribution operation on the financial data warehouse based on the target bucket expansion strategy.
[0009] To achieve the above objectives, according to another aspect of this application, a bucket expansion control device based on an integrated decision tree is also provided. This device includes: a collection unit, configured to collect data bucket performance data of a financial data warehouse and operational performance data of the associated financial system driving the operation of the financial data warehouse through a pre-deployed real-time monitoring agent, to obtain a decision basis; a generation unit, configured to generate a decision result using a decision tree regression algorithm and the decision basis, wherein the decision tree regression algorithm pre-configures N decision trees constructed through random feature selection based on an integrated framework, and the N decision trees generate N-dimensional decision suggestions based on the decision basis and perform average aggregation to obtain the decision result, where N is a first preset value; a determination unit, configured to determine a target bucket expansion strategy from M bucket expansion strategies with pre-determined priorities based on the decision result, where M is a second preset value; and an execution unit, configured to perform a data redistribution operation on the financial data warehouse based on the target bucket expansion strategy.
[0010] To achieve the above objectives, according to another aspect of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the bucket expansion control method based on the integrated decision tree described above.
[0011] To achieve the above objectives, according to another aspect of this application, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the bucket expansion control method based on the integrated decision tree described above.
[0012] To achieve the above objectives, according to another aspect of this application, a computer program product is also provided, including computer instructions, wherein when the computer instructions are executed by a processor, they implement the steps of the bucket expansion control method based on ensemble decision tree described in any one of the above claims.
[0013] This invention proposes a bucket expansion control method based on ensemble decision trees. First, a pre-deployed real-time monitoring agent collects data bucket performance data from a financial data warehouse, as well as performance data from related financial systems driving the data warehouse, to obtain decision-making criteria. Then, a decision result is generated using a decision tree regression algorithm and the decision criteria. Specifically, the decision tree regression algorithm pre-configures N decision trees constructed through random feature selection based on an ensemble framework. These N decision trees generate N-dimensional decision suggestions based on the decision criteria and are then averaged and aggregated to obtain the decision result, where N is a first preset value. Next, based on the decision result, a target bucket expansion strategy is determined from M bucket expansion strategies with pre-determined priorities, where M is a second preset value. Finally, a data redistribution operation is performed on the financial data warehouse based on the target bucket expansion strategy.
[0014] This invention employs an integrated decision tree machine learning approach, combining a real-time monitoring agent with a decision tree regression algorithm to achieve intelligent dynamic adjustment of the bucket index in a financial data warehouse. This results in a significant improvement in system performance and a high degree of resource utilization efficiency under scenarios of sudden data surges. Specifically, the real-time monitoring agent collects performance data of data buckets in the financial data warehouse, as well as the operational performance indicators of the supporting financial system. Then, using multiple decision trees within the integrated framework, each tree independently generates decision suggestions through a random feature selection strategy. Finally, a weighted average is used to aggregate suggestions from multiple dimensions to obtain a comprehensive and accurate bucket index augmentation decision result, fully considering the current and expected system load conditions to ensure... The scientific validity and effectiveness of the bucket expansion strategy are demonstrated by selecting the most suitable target strategy from multiple bucket expansion strategies with pre-set priorities based on the decision results. This balances the flexibility of automated decision-making with the necessity of manual intervention. Finally, a refined data redistribution operation is performed on the financial data warehouse according to the selected target bucket expansion strategy. This not only balances the data storage load but also optimizes query performance, ensuring the high availability and efficiency of the data warehouse. It overcomes the inherent static defects of traditional financial data warehouse indexing strategies, realizes dynamic adaptive adjustment of bucket indexes, and thus solves the technical problem in related technologies where financial data warehouse indexing strategies cannot adapt to changes in data volume in real time, resulting in uneven data distribution and low system resource utilization efficiency. Attached Figure Description
[0015] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0016] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a bucket control method based on an integrated decision tree is shown.
[0017] Figure 2 This is a flowchart of an optional bucket expansion control method based on an integrated decision tree according to an embodiment of the present invention;
[0018] Figure 3 This is a schematic diagram of an optional data table bucket index management system based on a random forest regression model in a real-time financial data warehouse scenario, according to an embodiment of the present invention.
[0019] Figure 4 This is a schematic diagram of an optional bucket expansion control device based on an integrated decision tree according to an embodiment of the present invention;
[0020] Figure 5 This is a structural block diagram of an electronic device that executes an expansion bucket control method based on an integrated decision tree according to an embodiment of the present invention. Detailed Implementation
[0021] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0023] To facilitate understanding of the present invention by those skilled in the art, some terms or nouns involved in the various embodiments of the present invention are explained below:
[0024] Ensemble Decision Trees, a method in ensemble learning, improves model accuracy and stability by combining the predictions of multiple decision trees. This invention utilizes an ensemble of decision tree regression algorithms, specifically Random Forest. Multiple decision trees are built, and a subset of features is randomly selected at each tree node split. Finally, the predictions from all decision trees are averaged to make a decision. This approach effectively captures non-linear relationships between data when dealing with complex, high-dimensional data, improving prediction accuracy while avoiding overfitting and enhancing the model's generalization ability.
[0025] Real-time Monitoring Agent, as described in this invention, is a software component responsible for collecting and transmitting system performance data in real time. Deployed on key nodes of the financial data warehouse and its associated financial systems, it can continuously monitor and record key indicators such as data bucket write rate, storage capacity, system CPU utilization, and memory usage, and transmit this data to the decision tree regression algorithm module for analysis.
[0026] It should be noted that the bucket expansion control method and apparatus based on ensemble decision tree in this application can be used in the fields of big data and distributed computing technology for real-time performance monitoring and intelligent resource allocation of financial real-time data warehouses, and can also be used in any field other than big data and distributed computing technology for real-time performance monitoring and intelligent resource allocation of financial real-time data warehouses. This application does not limit the application field of the bucket expansion control method and apparatus based on ensemble decision tree.
[0027] It should be noted that all relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in this application are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, processing, transmission, provision, disclosure, use, and handling of such data comply with the laws, regulations, and standards of the relevant regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding access points are provided for users to choose to authorize or refuse access. For example, this system has interfaces with relevant users or organizations. Before obtaining relevant information, a request to obtain the information needs to be sent to the aforementioned user or organization through the interface, and the relevant information is obtained only after receiving consent from the aforementioned user or organization.
[0028] The information collection (e.g., user voice, video, and text collection) and analysis operations involved in this application have provided users with corresponding operation entry points during execution, allowing users to choose to agree to or reject the automated decision results; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0029] The following embodiments of the present invention can be applied to various systems / applications / devices that require real-time data processing and dynamic resource allocation, enabling intelligent bucket expansion control and resource optimization based on ensemble decision trees. The present invention uses an ensemble decision tree regression algorithm to perform in-depth analysis and prediction of the performance data of a financial data warehouse, and then intelligently selects a bucket expansion strategy and executes data redistribution operations based on the prediction results. This better addresses dynamic changes in data volume while optimizing resource utilization, improving data processing efficiency and system stability.
[0030] This invention also uses a real-time monitoring and feedback mechanism to precisely adjust the number of data buckets, efficiently manage data flow, and ensure that the data warehouse maintains high performance and low latency in high-concurrency scenarios, providing real-time and accurate data support for financial decision-making.
[0031] The present invention will now be described in detail with reference to various embodiments.
[0032] Example 1
[0033] According to an embodiment of the present invention, an embodiment of a bucket expansion control method based on an integrated decision tree is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0034] The bucket expansion control method based on integrated decision tree provided in Embodiment 1 of the present invention can be executed on a mobile terminal, computer terminal or similar computing device. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing a bucket control method based on an ensemble decision tree is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0035] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the bucket expansion control method based on integrated decision tree in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned bucket expansion control method based on integrated decision tree. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0037] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0038] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0039] Under the above operating environment, the present invention provides, as follows: Figure 2 The bucket expansion control method shown is based on ensemble decision trees. The main implementation component is an intelligent bucket expansion control system. Combining ensemble decision trees and random forest regression techniques, this method is applied to real-time financial data warehouse scenarios, particularly addressing performance bottlenecks and uneven resource allocation caused by sudden data surges. Through real-time monitoring and intelligent decision-making, it aims to intelligently and dynamically optimize the performance and resource utilization efficiency of the financial data warehouse, improving business continuity and the ability to cope with short-term data surges. Specifically, it includes the following steps: real-time collection of various performance indicators and system load data of the data warehouse; prediction of future data growth trends and potential performance requirements using machine learning models; dynamic adjustment of bucket indexing strategies based on prediction results, including increasing the number of buckets and redistributing data; implementation of gradual data migration to minimize system performance fluctuations and maintain business continuity; and real-time evaluation of performance improvements and resource utilization after the bucket expansion strategy is implemented, with feedback used for model optimization.
[0040] The embodiments of the present invention will now be described in detail with reference to each specific step.
[0041] Figure 2 This is a flowchart of an optional bucket expansion control method based on an ensemble decision tree according to an embodiment of the present invention, such as... Figure 2 As shown, the method includes the following steps:
[0042] Step S201: Through a pre-deployed real-time monitoring agent, collect the performance data of the data buckets in the financial data warehouse and the operational performance data of the related financial systems that drive the operation of the financial data warehouse to obtain the basis for decision-making.
[0043] Specifically, the real-time monitoring agent is a series of lightweight software components that are pre-deployed on various computing nodes of the financial data warehouse and related financial systems. It is responsible for collecting and transmitting performance indicator data in real time, and specifically performs actions such as data acquisition, data preprocessing, and data transmission.
[0044] Data acquisition refers to monitoring and acquiring key indicators of the financial data warehouse, such as data bucket write rate, query response time, and bucket data volume, according to the collection cycle, as well as the operating status of related financial systems (such as transaction processing systems, market data sources, customer information systems, etc.), including CPU utilization, memory usage, and disk I / O load.
[0045] Data preprocessing refers to cleaning and preprocessing the collected data before sending it to the core analysis components to remove noise and outliers; data transmission refers to the secure and efficient transmission of the processed data to the central monitoring or analysis platform.
[0046] A financial data warehouse is a data storage and analysis system used to integrate massive amounts of transaction, market, and customer data from multiple sources. It supports real-time business analysis and decision-making. Compared to traditional data warehouses, financial data warehouses place greater emphasis on the real-time nature, consistency, and security of data, as well as the ability to perform high-concurrency queries and writes.
[0047] Data bucket performance data refers to the real-time performance metrics of data buckets (logical units used for data distribution and partitioning) in a financial data warehouse. These mainly include, but are not limited to: write rate (reflecting the data ingestion speed of each data bucket, typically measured in entries per second), query latency (measuring the time required for a data bucket to respond to a query request, measured in milliseconds), data capacity (the amount of data currently stored in each data bucket, measured in bytes or more commonly GB, TB, etc.), and the number and size of files (the number and average size of columnar files within the data bucket).
[0048] The related financial systems encompass all external systems upon which the financial data warehouse depends, directly or indirectly affecting the data warehouse's performance and data quality. These include, but are not limited to: transaction processing systems, which handle user transaction requests, and whose performance directly impacts the data warehouse's write load and data update frequency; market data sources, which provide real-time market data, and whose data accuracy and timeliness directly affect the data warehouse's timeliness; and customer information systems, which contain customer transaction history, account information, and other data, and whose data query needs influence the data warehouse's query patterns and resource allocation.
[0049] Performance data are key indicators reflecting the operational status and performance of related financial systems. Typical indicators include: CPU utilization, i.e., the load on the system's processors; memory usage, i.e., the memory resources currently consumed by the system, reflecting the pressure on data processing and caching; disk I / O load, i.e., the intensity of disk read and write activities, which is closely related to the write and query performance of the data warehouse; network bandwidth usage, i.e., the speed and efficiency of data transmission between systems, affecting the real-time data ingestion of the data warehouse and communication with external systems; and thread and process status, i.e., the number and status of active threads and processes in the system, revealing the situation of concurrent operations.
[0050] Based on the above concepts, step S201 collects the operational status of the financial data warehouse itself and related financial systems through a real-time monitoring agent. This not only reflects the current system load but also provides a basis for predicting future data growth trends and performance requirements. This ensures that the financial data warehouse can make informed resource adjustments and maintain an efficient and stable operating status when facing complex scenarios such as short-term data surges.
[0051] Optionally, in the bucket expansion control method based on integrated decision tree provided in this embodiment of the invention, the step of collecting data bucket performance data of the financial data warehouse and the operation performance data of the associated financial system driving the operation of the financial data warehouse includes: calling a first agent deployed in association with the metadata interface of the financial data warehouse, and obtaining data bucket operation logs through the first agent; extracting data bucket performance data from the data bucket operation logs, wherein the data bucket performance data includes at least the write rate and storage capacity of all data buckets in the financial data warehouse; calling a second agent deployed in association with the business server of the financial system, and collecting system operation logs through the second agent; extracting operation performance data from the system operation logs, wherein the operation performance data includes at least the CPU utilization and memory occupancy of all computing nodes in the financial system.
[0052] In this embodiment of the invention, the first agent refers to a group of monitoring components deployed and tightly integrated with the metadata interface of the financial data warehouse. These components are responsible for capturing and collecting the operational logs of the data buckets in real time. The logs contain key performance indicators for each data bucket when processing financial data, such as write rate and storage capacity. By calling the metadata interface, the first agent can seamlessly access the internal state information of the data tables, including but not limited to file size, number of records, and partition information, as a basis for evaluating the load and performance of the data buckets.
[0053] The collected data bucket performance data not only reveals the real-time write and storage status of each data bucket, but also helps the system identify potential data unevenness or hotspot areas. For example, if the write rate of a data bucket is much higher than the average, it may be due to data skew, where some transaction data or market information is concentrated in a few buckets, causing these buckets to be overloaded and thus affecting query response time and overall system performance.
[0054] In this embodiment of the invention, the second agent is deployed in association with the business server of the financial system to collect and analyze system operation logs and extract operation performance data, including but not limited to CPU utilization and memory usage of all computing nodes, providing a system-level perspective for intelligent bucket expansion control.
[0055] Among the collected performance data, indicators such as CPU utilization and memory usage are used to reflect the stress level of computing resources. For example, if the CPU utilization is detected to be close to saturation while the data write rate is still increasing, it indicates that the system is approaching its processing limit and requires immediate measures such as expanding the bucket or other resource optimization measures.
[0056] By combining the first and second agents, this embodiment of the invention can capture changes in performance indicators at the data bucket and system levels in real time through agent components in the intelligent bucket expansion strategy of financial data warehouses, ensuring the accuracy and timeliness of decision-making. The extracted data bucket performance data and operational performance data provide rich input information for the intelligent bucket expansion control algorithm, enabling the decision tree regression model to accurately predict performance change trends and propose reasonable bucket expansion strategy suggestions based on historical data and the current system status.
[0057] This invention also enables more granular resource allocation management and avoids resource waste through continuous monitoring of system and data bucket performance. For example, during off-peak periods, the number of data buckets can be reduced based on actual load to decrease storage overhead; while during peak periods, intelligent bucket expansion and resource scheduling ensure sufficient system resources and avoid costly over-configuration. Real-time monitoring of performance data also helps to provide early warnings of potential system performance problems. Intelligent bucket expansion strategies can be used to adjust data distribution in a timely manner, reducing the load on individual data buckets and ensuring that query delays or system crashes do not occur during data surges, thus guaranteeing business continuity and customer experience.
[0058] Optionally, in the bucket expansion control method based on integrated decision tree provided in this embodiment of the invention, after collecting the data bucket performance data of the financial data warehouse and the operational performance data of the associated financial system driving the operation of the financial data warehouse, the method further includes: performing time-series integrity verification on the data bucket performance data and operational performance data, and removing time period data with missing rates exceeding a preset threshold based on the verification results; normalizing all verified performance data according to indicator type, wherein the performance data of the same indicator type after normalization are in the same dimension range; performing feature analysis on the normalized performance data based on a preset sliding time window to obtain statistical features, wherein the statistical features include at least the mean, range, and rate of change within the sliding time window; combining the statistical features with the original performance data to obtain a feature matrix, and determining the feature matrix as the decision basis.
[0059] After collecting performance data from the data buckets of the financial data warehouse and the operational performance data of the associated financial systems, preprocessing can be performed. This includes time-series integrity verification, which checks the time-series integrity of the data and identifies and removes missing data caused by network latency, equipment failure, or other reasons. If the data missing rate for a certain period exceeds a preset threshold, it indicates that the performance indicators for that period are unreliable, and this data is ignored to prevent it from affecting decision-making results.
[0060] Data that has passed time-series integrity verification can be normalized to convert performance metrics of different dimensions to the same range or scale, eliminating the impact of dimensional differences on the analysis results. Normalization ensures that all performance data are within the same dimensional range. For example, metrics such as CPU utilization, write speed, and storage capacity can be converted to a value range of 0 to 1, and then directly used for mathematical operations and model training.
[0061] Furthermore, feature analysis can be performed on the time-series verified and normalized data, and statistical feature extraction can be performed using a sliding time window. Specifically, a preset sliding time window is used to perform feature analysis on the normalized performance data to extract statistical features (such as mean, range, rate of change, etc.).
[0062] It's important to note that sliding time windows can capture trends and patterns in data over time. For financial data warehouses, this helps analyze the fluctuations in real-time data streams and predict future trends in data volume and system load. Extracting statistical features is used to quantify data fluctuations and provide input data for decision tree regression models.
[0063] Furthermore, a feature matrix can be constructed, combining statistical features with raw performance data to form a feature matrix containing various indicators and statistics. This matrix serves as input to a machine learning model, integrating all relevant system status information to reflect the current operational status and historical trends of the financial data warehouse. Constructing the feature matrix provides comprehensive decision-making support, including but not limited to data bucket write rates, storage capacity, CPU utilization, and memory usage. This allows for more accurate performance predictions and bucket expansion strategy recommendations based on multi-dimensional data.
[0064] Step S202: Decision results are generated using the decision tree regression algorithm and decision basis. The decision tree regression algorithm pre-configures N decision trees constructed through random feature selection based on the ensemble framework. The N decision trees generate decision suggestions in N dimensions based on the decision basis and are averaged and aggregated to obtain the decision results. N is a first preset value.
[0065] Specifically, decision tree regression is a machine learning method for predicting continuous values. It predicts output variables by constructing a tree-structured model. In the intelligent bucket expansion strategy of financial data warehouses, decision tree regression is used to analyze real-time performance monitoring data, predict future data growth patterns and system load trends, and make bucket expansion strategy recommendations based on these predictions. It can handle multi-dimensional input data, including data volume, data write rate, system resource usage, etc. It progressively identifies key factors affecting system performance through the hierarchical structure of the tree, and finally outputs prediction results regarding bucket index adjustments.
[0066] Ensemble frameworks are methods in machine learning that improve prediction accuracy and model stability by combining the prediction results of multiple base models (such as decision trees). The ensemble framework used in this embodiment is based on the construction principle of random forests, which involves constructing multiple decision trees through random feature selection and random sample sampling. Each decision tree learns independently and provides prediction results. Ensemble frameworks can reduce the risk of overfitting by decreasing the correlation between models, while simultaneously improving prediction performance by increasing model diversity.
[0067] Decision trees are the basic building blocks of decision tree regression algorithms, classifying input data or predicting output values through a series of "feature condition tests." In the intelligent bucket expansion strategy of financial data warehouses, each decision tree is trained based on a subset of randomly selected features (such as CPU utilization, data bucket write rate, query response time, etc.). By learning from historical performance data and the effects of bucket expansion operations, it predicts the most suitable bucket adjustment scheme under specific system conditions. Decision trees can intuitively demonstrate the decision-making logic through path selection, providing a clear basis for choosing the bucket expansion strategy.
[0068] Decision recommendations are predictive suggestions from the decision tree regression algorithm based on real-time monitoring data and historical learning results regarding bucket expansion strategies in the current system state. These recommendations may include key strategy parameters such as whether bucket expansion is needed, the number of buckets to expand, and the timing of expansion. The process of generating decision recommendations involves analyzing and predicting trends in performance metrics such as current data volume, data write rate, and query load, as well as the impact of these metrics on system resource utilization and query efficiency, to determine the most effective bucket expansion action.
[0069] Average aggregation is a computational strategy employed by the ensemble framework when processing predictions from multiple decision trees. It integrates the predictions from all decision trees to arrive at a final decision recommendation. In the intelligent bucket expansion strategy, since each decision tree may provide different bucket expansion recommendations based on different feature subsets, calculating a weighted average of all decision tree recommendations can reduce decision uncertainty and improve decision stability and accuracy. Average aggregation is a key step in the random forest regression algorithm, effectively avoiding the bias or randomness that may arise from a single model by combining the predictive power of multiple models.
[0070] The decision outcome is a specific bucket expansion strategy derived from a comprehensive analysis of decision tree regression algorithm and real-time monitoring data. This includes whether the number of data buckets needs to be adjusted, the number of buckets to be adjusted (i.e., how many new buckets to add), when to perform the bucket expansion operation, and other resource scheduling suggestions.
[0071] Through the above steps S202, intelligent decision suggestions for data bucket adjustment can be generated based on real-time performance monitoring data and pre-trained decision tree regression models. This enables the financial data warehouse to achieve adaptive management when facing highly volatile data volumes and query loads, maintaining the efficient and stable operation of the system. It not only improves the flexibility and response speed of the financial data warehouse in real-time data stream processing scenarios, but also significantly enhances its intelligent operation and maintenance and resource optimization capabilities through the application of intelligent algorithms.
[0072] Optionally, in the bucket expansion control method based on integrated decision trees provided in this embodiment of the invention, the step of generating decision results using decision tree regression algorithm and decision basis includes: inputting decision basis into N parallel-trained decision trees, with each decision tree randomly extracting a preset proportion of feature subsets from the decision basis for decision calculation when splitting tree nodes, and outputting N decision suggestions, wherein each decision suggestion records the suggestion value of the corresponding decision tree for the number of data buckets; preprocessing the N suggestion values recorded in the N decision suggestions, wherein the preprocessing includes at least: eliminating outlier prediction values; calculating the weighted average of the preprocessed suggestion values, wherein the weight values for the weighted average calculation are predetermined by the decision tree regression algorithm; and mapping the weighted average to a preset bucket number range to obtain the final decision result.
[0073] In embodiments of the present invention, each decision tree does not use all features when performing decision calculations, but instead randomly selects a preset proportion of features from the decision criteria. By randomly selecting features, the learning process and decision logic of each decision tree are different, increasing model diversity, reducing correlation between models, and improving the generalization ability of the entire ensemble model.
[0074] Randomly selecting features can also prevent decision trees from over-relying on a certain set of features, reduce the risk of overfitting a single decision tree, and reduce computational complexity by using fewer features to perform calculations when splitting at each node of the decision tree, thereby speeding up the decision-making process.
[0075] Another point to note is that decision recommendations are generated independently by each decision tree, and each tree contains recommended values for at least the number of data buckets. However, the possibility of outliers or discrepancies in predictions due to biases in individual decision trees cannot be ruled out. To improve the reliability and accuracy of the decision results, these recommended values need to be preprocessed, including at least outlier removal, by identifying and eliminating recommended values that significantly deviate from other values through statistical analysis or machine learning algorithms.
[0076] The preprocessed suggested values are weighted and averaged, with the weights predetermined by the decision tree regression algorithm. The weights reflect the importance of each decision tree in the model and can be set based on the model's performance during training (e.g., accuracy, bias).
[0077] Finally, the weighted average value is mapped to a preset range of bucket numbers to obtain the final decision result, avoiding overly extreme suggestions to expand or shrink buckets, and maintaining system stability and appropriate resource utilization.
[0078] Through the aforementioned enhanced decision generation steps, this embodiment of the invention not only solves the static configuration problem in the prior art in the dynamic expansion management of financial data warehouses, but also achieves comprehensive technical breakthroughs such as decision robustness, balance between decision speed and quality, resource optimization, preventive maintenance, and model self-optimization.
[0079] Step S203: Based on the decision result, determine the target bucket expansion strategy from M bucket expansion strategies with predetermined priorities, where M is a second preset value.
[0080] The bucket expansion strategy in this embodiment of the invention is dynamically selected based on the real-time performance requirements and prediction results of the financial data warehouse. These strategies are assigned different priorities at the beginning of the design to ensure that the most efficient or most urgent strategy can be executed first among a variety of possible bucket expansion needs.
[0081] The predefined bucket expansion strategy covers a variety of bucket expansion scenarios and operation types. For example, it can dynamically adjust the number of data buckets based on real-time monitoring data and prediction results from decision tree regression models to adapt to sudden changes in data volume and prioritize handling situations such as data skew or excessive system load. It can also automatically increase the number of data buckets at pre-set time intervals or specific moments, such as before business peaks, to suit predictable periodic data growth patterns and ensure system stability and advance resource preparation. Furthermore, it allows system administrators or operations personnel to manually trigger bucket expansion operations based on experience and specific business needs, serving as a supplement to intelligent and scheduled strategies to handle resource adjustment needs in emergency situations or special circumstances.
[0082] Based on the decision results, the optimal expansion strategy is selected and implemented from multiple predefined expansion strategies, following a preset strategy priority order. This ensures that the strategy with the highest intelligence and immediate effect is executed first, thereby quickly responding to changes in system requirements and avoiding performance bottlenecks or resource waste. The priority order considers factors such as the strategy's real-time performance, resource consumption, operational complexity, and impact on system stability and business continuity. Once the target expansion strategy is determined, the execution mechanism is immediately initiated, performing specific expansion operations through the data table management API or similar interface, including data redistribution and necessary adjustments to the file system.
[0083] In a specific implementation scenario, suppose a financial data warehouse system detects a sudden increase in the transaction data write rate and a corresponding increase in query response time during real-time monitoring. The necessity of expanding the bucket count and the optimal number of buckets are determined based on predictions from a decision tree regression algorithm. Initially, the highest priority strategy is adopted, quickly alleviating data skew and system load by redistributing data and dynamically adjusting the number of buckets. If this fails to resolve the issue or if future data growth trends are predicted to exceed the current dynamic adjustment capabilities, the next priority strategy is considered, increasing the number of buckets in advance based on foreseeable future events. Finally, if the system administrator, based on experience, deems there to be special resource needs or an emergency, they can manually trigger the bucket expansion operation.
[0084] By employing a pre-defined bucket expansion strategy system, this invention not only intelligently handles sudden changes in data volume but also pre-allocates resources during foreseeable periodic peaks, while retaining the possibility of manual intervention. This ensures that the financial data warehouse can make the most reasonable and efficient resource adjustments when facing complex and ever-changing data processing demands, maintaining a high level of system performance. Furthermore, the strategy's priority ranking and dynamic coordination mechanism effectively avoids resource conflicts, ensuring the smooth execution of bucket expansion operations and minimizing interference with online business.
[0085] Optionally, in the bucket expansion control method based on integrated decision tree provided in this embodiment of the invention, the step of determining the target bucket expansion strategy from M bucket expansion strategies with predetermined priorities based on the decision result includes: calculating the difference between the bucket number suggestion value indicated by the decision result and the current bucket number of the financial data warehouse; if the difference is less than a first preset threshold, determining the target bucket expansion strategy as a timed bucket expansion strategy; if the difference is greater than or equal to the first preset threshold and less than a second preset threshold, determining the target bucket expansion strategy as a machine learning bucket expansion strategy, wherein the machine learning bucket expansion strategy is used to indicate the generation of an automatic execution program that expands buckets in stages according to the difference, and the second preset threshold is greater than the first preset threshold; if the difference is greater than or equal to the second preset threshold or a manual intervention instruction is detected, determining the target bucket expansion strategy as a manual bucket expansion strategy, and sending a manual bucket expansion instruction to the manual operation and maintenance terminal.
[0086] In this embodiment of the invention, the difference between the suggested number of buckets indicated by the calculation decision result and the current number of buckets in the financial data warehouse provides a quantitative indicator of the bucket expansion demand, determines the urgency and scale of bucket expansion, and then decides which bucket expansion strategy to adopt and executes a tiered response of the bucket expansion strategy.
[0087] Level 1 response: Timed bucket expansion strategy.
[0088] When the difference is less than a first preset threshold, the timed bucket expansion strategy is activated, indicating that the current bucket expansion demand is relatively small and can be met through periodic, preset bucket expansion operations. The timed bucket expansion strategy can provide stable resource adjustment in scenarios with small data fluctuations or high predictability, reducing unnecessary system burden and resource waste.
[0089] Second-level response: Machine learning bucket expansion strategy.
[0090] When the difference is greater than or equal to the first preset threshold and less than the second preset threshold, a machine learning bucket expansion strategy is adopted. Based on the prediction results of the random forest regression model, an automatic execution program for phased bucket expansion is intelligently generated. This is suitable for scenarios where data fluctuations are relatively obvious but still within a controllable range. The number of buckets is dynamically adjusted to cope with the rapid growth of data volume, while avoiding the impact of a one-time large-scale bucket expansion on system performance and resources.
[0091] Level 3 response: Manual bucket expansion strategy.
[0092] When the difference is greater than or equal to the second preset threshold or a manual intervention command is detected, the target bucket expansion strategy is determined to be a manual bucket expansion strategy, and a manual bucket expansion command is sent to the manual operation and maintenance terminal. The manual bucket expansion strategy is a bucket expansion operation intervened by professionals when the data volume increases sharply or when unforeseen abnormal situations occur in the system. Although the response speed may not be as fast as the automatic strategy, it is more flexible and capable of handling complex problems, and is suitable for resource adjustment and performance optimization in extreme scenarios.
[0093] This invention achieves a tiered response to the bucket expansion strategy by comparing the difference with a preset threshold. This allows for the adoption of the most suitable and economical bucket expansion measures under different scales of demand, avoiding over-allocation or under-allocation of resources and achieving refined resource management. Phased bucket expansion also allows for gradual adjustments while avoiding sudden changes in system performance, ensuring that the read / write performance and query efficiency of the data warehouse are not significantly affected during the bucket expansion process, thus maintaining system stability and business continuity.
[0094] The embodiments of the present invention combine machine learning-based bucket expansion strategies with manual bucket expansion strategies, which not only leverages the automation and predictive capabilities of machine learning but also preserves the decision-making power of professionals in extreme situations. This achieves complementarity between intelligent decision-making and human experience, enhancing the system's ability to cope with complex scenarios.
[0095] Scheduled bucket expansion strategies reduce unnecessary bucket expansion operations when data fluctuations are small, thus lowering resource consumption and maintenance costs. Machine learning bucket expansion strategies can adjust resources more quickly and accurately when the data volume grows appropriately, avoiding resource waste. Manual bucket expansion strategies, on the other hand, avoid the blindness of system decision-making through the judgment of professionals in extreme cases, ensuring that resource adjustments match actual business needs.
[0096] Step S204: Perform data redistribution operation on the financial data warehouse based on the target bucket expansion strategy.
[0097] It should be noted that once the target bucket expansion strategy is determined, a data redistribution operation needs to be performed, which means redistributing the existing data to the newly added data buckets according to the new strategy. This ensures that the data can be evenly distributed when the data volume suddenly increases or the query pattern changes, thus avoiding the generation of hotspots and performance bottlenecks.
[0098] The data redistribution operation follows these steps: First, thoroughly analyze the details of the target bucket expansion strategy, clarifying the number of new data buckets, distribution rules, and any specific index fields or partitioning strategies. Based on the strategy analysis results, generate a detailed execution plan, including how data will be redistributed from the current buckets to the new buckets and whether data preprocessing or format conversion is required. Move or copy data from the original buckets to the new buckets according to the generated execution plan. To minimize the impact on running services, a gradual migration strategy is typically adopted, i.e., data migration is performed in stages without affecting query or write operations. After data migration is complete, update the data indexes to ensure they match the new data bucket layout. In addition to the physical data redistribution and index updates, metadata, including file location, data size, and partition information, also needs to be synchronized to ensure all system components are aware of the latest data layout and to avoid errors during queries or writes. Finally, perform performance verification to check whether the bucket expansion and data redistribution operations have achieved the expected results, such as significant improvements in query response time and write speed. If performance issues persist, further optimize the data layout or indexing strategy based on the actual situation.
[0099] It's important to note that incremental data migration is an efficient and low-interference data redistribution method that allows data to be redistributed while the system continues to run. By controlling the rate and scope of data migration, it ensures that data redistribution does not lead to a significant degradation in system performance. In financial data warehouse scenarios, incremental data migration strategies can achieve effective resource utilization and a smooth performance transition during sensitive periods such as peak business periods, avoiding negative impacts on real-time transaction or query services.
[0100] Optionally, in the bucket expansion control method based on integrated decision tree provided in this embodiment of the invention, the step of performing data redistribution operation on the financial data warehouse based on the target bucket expansion strategy includes: when the target bucket expansion strategy is a timed bucket expansion strategy, reading the preset time plan and preset number of adjustment buckets recorded in the timed bucket expansion strategy; when the start time point specified by the preset time plan is detected, generating a new bucket index mapping table according to the preset number of adjustment buckets; within the migration time period specified by the preset time plan, migrating the data buckets in the financial data warehouse that meet the data migration conditions based on the new bucket index mapping table; after the data migration is completed, updating all data bucket indexes through the metadata interface of the financial data warehouse.
[0101] In this embodiment of the invention, when the target bucket expansion strategy is determined to be a timed bucket expansion strategy, the preset time plan and preset adjustment bucket quantity recorded in the timed bucket expansion strategy are first read. The preset time plan is the trigger mechanism for the bucket expansion operation, while the preset adjustment bucket quantity guides the specific bucket expansion scale. Upon detecting that the current time has reached the start time specified by the preset time plan, a new bucket index mapping table is generated based on the preset adjustment bucket quantity. The new bucket index mapping table defines the data distribution rules from the original buckets to the new buckets, ensuring that data can be evenly distributed among the newly added buckets and avoiding data skew.
[0102] Data migration is conducted within a pre-defined migration timeframe. This timeframe is chosen considering the cyclical nature of financial business and the data warehouse's load, ensuring that bucket expansion will not significantly impact ongoing operations. Based on the new bucket index mapping table, data buckets in the financial data warehouse that meet the migration criteria are migrated sequentially or in batches, ensuring data integrity and consistency. After data migration is complete, all bucket indexes are updated via the metadata interface to confirm that the new buckets have been correctly created and the data has been correctly redistributed.
[0103] The scheduled bucket expansion strategy achieves precise control over the timing and scale of bucket expansion operations by pre-setting time plans and adjusting the number of buckets, avoiding resource waste and system instability caused by blind bucket expansion. Data migration during off-peak periods effectively utilizes idle computing resources, reduces the impact of the migration process on real-time services, ensures business continuity, and also lowers operational costs, achieving rational utilization and optimization of resources.
[0104] Through the above-described enhanced timed bucket expansion strategy execution steps, this embodiment of the invention not only solves the timing and scale issues of bucket expansion operations when facing foreseeable changes in data volume, but also achieves additional beneficial effects such as resource optimization, business continuity assurance, balanced data distribution, and maintenance of system state consistency.
[0105] Optionally, in the bucket expansion control method based on ensemble decision tree provided in this embodiment of the invention, the step of performing data redistribution operation on the financial data warehouse based on the target bucket expansion strategy further includes: when the target bucket expansion strategy is a machine learning bucket expansion strategy, configuring S migration stages and the target migration bucket number for each migration stage based on the difference to obtain a configuration document, where S is a positive integer greater than or equal to 2; writing an automatic execution program according to the configuration document, and compiling and executing the automatic execution program to obtain an execution result, wherein the execution result records the data redistribution result of the financial data warehouse.
[0106] In this embodiment of the invention, when the target bucket expansion strategy is determined to be a machine learning bucket expansion strategy, multiple migration stages and the target migration bucket number for each stage are configured based on the difference, and a configuration document is generated throughout the process to guide the bucket expansion operation.
[0107] Specifically, the bucket expansion operation is divided into multiple stages based on the difference output by the integrated decision tree model. The target number of buckets to migrate in each stage is determined by a comprehensive evaluation of the difference and the current system resource status. The goal is to smooth the bucket expansion process and reduce the impact of a single bucket expansion on system performance.
[0108] Detailed configuration documents are created based on phased planning, recording the target number of buckets, estimated execution time, and specific rules for data migration at each stage. This serves as a blueprint for bucket expansion operations, ensuring the orderly and traceable nature of the process. Corresponding automated execution programs are automatically generated based on the configuration documents. These programs automatically read the data warehouse's status information, perform data redistribution operations, and update metadata after each stage of migration is completed. The compilation and execution of these automated programs ensure the automation and efficiency of the bucket expansion process, reducing manual intervention and mitigating the risk of errors.
[0109] After the automatic execution program is completed, the data redistribution results of the financial data warehouse will be recorded in the execution results, including key information such as the number of data buckets before and after bucket expansion, the data distribution status, and changes in query performance.
[0110] This invention implements a gradual resource adjustment mechanism, effectively avoiding system instability and resource waste that may result from a large-scale, one-time bucket expansion, ensuring the smoothness of the bucket expansion operation and the stability of the data warehouse performance. The automatic execution of programs significantly reduces manual intervention, improving efficiency while ensuring operational consistency and accuracy, and reducing the risk of human error.
[0111] In addition, phased expansion can be dynamically adjusted according to actual needs and resource conditions, avoiding over-expansion or under-expansion, and achieving the best balance between cost and benefit. Especially when facing market fluctuations and changes in business needs, it can make more effective use of resources and reduce unnecessary expenses.
[0112] Through steps S201 to S204, an ensemble decision tree machine learning approach is employed. By combining a real-time monitoring agent with a decision tree regression algorithm, the goal of intelligently and dynamically adjusting the bucket index of the financial data warehouse is achieved. This results in a significant improvement in system performance and a high degree of optimization in resource utilization efficiency under scenarios of sudden data volume increases. Specifically, the real-time monitoring agent collects performance data of data buckets in the financial data warehouse and the operating performance indicators of the financial system supporting its operation. Then, using multiple decision trees within the ensemble framework, each tree independently generates decision suggestions through a random feature selection strategy. Finally, a weighted average is used to aggregate suggestions from multiple dimensions to obtain a comprehensive and accurate bucket index augmentation decision result, fully considering the current and expected system load. The system first assesses the load conditions to ensure the scientific validity and effectiveness of the bucket expansion strategy. Then, based on the decision results, it selects the most suitable target strategy from multiple bucket expansion strategies with pre-set priorities, balancing the flexibility of automated decision-making with the necessity of human intervention. Finally, it performs fine-grained data redistribution operations on the financial data warehouse according to the selected target bucket expansion strategy. This not only balances the data storage load but also optimizes query performance, ensuring the high availability and efficiency of the data warehouse. It overcomes the inherent static defects of traditional financial data warehouse indexing strategies, realizes dynamic adaptive adjustment of bucket indexes, and solves the technical problem in related technologies where financial data warehouse indexing strategies cannot adapt to changes in data volume in real time, resulting in uneven data distribution and low system resource utilization efficiency.
[0113] The present invention will now be described in conjunction with another specific embodiment.
[0114] Figure 3 This is a schematic diagram of an optional data table bucket index management system based on a random forest regression model in a real-time financial data warehouse scenario, according to an embodiment of the present invention. Figure 3 As shown, the system includes: a real-time performance monitoring module, a dynamic bucket expansion strategy module, and an intelligent data redistribution module. The real-time performance monitoring module transmits the captured performance data to the dynamic bucket expansion strategy module, which generates bucket expansion instructions and commands the intelligent data redistribution module to execute the instructions and collect execution status feedback.
[0115] Specifically, the real-time performance monitoring module is responsible for monitoring various system performance metrics in real time, including data volume, write rate, query load, and system resource usage. This module is the foundation of the entire dynamic adaptive bucket index management system, responsible for continuously monitoring various system performance metrics. It continuously collects, processes, and analyzes various metrics during system operation through multiple data acquisition methods to form a comprehensive system performance profile, providing decision-making basis for other modules.
[0116] Monitoring metrics include: bucket data volume, write rate, query response time, CPU utilization, memory usage, disk I / O utilization, garbage collection time, and number of threads.
[0117] The dynamic bucket expansion strategy module implements three complementary bucket expansion strategies to address data growth needs in different scenarios: scheduled bucket expansion, intelligent bucket expansion, and manual bucket expansion. The three strategies work together according to the following rules: 1. Priority: Intelligent expansion > Scheduled expansion > Manual bucket expansion; 2. Conflict handling: If multiple strategies are triggered simultaneously, they are executed according to priority, and conflict details are recorded; 3. Feedback mechanism: The result of each bucket expansion operation is fed back to the intelligent expansion model for continuous optimization of prediction accuracy.
[0118] The scheduled bucket expansion strategy refers to automatically triggering bucket expansion operations based on preset time points, which is suitable for scenarios with predictable periodic data growth. The general workflow of scheduled bucket expansion is as follows: preset the bucket expansion time points and expansion ratios, maintain a scheduled task queue, and automatically trigger the bucket expansion operation when the preset time points are reached.
[0119] The intelligent bucket expansion strategy refers to dynamically judging and executing bucket expansion operations based on real-time monitoring data and machine learning prediction models. It is suitable for scenarios with unpredictable sudden data growth. The main steps include:
[0120] S1, Data Acquisition: Real-time acquisition of system performance metrics (CPU utilization, memory usage, disk I / O utilization, query response time); collection of data table information (current number of buckets, data volume in each bucket, data growth rate, number of files); recording of historical bucket expansion operations and their effects.
[0121] S2, Feature Engineering: Data cleaning (removing outliers and handling missing data); Feature standardization (normalizing features of different dimensions to a unified range); Feature encoding (converting categorical features into numerical representations).
[0122] S3, Model Training: Construct a set of decision trees, each tree using a randomly selected subset of features; set the model input to system metrics, data table information, and time features; set the model output to predicted performance metrics and suggested bucket number adjustments; periodically retrain the model using newly collected data to adapt to dynamic changes in the system.
[0123] S4, Real-time Prediction: Collect current system state data periodically (e.g., every 5 minutes); apply feature engineering to the collected data; input the processed data into the trained random forest model; obtain prediction results, including performance predictions and suggested bucket numbers for a future period.
[0124] S5, Decision Execution: Evaluate the prediction results based on predefined performance thresholds; if the predicted performance is lower than the threshold, trigger the bucket expansion operation; calculate the number of buckets to be added based on the predicted optimal number of buckets and the current number of buckets; call the data table API to execute the bucket expansion operation; record the details of the bucket expansion operation, including operation time, number of buckets added, etc.
[0125] S6, Feedback Optimization: Monitor the actual performance changes after expanding the bucket; add the actual performance data to the training set for subsequent model updates and optimizations.
[0126] The data redistribution module is responsible for executing the actual redistribution of data in the data table according to the bucket expansion strategy generated by the dynamic bucket expansion strategy module. This ensures efficient, safe, and balanced distribution of data in the data table during the bucket expansion process, while minimizing the impact on system performance.
[0127] The invention will now be described in conjunction with another alternative embodiment.
[0128] Example 2
[0129] This invention also provides a bucket expansion control device based on an integrated decision tree. It should be noted that the bucket expansion control device based on an integrated decision tree in this invention includes multiple implementation units, which can be used to execute the bucket expansion control method based on an integrated decision tree provided in Embodiment 1 above. Each implementation unit corresponds to each implementation step in Embodiment 1 above.
[0130] Figure 4 This is a schematic diagram of an optional bucket expansion control device based on an integrated decision tree according to an embodiment of the present invention, as shown below. Figure 4 As shown, the device may include: a data acquisition unit 41, a generation unit 42, a determination unit 43, and an execution unit 44.
[0131] The system comprises the following components: Collection unit 41 collects performance data of data buckets in the financial data warehouse and performance data of related financial systems driving the data warehouse operation via a pre-deployed real-time monitoring agent, thus obtaining decision-making basis. Generation unit 42 generates decision results using a decision tree regression algorithm and the decision basis. The decision tree regression algorithm pre-configures N decision trees constructed through random feature selection based on an integrated framework. These N decision trees generate N-dimensional decision suggestions based on the decision basis and are then averaged and aggregated to obtain the decision result, where N is a first preset value. Determination unit 43 determines the target bucket expansion strategy from M bucket expansion strategies with pre-determined priorities based on the decision result, where M is a second preset value. Execution unit 44 performs data redistribution operations on the financial data warehouse based on the target bucket expansion strategy.
[0132] The aforementioned bucket expansion control device based on ensemble decision trees employs a machine learning approach using ensemble decision trees. By combining a real-time monitoring agent with a decision tree regression algorithm, it achieves intelligent and dynamic control of the bucket indexes in the financial data warehouse. This results in a significant improvement in system performance and a high degree of optimization in resource utilization efficiency under scenarios of sudden data volume increases. Specifically, the real-time monitoring agent collects performance data of data buckets in the financial data warehouse, as well as the operational performance indicators of the supporting financial system. Subsequently, it utilizes multiple decision trees within the ensemble framework, with each tree independently generating decision suggestions through a random feature selection strategy. Then, it aggregates these suggestions from multiple dimensions through a weighted average to obtain a comprehensive and accurate bucket index expansion decision result, fully considering both current and expected system load. The system first assesses the load conditions to ensure the scientific validity and effectiveness of the bucket expansion strategy. Then, based on the decision results, it selects the most suitable target strategy from multiple bucket expansion strategies with pre-set priorities, balancing the flexibility of automated decision-making with the necessity of human intervention. Finally, it performs fine-grained data redistribution operations on the financial data warehouse according to the selected target bucket expansion strategy. This not only balances the data storage load but also optimizes query performance, ensuring the high availability and efficiency of the data warehouse. It overcomes the inherent static defects of traditional financial data warehouse indexing strategies, realizes dynamic adaptive adjustment of bucket indexes, and solves the technical problem in related technologies where financial data warehouse indexing strategies cannot adapt to changes in data volume in real time, resulting in uneven data distribution and low system resource utilization efficiency.
[0133] Furthermore, the acquisition unit includes: a first invocation module, used to invoke a first agent deployed in association with the metadata interface of the financial data warehouse, and obtain data bucket operation logs through the first agent; a first extraction module, used to extract data bucket performance data from the data bucket operation logs, wherein the data bucket performance data includes at least the write rate and storage capacity of all data buckets in the financial data warehouse; a second invocation module, used to invoke a second agent deployed in association with the business server of the financial system, and collect system operation logs through the second agent; and a second extraction module, used to extract operation performance data from the system operation logs, wherein the operation performance data includes at least the CPU utilization and memory usage of all computing nodes in the financial system.
[0134] Furthermore, the bucket expansion control device based on integrated decision trees also includes: a feature engineering unit, which includes: a verification module, used to perform time-series integrity verification on the data bucket performance data and operational performance data after collecting the data bucket performance data of the financial data warehouse and the operational performance data of the related financial systems driving the operation of the financial data warehouse, and to remove time period data with missing rates exceeding a preset threshold based on the verification results; a normalization processing module, used to normalize all verified performance data according to indicator type, wherein the performance data of the same indicator type after normalization are in the same dimension range; a feature analysis module, used to perform feature analysis on the normalized performance data based on a preset sliding time window to obtain statistical features, wherein the statistical features include at least the mean, range, and rate of change within the sliding time window; and a combination module, used to combine the statistical features with the original performance data to obtain a feature matrix, and to determine the feature matrix as the basis for decision-making.
[0135] Furthermore, the generation unit includes: an input / output module, used to input the decision basis into N parallel-trained decision trees, whereby each decision tree randomly extracts a preset proportion of feature subsets from the decision basis when splitting at a tree node, performs decision calculation, and outputs N decision suggestions, wherein each decision suggestion records the suggestion value of the corresponding decision tree for the number of data buckets; a preprocessing module, used to preprocess the N suggestion values recorded in the N decision suggestions, wherein the preprocessing includes at least: eliminating outlier prediction values; a first calculation module, used to calculate the weighted average of the preprocessed suggestion values, wherein the weight values for the weighted average calculation are predetermined by the decision tree regression algorithm; and a mapping module, used to map the weighted average to a preset number of buckets range to obtain the final decision result.
[0136] Further, the determining unit includes: a second calculation module, used to calculate the difference between the suggested number of buckets indicated by the decision result and the current number of buckets in the financial data warehouse; a first determining module, used to determine the target bucket expansion strategy as a timed bucket expansion strategy when the difference is less than a first preset threshold; a second determining module, used to determine the target bucket expansion strategy as a machine learning bucket expansion strategy when the difference is greater than or equal to the first preset threshold and less than a second preset threshold, wherein the machine learning bucket expansion strategy is used to instruct the generation of an automated execution program that expands buckets in stages according to the difference, and the second preset threshold is greater than the first preset threshold; and a third determining module, used to determine the target bucket expansion strategy as a manual bucket expansion strategy when the difference is greater than or equal to the second preset threshold or a manual intervention instruction is detected, and to send a manual bucket expansion instruction to the manual operation and maintenance terminal.
[0137] Furthermore, the execution unit includes: a reading module, used to read the preset time plan and preset number of buckets to be adjusted recorded in the timed bucket expansion strategy when the target bucket expansion strategy is a timed bucket expansion strategy; a generation module, used to generate a new bucket index mapping table according to the preset number of buckets to be adjusted when the start time point specified by the preset time plan is detected; a migration module, used to migrate the data buckets in the financial data warehouse that meet the data migration conditions based on the new bucket index mapping table within the migration period specified by the preset time plan; and an update module, used to update all data bucket indexes through the metadata interface of the financial data warehouse after the data migration is completed.
[0138] Furthermore, the execution unit also includes: a configuration module, used to configure S migration stages and the target migration bucket number for each migration stage based on the difference when the target bucket expansion strategy is a machine learning bucket expansion strategy, to obtain a configuration document, where S is a positive integer greater than or equal to 2; and an execution module, used to write an automatic execution program according to the configuration document, and compile and execute the automatic execution program to obtain an execution result, wherein the execution result records the data redistribution result of the financial data warehouse.
[0139] It should be noted that the aforementioned acquisition unit 41, generation unit 42, determination unit 43, and execution unit 44 correspond to steps S201 to S204 in Embodiment 1. The instances and application scenarios implemented by the aforementioned units and corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should be noted that the aforementioned modules or units may be hardware or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The aforementioned modules or units may also be part of a device and can run in the computer terminal 10 provided in Embodiment 1.
[0140] The invention will now be described in conjunction with another alternative embodiment.
[0141] Example 3
[0142] The present invention can also provide an electronic device. Figure 5 This is a structural block diagram of an electronic device that implements an expansion bucket control method based on an integrated decision tree according to an embodiment of the present invention, such as... Figure 5 As shown, the electronic device may include: one or more ( Figure 5 Only one of the components is shown: processor 502, memory 504, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module, and display.
[0143] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the bucket expansion control method and apparatus based on integrated decision trees in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned bucket expansion control method based on integrated decision trees. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0144] The processor can access the information and application programs stored in the memory via a transmission device to execute the various implementation steps in Embodiment 1 above, including at least the following steps: Collecting data bucket performance data of the financial data warehouse and the operational performance data of the associated financial systems driving the operation of the financial data warehouse through a pre-deployed real-time monitoring agent to obtain decision-making basis; generating decision results using a decision tree regression algorithm and the decision-making basis, wherein the decision tree regression algorithm pre-configures N decision trees constructed through random feature selection based on an integrated framework, and the N decision trees generate decision suggestions in N dimensions based on the decision-making basis and perform average aggregation to obtain the decision result, where N is a first preset value; determining the target bucket expansion strategy from M bucket expansion strategies with pre-determined priorities based on the decision result, where M is a second preset value; and performing data redistribution operations on the financial data warehouse based on the target bucket expansion strategy.
[0145] This invention provides a bucket expansion and control scheme based on ensemble decision trees. By employing machine learning through ensemble decision trees, and combining real-time monitoring agents with decision tree regression algorithms, it achieves intelligent and dynamic control of the financial data warehouse bucket index. This significantly improves system performance and optimizes resource utilization efficiency under scenarios of sudden data volume increases. Furthermore, it solves the technical problem in related technologies where financial data warehouse indexing strategies cannot adapt to changes in data volume in real time, leading to uneven data distribution and low system resource utilization efficiency.
[0146] Those skilled in the art will understand that Figure 5 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 5 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 5The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 5 The different configurations shown.
[0147] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0148] The invention will now be described in conjunction with another alternative embodiment.
[0149] Example 4
[0150] This invention also provides a computer-readable storage medium. Optionally, in this invention, the computer-readable storage medium can be used to store the program code executed by the bucket expansion control method based on ensemble decision tree provided in Embodiment 1.
[0151] Optionally, in this embodiment of the invention, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0152] This invention also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of a bucket expansion control method based on an integrated decision tree: collecting data bucket performance data of a financial data warehouse and operational performance data of the associated financial systems driving the operation of the financial data warehouse through a pre-deployed real-time monitoring agent to obtain a decision basis; generating a decision result using a decision tree regression algorithm and the decision basis, wherein the decision tree regression algorithm pre-configures N decision trees constructed through random feature selection based on an integrated framework, and the N decision trees generate N-dimensional decision suggestions based on the decision basis and perform average aggregation to obtain the decision result, where N is a first preset value; determining a target bucket expansion strategy from M bucket expansion strategies with pre-determined priorities based on the decision result, where M is a second preset value; and performing a data redistribution operation on the financial data warehouse based on the target bucket expansion strategy.
[0153] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0154] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0155] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0156] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0157] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0158] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0159] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A bucket expansion control method based on ensemble decision tree, characterized in that, The method comprises the steps of: Collecting data bucket performance data of a financial data warehouse and operation performance data of an associated financial system driving the operation of the financial data warehouse through a pre-deployed real-time monitoring agent to obtain a decision basis; Generating a decision result using a decision tree regression algorithm and the decision basis, wherein the decision tree regression algorithm is pre-configured with N decision trees constructed through random feature selection based on an integrated framework, and the N decision trees generate N-dimensional decision suggestions based on the decision basis and perform average aggregation to obtain the decision result, N being a first preset value; Determining a target bucket expansion strategy among M bucket expansion strategies with predetermined priorities based on the decision result, wherein M is a second preset value; Performing a data redistribution operation on the financial data warehouse based on the target bucket expansion strategy.
2. The integrated decision tree based bucket expansion regulation method of claim 1, wherein, The step of collecting data bucket performance data of a financial data warehouse and operation performance data of an associated financial system driving the operation of the financial data warehouse comprises: Calling a first agent deployed in association with a metadata interface of the financial data warehouse to obtain data bucket operation logs through the first agent; Extracting the data bucket performance data from the data bucket operation logs, wherein the data bucket performance data at least includes the write rate and storage amount of all data buckets in the financial data warehouse; Calling a second agent deployed in association with a business server of the financial system to collect system operation logs through the second agent; Extracting the operation performance data from the system operation logs, wherein the operation performance data at least includes the CPU utilization and memory occupancy of all computing nodes in the financial system.
3. The integrated decision tree based bucket expansion regulation method of claim 1, wherein, After collecting data bucket performance data of a financial data warehouse and operation performance data of an associated financial system driving the operation of the financial data warehouse, the method further comprises: Performing time series integrity verification on the data bucket performance data and the operation performance data, and eliminating time period data with a missing rate exceeding a preset threshold based on the verification result; Performing normalization processing on all performance data after verification according to index types, wherein the performance data of the same index type after normalization processing is in the same dimension interval; Performing feature analysis on the performance data after normalization processing based on a preset sliding time window to obtain statistical features, wherein the statistical features at least include the mean, range, and change rate within the sliding time window; Combining the statistical features with the original performance data to obtain a feature matrix, and determining the feature matrix as the decision basis.
4. The integrated decision tree based bucket expansion regulation method of claim 1, wherein, The step of generating a decision result using a decision tree regression algorithm and the decision basis comprises: Inputting the decision basis into N parallel trained decision trees, and randomly extracting a preset proportion of feature subsets from the decision basis for decision calculation when each decision tree splits at a tree node, and outputting N decision suggestions, wherein each decision suggestion records the recommended value for the number of data buckets for the corresponding decision tree; Performing preprocessing on the N recommended values recorded in the N decision suggestions, wherein the preprocessing at least includes eliminating outlier predicted values. calculating a weighted average of the recommended values after preprocessing, wherein the weight values of the weighted average calculation are determined in advance by the decision tree regression algorithm; mapping the weighted average to a preset bucket number interval to obtain a final decision result.
5. The integrated decision tree based bucket expansion regulation method of claim 1, wherein, The step of determining a target bucket expansion strategy from the decision result among a plurality of bucket expansion strategies with predetermined priorities includes: calculating a difference between the bucket number recommended value indicated by the decision result and the current bucket number of the financial data warehouse; in the case that the difference is less than a first preset threshold, determining the target bucket expansion strategy as a timing bucket expansion strategy; in the case that the difference is greater than or equal to the first preset threshold and less than a second preset threshold, determining the target bucket expansion strategy as a machine learning bucket expansion strategy, wherein the machine learning bucket expansion strategy is used to indicate an automatic execution program for expanding buckets in stages according to the difference, and the second preset threshold is greater than the first preset threshold; in the case that the difference is greater than or equal to the second preset threshold or an artificial intervention instruction is detected, determining the target bucket expansion strategy as an artificial bucket expansion strategy, and sending an artificial bucket expansion instruction to an artificial operation and maintenance end.
6. The integrated decision tree based bucket expansion regulation method of claim 5, wherein, The step of performing a data redistribution operation on the financial data warehouse based on the target bucket expansion strategy includes: in the case that the target bucket expansion strategy is the timing bucket expansion strategy, reading a preset time plan and a preset adjustment bucket number recorded in the timing bucket expansion strategy; in the case that a start time point specified by the preset time plan is detected, generating a new bucket index mapping table according to the preset adjustment bucket number; in a migration time period specified by the preset time plan, migrating data buckets that meet a migration data condition in the financial data warehouse based on the new bucket index mapping table; after completing the data migration, updating all data bucket indexes through a metadata interface of the financial data warehouse.
7. The integrated decision tree based bucket expansion regulation method of claim 5, wherein, The step of performing a data redistribution operation on the financial data warehouse based on the target bucket expansion strategy further includes: in the case that the target bucket expansion strategy is the machine learning bucket expansion strategy, configuring S migration stages and a target migration bucket number of each migration stage based on the difference to obtain a configuration document, wherein S is a positive integer greater than or equal to 2; compiling and executing the automatic execution program according to the configuration document to obtain an execution result, wherein the execution result records a data redistribution result of the financial data warehouse.
8. An integrated decision tree based bucket extension regulating device, characterized in that, includes: a collection unit configured to collect data bucket performance data of a financial data warehouse and operation performance data of an associated financial system driving the financial data warehouse through a pre-deployed real-time monitoring agent to obtain decision basis; a generation unit configured to generate a decision result by using a decision tree regression algorithm and the decision basis, wherein the decision tree regression algorithm is pre-configured with N decision trees constructed by random feature selection based on an integrated framework, the N decision trees generate decision recommendations in N dimensions based on the decision basis and perform average aggregation to obtain the decision result, and N is a first preset value; A determining unit is configured to determine a target bucket expansion strategy from M bucket expansion strategies with predetermined priorities based on the decision result, where M is a second preset value; An executing unit is configured to perform a data redistribution operation on the financial data warehouse based on the target bucket expansion strategy.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored computer program, wherein the computer readable storage medium controls a device in which the computer readable storage medium is located to perform the integrated decision tree based bucket expansion regulation method in any one of claims 1 to 7 when the computer program is running.
10. An electronic device, comprising: The device comprises one or more processors and a memory, and the memory is configured to store one or more programs, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement the integrated decision tree based bucket expansion regulation method in any one of claims 1 to 7.
11. A computer program product, characterised in that, The device comprises computer instructions, wherein the computer instructions, when executed by a processor, implement the steps of the integrated decision tree based bucket expansion regulation method in any one of claims 1 to 7.